[kernels] Add Tests & CI for kernels #41765

MekkCyber · 2025-10-21T11:00:27Z

What does this PR do?

Adds tests for kernels, and proper daily CI, and slack notifications

run example : https://github.com/huggingface/transformers/actions/runs/18688016017/job/53285883834

HuggingFaceDocBuilderDev · 2025-10-21T11:14:15Z

The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update.

ydshieh

A few comments to address. I saw you triggered a few runs. Is this working well?

ydshieh · 2025-10-21T14:22:05Z

.github/workflows/self-scheduled.yml

+      - name: Install kernels
+        working-directory: /transformers
+        run: python3 -m pip install -U kernels


we should move this before the step

name: Show installed libraries and their versions

good catch! thanks

.github/workflows/self-scheduled.yml

ydshieh · 2025-10-21T14:26:28Z

tests/kernels/test_kernels.py

+        # Free accelerator memory/cache and trigger GC
+        gc.collect()
+        backend_empty_cache(torch_device)
+        gc.collect()


We could simply use the cleanup

from transformers.testing_utils

like

cleanup(torch_device, gc_collect=True)

ydshieh · 2025-10-21T14:27:20Z

tests/kernels/test_kernels.py

+        self.model_kernelized = AutoModelForCausalLM.from_pretrained(
+            self.model_id, use_kernels=True, device_map=torch_device
+        )
+        self.model_not_kernelized = AutoModelForCausalLM.from_pretrained(
+            self.model_id, use_kernels=False, device_map=torch_device


is this loading fast enough? We can probably do it once in setupClass

Yes it's really fast, since we use a 1B model, and it's cached for all runs, but setupClass seems like a good way to do it too

then it's up to you, fine for me :)

but we still need to add @slow

yes done, I added @slow and used setUpClass

ydshieh · 2025-10-21T14:29:10Z

tests/kernels/test_kernels.py

test class or method that use real model checkpoint should use @slow.

If all test methods in a class quality slow in this sense, you can add @slow to the class and not each method

ydshieh · 2025-10-21T14:30:09Z

tests/kernels/test_kernels.py

Maybe also ask Daniel to have a look here too.

ydshieh · 2025-10-22T14:18:59Z

tests/kernels/test_kernels.py

+        )
+        cls.input = "Hello"
+
+    def tearDown(self):


this should be tearDownClass now.

we should still keep tearDown with only cleanup(torch_device, gc_collect=True).

ydshieh · 2025-10-22T14:19:30Z

tests/kernels/test_kernels.py

+        # Clear any temporary kernel module cache entries populated by tests
+        try:
+            keys_to_remove = [
+                k for k, v in list(_KERNEL_MODULE_MAPPING.items()) if v is None or isinstance(v, types.ModuleType)
+            ]
+            for k in keys_to_remove:
+                _KERNEL_MODULE_MAPPING.pop(k, None)
+        except Exception:
+            pass


Could you explain why we needs this?

when using kernels _KERNEL_MODULE_MAPPING gets populated by kernels, we are just cleaning it here for other tests

ydshieh

Still prefer @danieldk to take a look on the new test file (the logic of the testing kernels)

But LGTM from my side for testing

* first commit * add tests * add kernel config * add more tests * add ci * small fix * change branch name * update tests * nit * change test name * revert jobs * addressing review * reenable all jobs * address second review

* remove attributes and add all missing sub processors to their auto classes * remove all mentions of .attributes * cleanup * fix processor tests * fix modular * remove last attributes * fixup * fixes after merge * fix wrong tokenizer in auto florence2 * fix missing audio_processor + nits * Override __init__ in NewProcessor and change hf-internal-testing-repo (temporarily) * fix auto tokenizer test * add init to markup_lm * update CustomProcessor in custom_processing * remove print * nit * fix test modeling owlv2 * fix test_processing_layoutxlm * Fix owlv2, wav2vec2, markuplm, voxtral issues * add support for loading and saving multiple tokenizer natively * remove exclude_attributes from save_pretrained * Run slow v2 (#41914) * Super * Super * Super * Super --------- Co-authored-by: ydshieh <[email protected]> * Fix `detectron2` installation in docker files (#41975) * detectron2 - part 1 * detectron2 - part 2 --------- Co-authored-by: ydshieh <[email protected]> * Fix `autoawq[kernels]` installation in quantization docker file (#41978) fix autoawq[kernels] Co-authored-by: ydshieh <[email protected]> * add support for saving encoder only so any parakeet model can be loaded for inference (#41969) * add support for saving encoder only so any decoder model can be loaded Signed-off-by: nithinraok <[email protected]> * use convolution_bias * convert modular * convolution_bias in convertion script --------- Signed-off-by: nithinraok <[email protected]> Co-authored-by: Eustache Le Bihan <[email protected]> Co-authored-by: eustlb <[email protected]> * Use indices as position_ids in modernebert (#41789) * Use indices as position_ids in modernebert * Move position_ids init to the branch * test tensor parallel: make tests for dense model more robust (#41968) * make test forward and backward more robust * refactor compile part of test tensor parallel * linting * pass rank around instead of calling it over and over * Run slow v2 (#41914) * Super * Super * Super * Super --------- Co-authored-by: ydshieh <[email protected]> * Fix `detectron2` installation in docker files (#41975) * detectron2 - part 1 * detectron2 - part 2 --------- Co-authored-by: ydshieh <[email protected]> * Fix `autoawq[kernels]` installation in quantization docker file (#41978) fix autoawq[kernels] Co-authored-by: ydshieh <[email protected]> * add support for saving encoder only so any parakeet model can be loaded for inference (#41969) * add support for saving encoder only so any decoder model can be loaded Signed-off-by: nithinraok <[email protected]> * use convolution_bias * convert modular * convolution_bias in convertion script --------- Signed-off-by: nithinraok <[email protected]> Co-authored-by: Eustache Le Bihan <[email protected]> Co-authored-by: eustlb <[email protected]> --------- Signed-off-by: nithinraok <[email protected]> Co-authored-by: Yih-Dar <[email protected]> Co-authored-by: ydshieh <[email protected]> Co-authored-by: Nithin Rao <[email protected]> Co-authored-by: Eustache Le Bihan <[email protected]> Co-authored-by: eustlb <[email protected]> * fix: dict[RopeParameters] to dict[str, RopeParameters] (#41963) * docs: add continuous batching page (#41847) * docs: add continuous batching page * docs(cb): add `generate_batch` example * docs(cb): add `opentelemtry` and `serving` section * feat: add `TODO` note about opentelemetry dependency * docs(cb): add supported features * docs(cb): add unsupported features * docs(cb): add `ContinuousBatchingManager` example * docs(cb): x reference CB in optimizing inference * Fix `torchcodec` version in quantization docker file (#41988) check Co-authored-by: ydshieh <[email protected]> * [kernels] Add Tests & CI for kernels (#41765) * first commit * add tests * add kernel config * add more tests * add ci * small fix * change branch name * update tests * nit * change test name * revert jobs * addressing review * reenable all jobs * address second review * Move the Mi355 to regular docker (#41989) * Move the Mi355 to regular docker * Disable gfx950 compilation for FA on AMD * More data in benchmarking (#41848) * Reduce scope of cross-generate * Rm generate_sall configs * Workflow benchmarks more * Prevent crash when FA is not installed * fix (CI): Refactor SSH runners (#41991) * Change ssh runner type * Add wait step to SSH runner workflow * Rename wait step to wait2 in ssh-runner.yml * Remove wait step from ssh-runner.yml Removed the wait step from the SSH runner workflow. * Update runner type for single GPU A10 instance * Update SSH runner version to 1.90.3 * Add sha256sum to ssh-runner workflow * Update runner type and remove unused steps * fix 3 failed test cases for video_llama_3 model on Intel XPU (#41931) * fix 3 failed test cases for video_llama_3 model on Intel XPU Signed-off-by: Liu, Kaixuan <[email protected]> * update Signed-off-by: Liu, Kaixuan <[email protected]> * adjust format Signed-off-by: Liu, Kaixuan <[email protected]> * update code Signed-off-by: Liu, Kaixuan <[email protected]> --------- Signed-off-by: Liu, Kaixuan <[email protected]> * Integrate colqwen2.5 using colqwen2 modelling code (#40600) * adding option for 2.5 * minor - arg in conversion script * getting started on modelling.py * minor - shouldve been using modular * adressing comments + fixing datatype/device _get method * minor * commiting suggestion Co-authored-by: Yoni Gozlan <[email protected]> * docs + first test * ruff fix * minor fix * ruff fix * model fix * model fix * fine-grained check, with a hardcoded score from the original Hf implementation. * minor ruff * update tests values with CI hardware * adding 2.5 to conversion script * Apply style fixes --------- Co-authored-by: Sahil Kabir <[email protected]> Co-authored-by: Yoni Gozlan <[email protected]> Co-authored-by: yonigozlan <[email protected]> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> * Fixed wrong padding value in OWLv2 (#41938) * Update image_processing_owlv2_fast.py fixed padding value * fixed padding value * Change padding constant value from 0.5 to 0.0 * Fixed missed padding value in modular_owlv2.py --------- Co-authored-by: Yoni Gozlan <[email protected]> * Fix `run slow v2`: empty report when there is only one model (#42002) fix Co-authored-by: ydshieh <[email protected]> * [kernels] change import time in KernelConfig (#42004) * change import time * style * DOC Fix typo in argument name: pseudoquant (#41994) The correct argument name is pseudoquantization. Since there is no error on passing wrong arguments name (which is arguably an anti-pattern), this is difficult for users to debug. * Fix `torch+deepspeed` docker file (#41985) * fix * delete --------- Co-authored-by: ydshieh <[email protected]> * Correct syntax error in trainer.md (#42001) A comma is missing between two parameters in the signature of compute_loss function. * Reduce the number of benchmark in the CI (#42008) Changed how benchmark cfgs are chosen * Fix continuous batching tests (#42012) * Fix continuous batching tests * make fixup * add back `logging_dir` (#42013) * add back * Apply style fixes --------- Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> * Fix issue with from pretrained and kwargs in image processors (#41997) * accept kwargs in image proc from_pretrained * only use kwargs that are in cls.valid_kwargs * remove specific logic for _from_auto * add image_seq_length to Images_kwargs for backward compatibility * fix missing image kwargs in pix2struct * Fix default image_rows and image_cols initialization in Idefics3 and SmolVLM processors (#41871) * Fix default image_rows and image_cols initialization in Idefics3 and SmolVLM processors * Fix default initialization of image_rows and image_cols in Idefics3 and SmolVLM processors * Add GLPNImageProcessorFast (#41725) * Add GLPNImageProcessorFast for torch backend * Address review feedback - Simplified to_dict() method - Keep tensors as torch instead of converting to numpy for heterogeneous shapes - Removed unnecessary shape guards in post_process_depth_estimation - Improved variable names (tgt -> target_size, d -> resized) - Removed unnecessary GLPNImageProcessorKwargs class * Address review feedback - Simplified to_dict() method - Keep tensors as torch instead of converting to numpy for heterogeneous shapes - Removed unnecessary shape guards in post_process_depth_estimation - Improved variable names (tgt -> target_size, d -> resized) - Removed unnecessary GLPNImageProcessorKwargs class * commits after 2nd review * Address all review feedback and add explicit batched test - Simplified to_dict() with descriptive variable names (d->output_dict) - Fixed resize operation: changed from crop to proper resize with interpolation - Added padding for heterogeneous batch shapes in both slow and fast processors - Fused rescale and normalize operations for efficiency - Improved all variable names (tgt->target_size, d->depth_4d->resized) - Added GLPNImageProcessorKwargs class in slow processor and imported in fast - Renamed test_equivalence_slow_fast to test_slow_fast_equivalence - Added explicit test_slow_fast_equivalence_batched test - All 20 tests passing * using padding from utils * simplify glpn image processor fast * fix docstring --------- Co-authored-by: yonigozlan <[email protected]> Co-authored-by: Yoni Gozlan <[email protected]> * add fuyu fast image processors (#41817) * added fast processor for fuyu (#36978) * updated docs for fuyu model (#36978) * updated test_image_processing and image_processing_fuyu_fast * updated fuyu.md and image_processing_fuyu_fast (#36978) * updated test_image_processing_fuyu (#36978) * formatted image_processing_fuyu_fast and test_image_processing_fuyu (#36978) * updated tests and fuyu fast image processing (#36978) * Merge branch 'fuyu-fast-image-processors' of https://github.com/DeXtAr47-oss/transformers into fuyu-fast-image-processors * fixed format (#36978) * formatted files (#36978) * formatted files * revert unnecessary changes * clean up and process by group --------- Co-authored-by: yonigozlan <[email protected]> * [kernels] Fix XPU layernorm kernel (#41583) * fix * add comment * better fix * style * Update src/transformers/modeling_utils.py Co-authored-by: Marc Sun <[email protected]> --------- Co-authored-by: Marc Sun <[email protected]> * [v5] Deprecate Text2Text and related pipelines (#41996) * Deprecate Text2Text and related pipelines * Try a restructure * make fixup * logging -> logger * [FPQuant] MXFP8 and MXFP4 backwards support (#41897) * FP-Quant backwards * fp-quant v0.3.0 docker * availability version bump * fp_quant==0.3.1 * fp_quant v0.3.2 * add working auto_docstring for processors * add auto_docstring to processors first part * add auto_docstring to processors part 2 * modifs after review * fully working auto_docstring and check_docstring with placeholder docstrings * Working check_docstrings for Typed dicts * Add recurring processor args to auto_docstring and add support for removing redundant docstring and placeholders * replace placeholders with real docstrings * fix copies * fixup * remove unwanted changes * fix unprotected imports * Fix unprotected imports * fix unprotected imports * Add __call__ to all docs of processors * nits docs --------- Signed-off-by: nithinraok <[email protected]> Signed-off-by: Liu, Kaixuan <[email protected]> Co-authored-by: Yih-Dar <[email protected]> Co-authored-by: ydshieh <[email protected]> Co-authored-by: Nithin Rao <[email protected]> Co-authored-by: Eustache Le Bihan <[email protected]> Co-authored-by: eustlb <[email protected]> Co-authored-by: Rémi Ouazan <[email protected]> Co-authored-by: Ferdinand Mom <[email protected]> Co-authored-by: Ryan Mullins <[email protected]> Co-authored-by: Luc Georges <[email protected]> Co-authored-by: Mohamed Mekkouri <[email protected]> Co-authored-by: Guillaume LEGENDRE <[email protected]> Co-authored-by: kaixuanliu <[email protected]> Co-authored-by: Sahil Kabir <[email protected]> Co-authored-by: Sahil Kabir <[email protected]> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: James <[email protected]> Co-authored-by: Benjamin Bossan <[email protected]> Co-authored-by: Yacklin Wong <[email protected]> Co-authored-by: Matt <[email protected]> Co-authored-by: Marc Sun <[email protected]> Co-authored-by: MilkClouds <[email protected]> Co-authored-by: ARAVINDHAN T <[email protected]> Co-authored-by: Pritam Das <[email protected]> Co-authored-by: Andrei Panferov <[email protected]>

MekkCyber force-pushed the add_tests_for_kernels branch from 8eda347 to 32a4773 Compare October 21, 2025 11:04

MekkCyber mentioned this pull request Oct 21, 2025

[kernels] Add version to function mapping #41685

Merged

MekkCyber added 8 commits October 21, 2025 11:30

first commit

eb61415

add tests

36ea8a2

add kernel config

6441a0d

add more tests

b2a6e25

add ci

d671ef5

small fix

d9dfd7f

change branch name

32310d8

update tests

4eda7fe

MekkCyber force-pushed the add_tests_for_kernels branch from c57e641 to 4eda7fe Compare October 21, 2025 11:30

MekkCyber added 3 commits October 21, 2025 11:35

nit

b527815

change test name

2a680e5

revert jobs

fcfd640

MekkCyber requested review from ArthurZucker and ydshieh October 21, 2025 11:55

ydshieh reviewed Oct 21, 2025

View reviewed changes

MekkCyber added 2 commits October 21, 2025 14:52

addressing review

a09d2cb

reenable all jobs

6d64718

ydshieh reviewed Oct 22, 2025

View reviewed changes

address second review

24ae781

ydshieh approved these changes Oct 23, 2025

View reviewed changes

MekkCyber merged commit a623cda into main Nov 3, 2025
24 checks passed

MekkCyber deleted the add_tests_for_kernels branch November 3, 2025 15:36

[kernels] Add Tests & CI for kernels #41765

[kernels] Add Tests & CI for kernels #41765

Uh oh!

Conversation

MekkCyber commented Oct 21, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

What does this PR do?

Uh oh!

HuggingFaceDocBuilderDev commented Oct 21, 2025

Uh oh!

ydshieh left a comment

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

ydshieh left a comment

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

MekkCyber commented Oct 21, 2025 •

edited

Loading