Releases: huggingface/transformers
Release list
Release 5.17.0
Release v5.17.0
New Model additions
HYV4
Hy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters per
token. Each MoE layer holds 256 routed experts plus one always-active shared expert and routes every
token to 8 of them. The context window is 1M tokens.
The architecture combines four features:
- Multi-head Latent Attention (MLA) compresses keys and values into a low-rank latent
(kv_lora_rank) thatkv_b_projexpands back to one key/value per query head. - DeepSeek Sparse Attention (DSA) selects
index_topkkeys per query with a lightweight indexer.
Following IndexShare, only the layers marked"full"
inindexer_typesrun an indexer;"shared"layers reuse the previous full layer's selection. - Gated MLA with learnable attention sinks, where each head owns a sink logit that participates
in the softmax and contributes no value, as in GPT-OSS. - Independent Hyper-Connections (iHC) replace the plain residual path with
hc_multparallel
residual streams that are collapsed before, and redistributed after, every sublayer.
The implementation does not execute the multi-token prediction (MTP) layers. Released checkpoints
keep those weights so that other runtimes can use them for speculative decoding; they are ignored
at load time.
Links: Documentation
- Add h4 (#48473) by @ArthurZucker in #48473
VibeVoice
VibeVoice is a novel framework for synthesizing high-fidelity, long-form speech with multiple speakers by employing a next-token diffusion approach within a Large Language Model (LLM) structure. It's designed to capture the authentic conversational "vibe" and is particularly suited for generating audio content like podcasts and multi-participant audiobooks.
Links: Documentation
- Implement VibeVoice (#40546) by @pengzhiliang in #40546
NeoMME
NeoMME is a family of efficient 260M and 800M parameter multimodal-native multilingual foundation encoders from H Company. It processes multilingual text tokens and raw image patches in a single bidirectional Transformer encoder, without a separately pretrained vision tower or causal language model.
NeoMME-Retriever is a model fine-tuned from the NeoMME backbone for visual document retrieval with joint late-interaction and dense objectives. It takes text queries and documents (text or page screenshots) and produces multi-vector embeddings for MeanMaxSim scoring (late-interaction) and mean-pooled embeddings for cosine similarity (dense).
Links: Documentation
Fun-ASR-Nano
Fun-ASR-Nano is an 800M-parameter end-to-end speech recognition model developed by Alibaba DAMO Academy's FunAudioLLM team. It achieves state-of-the-art performance on Chinese, English, and Japanese ASR benchmarks while being significantly smaller than comparable models.
Key features are
- Chinese, English, and Japanese, including 7 Chinese dialects and 26 regional accents
- Hotword customization for domain-specific vocabulary
- Native punctuation output (no separate punctuation model needed)
Links: Documentation
KimiLinear
Kimi Linear is a hybrid linear attention architecture from Moonshot AI, introduced in
Kimi Linear: An Expressive, Efficient Attention Architecture.
At its core is Kimi Delta Attention (KDA), a refinement of Gated DeltaNet
that gives each key channel its own forget gate, so the recurrent state decays per channel instead of per head. KDA is
used in most layers; every fourth layer keeps a full-attention block that reuses DeepSeek-V3's Multi-head Latent
Attention (MLA), and the feed-forward blocks are DeepSeek-V3-style MoE with a shared expert.
Links: Documentation
Canary
Canary-1B-v2, a fast, robust multilingual model for Automatic Speech Recognition (ASR) and Speech-to-Text Translation (AST):
Canary reuses the Fast Conformer encoder from Parakeet (loaded through [ParakeetEncoder] / [ParakeetEncoderConfig]) and pairs it with a Transformer decoder that uses fixed sinusoidal positional embeddings, cross-attention to the encoder outputs and tied input/output embeddings. The task is selected through a decoder prompt prefix built by [CanaryProcessor] of the form <|startofcontext|> <|startoftranscript|> <|emo:undefined|> <source_lang> <target_lang> <pnc|nopnc> <|noitn|> <|notimestamp|> <|nodiarize|>, where source_lang == target_lang selects transcription and otherwise selects translation.
Links: Documentation
- model: Add NVIDIA Canary-1B-v2 to Transformers (#46825) by @harshaljanjani in #46825
NeuCodec
The NeuCodec model was proposed in Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates.
NeuCodec is a neural audio codec extending on XCodec2. It takes advantage of the following features:
- Finite Scalar Quantization (FSQ) quantisation resulting in a single codebook, making it ideal for downstream modeling with Speech Language Models.
- Trained with CC data such that there are no Non-Commercial data restrictions.
- At 50 tokens/sec and 16 bits per token, the overall bit-rate is 0.8kbps.
- The codec takes in 16kHz input and outputs 24kHz using an upsampling decoder.
- The FSQ encoding scheme allows for bit-level error resistance suitable for unreliable and noisy channels.
Links: Documentation
- Add support for NeuCodec (#47143) by @harryjulian in #47143
Breaking changes
Vision rotary embeddings (2D/3D) have been standardized into a unified RoPE frequency computation module, so users with custom vision models relying on attention-layer-level or model-specific RoPE grid interleaving logic must migrate to the new centralized modeling_rope_utils.py implementation.
- 🚨 Vision (2d/3d) rotary embeddings (#48105) by @zucchini-nlp
Generation
Generation improvements include a performance optimization that avoids unnecessary accelerator synchronization on every decode step (reducing per-step overhead), and a fix to prevent unconditional downloading of remote hub files during generation. Several correctness fixes were also applied, including enforcing auto-compile cache checks for encoder-decoder models, standardizing past_key_values naming in AfMoE, and resolving flaky export and integration test failures.
- [
Generate] Avoid unconditionally downloading remote hub file (#48620) by @vasqu in [#48620] - [generate] stop synchronizing the accelerator on every decode step (#47975) by @SunMarc in [#47975]
- [AfMoE] Standardize past_key_values argument naming across forward and generate (#48430) by @shenhuaqingshi in [#48430]
- Fix MTP generation test regex gate for escaped layer ignore keys (#48003) (#48262) by @Noxtimo in [#48262]
- fix(generation): Enforce the auto-compile cache check for encoder-decoder models (#48364) by @harshaljanjani in [#48364]
- [serge] Fix 4 integration tests for model
generationfailing withoutput_mismatch(list output differs (4)) (#48133) by @sergereview[bot] in [#48133] - [VibeVoice] Skip generate export tests (flaky) (#48396) by @ydshieh in [#48396]
Cache
Fixed several cache-related bugs, including a quantized cache issue in VibeVoice, incorrect rejection of non-static cache implementations in VoxtralRealtime, missing auto-compile cache checks for encoder-decoder models, and a silent failure when paged attention is called without a cache. Documentation was also updated to clarify ContinuousBatchingConfig usage and sliding window model limitations.
- vibevoice: fix bug for quant cache (#48487) by @kaixuanliu in [#48487]
- Fix VoxtralRealtime rejecting non-static cache implementations (#48082) by @jiqing-feng in [#48082]
- Raise when a paged attention forward is called with no cache (#48297) by @qgallouedec in [#48297]
- [docs] Pass ContinuousBatchingConfig and sliding window models (#48381) by @stevhliu in [#48381]
- Retry get_daily_ci_runs on stale GitHub API cache (#48374) by @ydshieh in [#48374]
Kernels
Kernel support was improved with fixes for nested FLA kernel imports when only fla-core is installed, a warning when hub-kernel functions silently fall back to slower pure-PyTorch reference implementations, and the abi...
Release v5.16.1
Release v5.16.1
This is a special release as we include GLM! (and a few small fixes)
GLM-5.3-Flash
GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.
GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute.
Links: Documentation
Small patch fixes
Mainly BC behavior for TP and pinning a hf kernel for security reasons 🤗
- Restore BC for the tensor-parallel API (#48300) by @ArthurZucker
- Fix kernel commit and repo paths for ESMFold2 (#48186) by @Rocketknight1
Full Changelog: v5.16.0...v5.16.1
Release: v5.16.0
Release v5.16.0
New Model additions
Qwen4-Exp
Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE).
GR is a Qwen-developed residual architecture that combines Hyper-Connection with GatedNorm. It mixes multiple residual streams with fine-grained elementwise gating before each attention and Mixture-of-Experts (MoE) block, then controls how much of the block output is injected back into each stream.
QSA uses multiple query heads to score compressed key blocks, selects the most relevant contiguous token blocks, and keeps the incomplete trailing block uncompressed. This block-level selection reduces indexing overhead and improves memory locality for long sequences. Combined with Gated DeltaNet, QSA makes Qwen4-Exp the first hybrid architecture to integrate linear and sparse attention, substantially improving inference efficiency for long-context workloads.
PLE enriches selected decoder layers with layer-specific lexical features derived from hashed token n-grams and a dilated depthwise convolution.
Links: Documentation
- Add Qwen4Exp model (#48337) by @Cyrilvallez in #48337
GraniteSpeech5
Granite Speech 5.0 Turbo CTC is a lightweight (~470M parameters) conformer encoder for automatic speech recognition, trained with Connectionist Temporal Classification (CTC) on BPE targets. It is a fast, encoder-only member of the Granite Speech family: transcription requires a single forward pass followed by greedy CTC decoding, with no autoregressive decoder.
Architecturally, it extends the Granite Speech conformer CTC encoder with:
-
Frame stacking + block-wise time subsampling: the feature extractor stacks pairs of log-mel(+delta) frames (2x), and the first two conformer blocks each subsample time by 2 through a stride-2 depthwise convolution (with a mean-pooled residual), for a total 8x time reduction at 10 ms mel hop.
-
Block attention with Shaw's relative positional embeddings: attention is computed over fixed-size blocks (the sequence is right-padded to a whole number of blocks, with padded frames masked out), using separate bias-free query/key/value projections.
-
Self-conditioned CTC: the CTC posteriors of the middle layer are projected and fed back into the hidden states, and the CTC head is shared between this mid-layer self-conditioning and the final prediction.
Links: Documentation
Step3p7
Step-3.7-Flash was proposed in Step 3.7 Flash by StepFun. It is a 198B-parameter sparse Mixture-of-Experts vision-language model, pairing a 196B-parameter MoE language backbone with a 1.8B-parameter vision encoder for native image understanding.
StepFun hasn't published a technical report for Step-3.7-Flash, so the details below are drawn from the released checkpoint's configuration rather than a paper.
- Sparse MoE decoder: all but the first 3 decoder layers route through a MoE block of 288 routed experts (top-8 per token) plus a single shared expert. The router scores experts with a sigmoid and a learned per-expert bias instead of an auxiliary load-balancing loss, the same strategy as DeepSeek-V3.
- Gated attention: each attention layer adds an extra projection whose sigmoid output gates the attention output per head, before the output projection — the same Gated Attention mechanism used in Qwen3-Next. A subset of layers use fewer heads and a sliding window instead of full attention.
- Multi-token prediction: some checkpoints ship extra decoder layers trained for multi-token prediction, which [
~GenerationMixin.generate] can use for speculative decoding viause_mtp=True. - Vision encoder: a SigLIP-style ViT with 2-D rotary position embeddings and a learned per-layer scale on the attention and MLP branches. Its output is downsampled 4x by two stride-2 convolutions before a linear projector maps it into the text model's hidden size.
- Dynamic image tiling: instead of a fixed tile grid, the image processor picks its tiling window from each image's own aspect ratio, producing one downscaled global view plus zero or more local high-resolution crops per image.
Links: Documentation
CohereCompass
CohereCompass is the base architecture for small, specialized (vision-)language models trained by Cohere.
Links: Documentation
ESMC and ESMFold2
ESMC and ESMFold2 are new state-of-the-art protein language and folding models from BioHub. ESMC is trained with a masked language modeling objective, and it can be easily transferred to sequence and token classification tasks for proteins. Checkpoints exist in various sizes, from 300M parameters up to 6B parameters. It works as a drop-in replacement for older ESM-2 and ESM-3 models, with significantly higher accuracy.
ESMFold2 is a state-of-the-art protein folding model which produces high accuracy predictions. It uses an iterated diffusion approach that is significantly different from the original ESMFold, offering huge improvements in accuracy for more complex structures.
Links: Documentation ESMC, Documentation ESMFold2
- Port ESMC and ESMFold2 to Transformers (#46419) by @Rocketknight1 in #46419
Breaking changes
The legacy tensor-parallel implementation has been replaced with a DTensor-native backend, so users relying on the previous TP API for inference or training must migrate to the new DTensor-based interface.
- 🚨 TP dtensor API inference + training (#47579) by @3outeille
attn_implementation="sdpa" dispatch is now properly supported for wav2vec2_conformer, wav2vec2-bert, and SeamlessM4T/v2 models, which may change initialization behavior for users who previously worked around this limitation.
- 🚨[wav2vec2] Support attn_implementation=sdpa dispatch (#46196) by @YangKai0616
FuyuProcessor no longer returns the image_patch_indices output, so any code that depends on this field must be updated to remove references to it.
- 🚨 Leftover processors (#47924) by @zucchini-nlp
Cache
Several cache-related bugs were fixed in this release, including an off-by-one error in the sliding window cache, Whisper speculative decoding cache corruption, CpmAnt use-cache failures, Qwen2.5-Omni/Qwen3-Omni-MoE generation with compilable caches, and compressed-tensors loading for KV-cache-only quantized models. Documentation was also added for cache token removal using negative values, and per-layer cache configuration support (allowing models to use different cache settings per layer) was introduced.
- Cpmant fix use cache (#48013) by @jiqing-feng in [#48013]
- [docs] Cache crop (#47950) by @stevhliu in [#47950]
- Revert "Support per-layer cache configuration and attention-mask selection" (#48175) by @Cyrilvallez in [#48175]
- Support per-layer cache configuration and attention-mask selection (#47901) by @eladsegal in [#47901]
- Fix Qwen2.5-Omni / Qwen3-Omni-MoE generation with a compilable cache (#47872) by @jiqing-feng in [#47872]
- Fix sliding window cache index off-by-one on wraparound (#47708) by @hameedibrh in [#47708]
- [Whisper] Fix speculative decoding: UnboundLocalError, cache corruption, and speed regression (#48000) by @ydshieh in [#48000]
- Fix compressed-tensors loading for KV-cache-only quantized models (#47904) by @kylesayrs in [#47904]
Generation
This release fixes several generation bugs across multiple models, including Whisper speculative decoding issues (UnboundLocalError, cache corruption, speed regression, and left-padded batch position IDs), broken image generation in Emu3, garbage output in OLMo/GPTNeoX, and Qwen2.5-Omni/Qwen3-Omni-MoE generation with compilable caches. Additionally, logit distributions for candidate generators using sampling are now aligned by returning logits after applying logit processors.
- [Whisper] Fix speculative decoding: preserve cleared suppress tokens through super().generate() (#48108) by @ydshieh in [#48108]
- [Whisper] Fix decoder position IDs for left-padded batches in longform generation (#48028) by @ydshieh in [#48028]
- [serge] Fix 2 integration tests for model
generationfailing withimport_or_config(other (2)) (#48061) by @sergereview[bot] in [#48061] - Align logit distributions for CandidateGenerators using sampling (#48007) by @Cyrilvallez in [#48007]
- [GPTNeoX] Fix post_processor not overridden when loading from pretrained (OLMo garbage generation) (#47988) by @ydshieh in [#47988]
- [emu3] 🦮 Black Labrador is back! Fix image generation broken since #37033 (#47948) by @ydshieh in [#47948]
At...
Patch release: v5.15.1
Patch release v5.15.1
This patch most notably solves a few issues with DFlash and MTP candidate generators, as well as an issue where images could sometimes not be processed on accelerator if using Lanczos filter.
It contains the following commits:
- Fix DFlash candidate token device mismatch with device_map="auto" (#47877) by @sywangyi and @Cyrilvallez
- Align logit distributions for CandidateGenerators using sampling (#48007) by @Cyrilvallez
- Fix MTP config when mlp_layer_types is absent (#48015) by @Cyrilvallez
- Fallback from 'lanczos' to 'bicubic' when on cuda (#48026) by @zucchini-nlp
- Fix gemma4 video to device (#47896) by @guarin
Release: v5.15.0
Release v5.15.0
New Model additions
Meta Muse Glimmer
Muse Glimmer, released today, is Meta’s new multimodal model, especially designed for agentic use cases. Distilled from Muse to 30B parameters, and released under the Apache 2.0 license, it can be deployed to local setups for privacy-aware applications such as coding, document analysis, personal assistants, Claw- or Hermes-like setups.
Muse Glimmer is a dense 30B parameter model consisting of:
- 2B ViT-style encoder for vision (Perception Encoder)
- 28B parameter text decoder
We're covering it in the following blogpost: http://hf.co/blog/muse-glimmer
GraniteMoeSWA & GraniteSWA
Links: Documentation
Links: Documentation
A.X-K1 & A.X-K2
Links: Documentation
Links: Documentation
Cosmos3 Edge
Links: Documentation
- Add Cosmos3 Edge model support (#47181) by @atharvajoshi10 in #47181
Breaking changes
Kernels are now opt-in rather than mandatory for linear attention models (Mamba, GDN, Conv-only, etc.), so users who relied on automatic kernel selection must explicitly enable kernels to maintain previous behavior.
The cache cropping API now only accepts negative values (relative offsets) instead of absolute sizes, so users calling crop methods directly must update their code to pass negative values accordingly.
- 🚨 [cache] Cropping can only be done with negative values (#47720) by @Cyrilvallez
T5 and its model family (MT5, LongT5, etc.) now support SDPA and other attention backends via ALL_ATTENTION_FUNCTIONS, meaning the default attention implementation may change and users relying on the previous eager-only path should explicitly set attn_implementation="eager" if needed.
- 🚨 Enable SDPA (and other attention backends) for T5 and propagate to the T5 family (#47014) by @jiqing-feng
Several small private helper functions (e.g., _is_url, _build_image_tokens) have been removed from multimodal processor files, so users or downstream libraries that imported these private functions directly must remove or replace those references.
- 🚨 Processors update the rest (#46556) by @zucchini-nlp
Attention
This release includes several attention fixes and improvements, including correcting Multi-Head Latent Attention (MLA) cache compression, optimizing Flash Attention max sequence length computation in vision models, and fixing bugs in CTRL flex-attention and SDPA prefill with position bias. Additional changes refactor linear attention models for better maintainability, make Gemma 4's heterogeneous attention config explicit, and improve MPS support via metal-flash-sdpa integration.
- [Fix] Fix multi-head latent attention (MLA) (#47761) by @remi-or in [#47761]
- Refactor all linear attention models to latest best standards for convolution (#47452) by @Cyrilvallez in [#47452]
- Allow metal-flash-sdpa for OpenAIPrivacyFilter on MPS (#46740) by @ArthurZucker in [#46740]
- Use new
per_layer_configfor Gemma 4 so that heterogeneous attention config is explicit (#47384) by @hmellor in [#47384] - add paged attention tests support for XPU (#47163) by @kaixuanliu in [#47163]
- Move
valuepadding into the attention interfaces that need it (#47451) by @hmellor in [#47451] - Simplify function dispatch for linear attention (#47450) by @Cyrilvallez in [#47450]
- Optimize flash attention max seqlen computation in vision attention (#47170) by @ShareLer in [#47170]
- Fix
BlockMaskcrash in CTRL flex-attention generation (#46854) by @jiqing-feng in [#46854] - [CB] Automatically switch attention implementation to flash (#47330) by @remi-or in [#47330]
- Fix sdpa prefill with position_bias (#47359) by @Cyrilvallez in [#47359]
Vision
Vision improvements in this release include performance optimizations such as faster image preprocessing for vision-language models (GLM4V, MiniMaxM3-VL, and others) by eliminating redundant tensor copies, and more efficient Flash Attention variable-length paths by precomputing maximum sequence lengths once per forward pass. Several bug fixes were also applied, including correcting dtype alignment in Kosmos2/Kosmos2_5 embedding merges, fixing a position-embedding initialization fallback in Phi4Multimodal, resolving PIL resize parity in Hunyuan-VL, and patching stop-sequence handling in the image-text-to-text pipeline.
- Modularize qwen-format vision processors (#47573) by @zucchini-nlp in [#47573]
- Update daily CI Docker image to torch 2.13.0 / CUDA 13.0 (#47738) by @ydshieh in [#47738]
- Align image feature dtype in kosmos2 and kosmos2_5 embedding merge (#47691) by @ in [#47691]
- Speed up image preprocessing for vision-language models (#47453) by @labAxiaoming in [#47453]
- Fix vision position-embedding init width fallback in Phi4Multimodal (#47509) by @ in [#47509]
- Fix Hunyuan-VL PIL image resize parity with reference preprocessing (#47233) by @IMvision12 in [#47233]
- Fix image-text-to-text stop_sequence handling (#47032) by @Sunt-ing in [#47032]
- Refactor image loading in tests to use load_test_image helper (#47218) by @LevelVoid in [#47218]
Generation
Several generation improvements and bug fixes were made, including enabling batched audio generation for Qwen2.5/3-Omni, allowing sliding window cache layers to work with speculative decoding, and fixing memory overhead from static cache persistence across generate() calls. Multiple model-specific bugs were also resolved, including crashes in KyutaiSpeechToText, MusicgenForCausalLM, CTRL flex-attention, and assisted decoding for EncoderDecoder cache and OlmoHybrid models.
- Align OlmoHybrid to use a native cache in generate (#47604) by @Cyrilvallez in [#47604]
- [generate] Stop setting the static cache as an attribute to save memory (#47731) by @Cyrilvallez in [#47731]
- Add support for batched Qwen2.5/3-Omni audio generation (#47186) by @IMvision12 in [#47186]
- [cache] Allow sliding window layers to be roll-backed for speculative decoding (#47447) by @Cyrilvallez in [#47447]
- Fix shape mismatch in KyutaiSpeechToText
generate()last window (#46952) by @jiqing-feng in [#46952] - Fix typo in
MusicgenForCausalLM.generate()(#46974) by @jiqing-feng in [#46974] - Fix assisted decoding for models with EncoderDecoder cache & OlmoHybrid (#47361) by @Cyrilvallez in [#47361]
Cache
Several cache-related bugs were fixed, including correcting NemotronH's missing "mlp" layer-type mapping, resolving recurrent-layer padding masks being skipped during chunked prefill and cache continuation for hybrid models, and fixing assisted decoding for models with EncoderDecoderCache and OlmoHybrid. Additional improvements include aligning OlmoHybrid to use a native cache, enabling sliding window layers to support speculative decoding rollback, and stopping the static cache from being stored as a model attribute to reduce unexpected memory overhead.
- [docs] MPS graph cache (#47304) by @stevhliu in [#47304]
- Fix NemotronH: Register
"mlp"in the cache layer-type mappings (#47535) by @qgallouedec in [#47535] - Fix recurrent-layer padding mask being skipped on continued forwards (chunked prefill, cache continuation) (#47087) by @abcgco in [#47087]
Kernels
kernels python package will very likely be a required dependency for transformers[torch] in the near future. This will help us deliver maximum performance to all users; kernels will only be downloaded from trusted publishers manually approved by the HF team. Please let us know of any issues you're facing beforehands so that we may solidify our integration.
Improved robustness of the kernels integration by refactoring function handling to use layer repos, fixing CI EROFS fallback patches for kernel downloads via HfApi, resolving a positional argument collision in causal_conv1d_fn, and bumping the FP8 kernels version to prevent NaNs.
- [conftest] Fix EROFS fallback for kernel downloads (correct interception point) (#47794) by @ydshieh in [#47794]
- [conftest] Fix EROFS fallback for kernel downloads via HfApi (#47791) by @ydshieh in [#47791]
- [
Kernels] Refactor function handling (#46883) by @vasqu in [#46883] - Kernels and loaders robustification (#47334) by @IlyasMoutawwakil in [#47334]
- Fix
causal_conv1d_fnpositionalactivationcolliding with hub kernel's `seq_...
Patch release: v5.14.1
Patch release v5.14.1
This patch solves a few issues which appeared when integrating Inkling model, most notably an issue affecting models using EncoderDecoderCache during assisted generation. It also fixes an issue that could appear during prefill with StaticCache and sdpa without padding for Inkling which uses a position_bias.
It contains the following commits:
- Fix sdpa prefill with position_bias (#47359) by @Cyrilvallez
- Fix assisted decoding for models with EncoderDecoder cache & OlmoHybrid (#47361) by @Cyrilvallez
- [FP8] Bump kernels version (#47344) by @vasqu
- Fix deepgemm on multiple devices (#47323) by @IlyasMoutawwakil
Release v5.14.0
Release v5.14.0
New Model additions
Inkling (fresh from Thinking Machines): 975B total, 41B active
- Add Inkling model #47347 by @molbap @Cyrilvallez @eustlb and @zucchini-nlp
Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and
generates text outputs. It is intended for use in English and other languages, and across
multiple coding languages. The model is designed to be used by developers building AI-
powered applications, including agentic and tool-use systems, coding assistants, chatbots, and
retrieval-augmented generation systems, and is suitable for general-purpose conversational
use, instruction-following, and other natural language and multimodal tasks. It is released with
open weights to support research, fine-tuning and integration into third-party products by
downstream developers.
TIPSv2
Links: Documentation
- Add TIPSv2 (#46347) by @Ternura143 in #46347
TIPSv2 DPT
Links: Documentation
- Add TIPSv2 (#46347) by @Ternura143 in #46347
🚨 Breaking changes
GPTNeoX now remaps embed_out to lm_head and GPTBigCode has _supports_attention_backend = True enabled for vLLM compatibility; users relying on the previous weight naming or attention backend behavior for these models should update their code accordingly.
Kernels
Several kernel-related fixes and improvements were made, including pinning the kernels dependency to a compatible version in the benchmark workflow, removing a deprecated package_name argument from LocalLayerRepository, and making the DeepGEMM Triton fallback more robust when CUDA_HOME is unset or misconfigured. Additionally, SDPA prefill was updated to leverage the FlashAttention kernel with StaticCache, yielding significant performance gains (up to 260% faster for large input sizes).
- Pin kernels to compatible version in benchmark workflow (#47339) by @tarekziade in [#47339]
- [Fix] Remove deprecated argument from
kernelscall (#47100) by @remi-or in [#47100] - [Fix] Make DeepGEMM triton fallback more robust (#47126) by @remi-or in [#47126]
- [sdpa] Allow prefill to use FA kernel with StaticCache (#47094) by @Cyrilvallez in [#47094]
Generation
Generation improvements include adding Multi-Token Prediction (MTP) decoding support, static ensemble verification for speculative decoding to improve draft token acceptance rates, and a fix for crashes in greedy assisted generation with different tokenizers. A misleading double-negative warning message for synced_gpus in continuous batching mode was also corrected.
- [generation] Fix misleading synced_gpus warning in continuous batching (#47158) by @Partha-Shankar in [#47158]
- [generate] Add proper MTP support (#46229) by @Cyrilvallez in [#46229]
- Fix crash in greedy assisted generation with different tokenizers (#46936) by @Sunt-ing in [#46936]
- [Generation] Add static ensemble verification for lossy speculative decoding (#45979) by @kasakh in [#45979]
Performance
Fixed a Flash Attention performance regression affecting models like Qwen3-VL and resolved a MoE decode optimization bug where the grouped-to-batched matrix multiplication switch was not applied to experts residing in submodels (e.g., VLMs with a nested text config).
- Fix FA performance regression (#47134) by @andreasgoulas in [#47134]
- Fix MoE decode optimization for experts living in a submodel (#47107) by @IlyasMoutawwakil in [#47107]
- Make doc builds faster (#47099) by @mishig25 in [#47099]
Cache
Cache dispatch logic was simplified by introducing explicit layer-type mappings for sliding and static layers, reducing complexity in cache routing. Additionally, fixes were made for read-only cache failures in CPU CI environments and for MPS graph cache growth during variable-length batch training on Apple Silicon.
- Fix CI read-only cache failures by patching cached_files in conftest (#47043) by @ydshieh in [#47043]
- trainer: clear MPS graph cache via torch_empty_cache_steps (#45818) by @anagnorisis2peripeteia in [#45818]
- [cache] Simplify cache dispatch based on layer_types (#47118) by @Cyrilvallez in [#47118]
Bugfixes and improvements
- ci: cover xet as well (runtime error) (#47338) by @tarekziade in [#47338]
- [docs] TokenizersBackend fallback (#47302) by @stevhliu in [#47302]
- Resolve continuous batching XPU availability checks at runtime (#47185) by @kaixuanliu in [#47185]
- [Nit] Add kernels_fallback_ok kwarg to is_flash_attn_N_available (#47318) by @remi-or in [#47318]
- [Nit] Add expectations for gemma4 tests on H100 (#47311) by @remi-or in [#47311]
- [docs] DeepGEMM requirements (#47324) by @stevhliu in [#47324]
- DeepGEMM shouldn't pad on SM90 (#47313) by @IlyasMoutawwakil in [#47313]
- Fix half-precision torch.compile crash in DETR-family sine position embeddings (#47238) by @David-Wu1119 in [#47238]
- Fix hardcoded paths in siglip checkpoint/vocab loading (#47178) by @XanxusCrypto in [#47178]
- Update AMD CI runner groups to amd-mi300 (#47307) by @Abdennacer-Badaoui in [#47307]
- Point to Gemma 4 model in Gemma4ForCausalLM docstring example (#47255) by @lefft in [#47255]
- Fix Qwen Omni batched text postprocessing (#47197) by @Sunt-ing in [#47197]
- Fix AqlmConfig error messages to say "int" instead of "float" (#47089) by @Sreekant13 in [#47089]
- Fix check for interactive stdout in _style function (#47283) by @smart8986 in [#47283]
- Fix get_json_schema crash on non-string docstring choices (#47072) by @Sreekant13 in [#47072]
- Make
MODEL_IDS_TO_TOKENIZERS_BACKENDcapture all DeepSeek R1 distills (#47296) by @hmellor in [#47296] - Update doc preprocessing regex to prevent ReDoS (#47187) by @WilliamRoyNelson in [#47187]
- Shard on read Dtensor aware (#46717) by @3outeille in [#46717]
- Switch AMD daily CI to mi300 runners (#47259) by @Abdennacer-Badaoui in [#47259]
- tests: reduce processor test memory usage by using tiny Hub checkpoints (#47213) by @ydshieh in [#47213]
- Torch compile backend defaults to "neuron" (#47035) by @michaelbenayoun in [#47035]
- Fix flash-attn Docker build broken by setuptools 83 removing pkg_resources (#47251) by @ydshieh in [#47251]
- Add heterogeneous config support (per-layer configuration) (#45333) by @eladsegal in [#45333]
- [fix] update integration test values (#47146) by @eustlb in [#47146]
- Fix DeepSpeed SP loss aggregation and LocalLayerRepository kwargs (#47073) by @sshivampeta in [#47073]
- tests only for the top 10 download models (#47244) by @3outeille in [#47244]
- Fix InputTokensDetails missing cache_write_tokens for openai>=2.34.0 (#47248) by @ydshieh in [#47248]
- Revert "Trigger a scheduled run" (#47249) by @ydshieh in [#47249]
- Remove Executorch from CI until latest version is supported and fully tested on CI env (#47242) by @IlyasMoutawwakil in [#47242]
- Be more defensive with
remap_legacy_layer_typesfor custom models (#47245) by @hmellor in [#47245] - Fix DistributedConfig docstring for unimplemented sp_plan (#47237) by @3outeille in [#47237]
- Switch mlinter to 0.1.2 (#47172) by @tarekziade in [#47172]
- Trigger a scheduled run (#47209) by @ydshieh in [#47209]
- Make executorch exporter tests always use xnnpack backend (#47201) by @tarekziade in [#47201]
- No agent PR descriptions (#45790) by @Rocketknight1 in [#45790]
- Clarify that max_steps is required for datasets without len (#47155) by @albertvillanova in [#47155]
- Cleanup pipelines, stop materializing generators (#47142) by @Rocketknight1 in [#47142]
- Fix device_map computation when the no_split_modules have different sizes (#47203) by @Cyrilvallez in [#47203]
- Add native FSDP2 module + migration (#46707) by @3outeille in [#46707]
- Fix experts implementation in two spots (#47097) by @remi-or in [#47097]
- [Fix] Remove old automatic cross attn pattern from output recorders (#47117) by @remi-or in [#47117]
- 🌐 [i18n-KO] Translate accelerator_selection.md to Korean (#47157) by @kkwjk2718 in [#47157]
- [i18n-KO] Translate optimum.md to Korean and fix Furiosa typo (#47156) by @kkwjk2718 in [#47156]
- [docs] fix curly quotes rendering to straight quotes (#47135) by @clijo in [#47135]
- Fix custom code which doesn't know about the new linear layer type names (#47174) by @hmellor in [#47174]
- Reject path traversal in the
transformers_weightsconfig field (#46890) by @LinZiyuu in [#46890] - [docs] Custom code conversion mapping (#47114) by @stevhliu in [#47114]
- Add exporters min version requirements and test skip (#47161) by @IlyasMoutawwakil in [#47161]
- tests: reduce processor test memory usage and use tiny test assets (#47168) by @ydshieh in [#47168]
- Clarify input device placement in the Quicktour inference example (#47136) by @samyuktahegde in [#47136]
- Extend continuous batching memory prediction test to XPU (#47159) by @sywangyi in [#47159]
- Fix case where
_LazyAutoMapping.registeris passed astrkey (#47148) by @hmellor in [#47148] - [docs] MoE decode switching (#47149) by @stevhliu in [#47149]
- add XPU output expectations for minicpm3 tests (#47092) by @kaixuanliu in [#47092]
- Dif...
Patch release v5.13.1
Patch release v5.13.1
This patch is focused on enabling transformers for the latest release of vllm!
Release v5.13.0
Release v5.13.0
New Model additions
KimiK 2.5, 2.6, and 2.7
This release includes the architecture for Kimi 2.5 which is used by 2.5-2.7:
Kimi K2.5 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration. The model was proposed in Kimi K2.5: Visual Agentic Intelligence and further improved in [Kimi K2.6: Advancing Open-Source Coding](Kimi K2.5: Visual Agentic Intelligence).
Kimi K2.5 achieves significant improvements on complex, end-to-end coding tasks, generalizing robustly across programming languages (Rust, Go, Python) and domains spanning front-end, DevOps, and performance optimization. The model is capable of transforming simple prompts and visual inputs into production-ready interfaces and lightweight full-stack workflows, generating structured layouts, interactive elements, and rich animations with deliberate aesthetic precision.
Links: Documentation
- Add new model: Kimi2-6 (#45630) by @zucchini-nlp in #45630
MiMo-V2-Flash
MiMo-V2-Flash is a Mixture-of-Experts (MoE) language model developed by the Xiaomi MiMo team. Designed to establish a new balance between long-context modeling capabilities and inference efficiency, the model is built for strong performance in complex reasoning and agentic tasks. Trained on 27T tokens with native 32k sequence lengths, MiMo-V2-Flash seamlessly supports an extended 256K context window while significantly reducing KV-cache storage compared to standard global attention models.
Links: Documentation
Nemotron 3.5 ASR
Nemotron 3.5 ASR is a 600M-parameter multilingual speech recognition model from NVIDIA, built for high-quality transcription in both low-latency streaming and high-throughput batch settings, with native punctuation and capitalization. For streaming, it offers configurable chunk sizes—80ms, 160ms, 560ms, and 1120ms, letting users trade off latency against accuracy to suit their application. Its cache-aware FastConformer-RNNT architecture is central to this capability: unlike traditional buffered streaming, which repeatedly reprocesses overlapping audio windows, the model processes only each new incoming chunk while reusing cached encoder context from prior chunks. This eliminates redundant computation, significantly improves efficiency, and minimizes end-to-end delay without sacrificing accuracy, making it well suited to real-time transcription workloads.
Links: Documentation
NemotronAsrStreaming
Nemotron ASR Streaming is a 600M-parameter English speech recognition model from NVIDIA, built for high-quality transcription in both low-latency streaming and high-throughput batch settings, with native punctuation and capitalization. For streaming, it offers configurable chunk sizes—80ms, 160ms, 560ms, and 1120ms, letting users trade off latency against accuracy to suit their application. Its cache-aware FastConformer-RNNT architecture is central to this capability: unlike traditional buffered streaming, which repeatedly reprocesses overlapping audio windows, the model processes only each new incoming chunk while reusing cached encoder context from prior chunks. This eliminates redundant computation, significantly improves efficiency, and minimizes end-to-end delay without sacrificing accuracy, making it well suited to real-time transcription workloads.
Links: Documentation
Qwen3 ASR
Qwen3 ASR is an automatic speech recognition model from Alibaba's Qwen team that combines a Whisper-style audio encoder with a Qwen3 language model decoder for speech-to-text transcription. The model supports automatic language detection and multilingual transcription.
A forced aligner model is also included. It can be used to timestamp a provided transcript and its audio. It uses the same audio encoder model with a classification head that predicts a word's length. This model can be used with the transcript from any ASR model (see the example below with Parakeet CTC).
Links: Documentation
- Qwen3 ASR and Forced Aligner (#43838) by @mbtariq82 in #43838
ZAYA
ZAYA1 is a 760M active / 8.4B total parameter MoE language model trained by Zyphra. It combines Compressed
Convolutional Attention (CCA), a nonlinear ZAYA1 router, and residual scaling.
Links: Documentation
VideoPrism
The VideoPrism model was proposed in the paper VideoPrism: A Foundational Visual Encoder for Video Understanding by Google DeepMind (blog post).
VideoPrism is a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. The model is pretrained on a large-scale heterogeneous corpus containing 36M high-quality video-caption pairs and 582M video clips with noisy parallel text (e.g., ASR transcripts). The pretraining approach improves upon masked autoencoding through global-local distillation of semantic video embeddings and a token shuffling scheme, enabling the model to focus primarily on the video modality while leveraging text associated with videos. VideoPrism achieves state-of-the-art performance on 31 out of 33 video understanding benchmarks across four broad task groups, from web video question answering to computer vision for science.
Links: Documentation
RADIO
RADIO (Reduce All Domains Into One) is a family of vision foundation models from NVIDIA trained by multi-teacher distillation (e.g. CLIP, DINOv2, SAM) into a single ViT backbone. It produces both an image-level summary embedding and dense spatial features, and supports variable input resolutions through a Cropped Position Embedding (CPE) patch generator.
Links: Documentation
- Add support for RADIO models (#46425) by @meatybobby in #46425
MiniCPM3
MiniCPM3 is the third-generation MiniCPM dense language model from OpenBMB. The 4B variant
(openbmb/MiniCPM3-4B) outperforms many 7B–9B open
models on standard benchmarks while remaining lightweight enough for on-device usage.
MiniCPM3 combines several architectural ideas:
- Multi-head Latent Attention (MLA) from DeepSeek-V2, which compresses the key/value cache
into a low-rank latent representation while still using rotary embeddings on a portion of the
query/key heads. - A standard SwiGLU MLP (no MoE).
- Three scalar scaling factors that govern signal flow:
scale_emb— scales input embeddings.scale_depth / sqrt(num_hidden_layers)— scales residual connections.hidden_size / dim_model_base— scales hidden states before the language model head.
Links: Documentation
Breaking changes
A broad set of modeling changes have been made to standardize layer declarations, mask/cache construction, and hybrid-attention handling, making many models cleanly exportable (ONNX, torch.export, ExecuTorch) and fullgraph-compilable — users relying on internal modeling APIs may need to update their code accordingly.
- 🚨 Modeling changes for export, compile, and hybrid-attention standardization (#46738) by @IlyasMoutawwakil
Attention masking for image tokens in Gemma 3/4 models has been fixed to correctly respect sliding window boundaries in local layers, which changes model behavior and may affect reproducibility of previous results.
- 🚨 [gemma 3/4] Fix bidirectional attention masking crossing sliding window boundaries (#46850) by @douglas-reid
The Expert Parallelism (E...
Patch release v5.12.1
Patch release v5.12.1
Updated the lower bound for PEFT and a fix for auto tokenizer to properly resolve the mistral tokenizer (when mistral-common is installed). This is similar to v.5.10.3 minus the fixes that were already included in the main release - vLLM will first target 5.10.3 🤗
- Fix
peftlower bound #46605 by @hmellor (#46605) - mistral common backend fix #46667 by @itazap (#46667)
Full Changelog: v5.12.0...v5.12.1