-
Notifications
You must be signed in to change notification settings - Fork 356
Pull requests: microsoft/onnxruntime-genai
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Infer multimodal feature width from model I/O shape
#2608
opened Sep 22, 2026 by
Xiaoyu Z (xiaoyu-work)
Member
Loading…
fix: size fp32 logits buffer with last-token shape during prefill
#2607
opened Sep 22, 2026 by
Zhong, GenMing (amd-genmingz)
Loading…
fix: allocate multimodal pipeline buffers on the session that writes them
#2605
opened Sep 21, 2026 by
Yuri Khrustalev (ykhrustalev)
Contributor
Loading…
feat: add LFM2-Audio / LFM2.5-Audio speech input and output support
#2601
opened Sep 20, 2026 by
Yuri Khrustalev (ykhrustalev)
Contributor
Loading…
feat(builder): support INT2 DFlash 2 drafters
#2600
opened Sep 20, 2026 by
Tianlei Wu (tianleiwu)
Contributor
Loading…
Standardize telemetry events and dependency policy
#2599
opened Sep 20, 2026 by
bmehta001
Contributor
Loading…
Reduce intermediate memory in quantized weight packing
#2596
opened Sep 20, 2026 by
野生の男 (Yasei-no-otoko)
Loading…
Load model builders and Transformers architectures lazily
#2594
opened Sep 18, 2026 by
xieofxie
Loading…
Feat: Add safetensors LoRA adapter loading to the C model benchmark
#2593
opened Sep 18, 2026 by
Chang Liu (cliu1003)
Contributor
Loading…
4 tasks
Add structured model builder configuration
#2592
opened Sep 18, 2026 by
Tianlei Wu (tianleiwu)
Contributor
Loading…
4 tasks done
Adding word level timestamps to Nemotron Streaming Pipeline
#2591
opened Sep 17, 2026 by
Mason Corey (themason2011)
Loading…
Centralize packed position-ID plane detection and validate generated ranges
#2589
opened Sep 17, 2026 by
aciddelgado
Contributor
Loading…
Rebuild AMDGPU device allocator per model instead of caching across models
#2584
opened Sep 17, 2026 by
KenLagos
Loading…
Add prefix caching to the dynamic engine
#2573
opened Sep 16, 2026 by
bmehta001
Contributor
Loading…
[WebGPU] Support paged attention models
#2570
opened Sep 15, 2026 by
Prathik Rao (prathikr)
Contributor
Loading…
Add an API to release cached device resources after model unload
#2569
opened Sep 15, 2026 by
aciddelgado
Contributor
Loading…
Add a native Gemma 3n multimodal processor
#2567
opened Sep 15, 2026 by
Anirudh Swaminathan (Anirudh-Swaminathan)
Loading…
Previous Next
ProTip!
Follow long discussions with comments:>50.