-
-
Notifications
You must be signed in to change notification settings - Fork 20.6k
Pull requests: vllm-project/vllm
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[Core] Add per-request prefix-cache write policy
frontend
performance
Performance-related issues
#51981
opened Aug 12, 2026 by
NancyFyong
•
Draft
[Bugfix][ROCm][MoE] Update AITER MXFP4 W4A16 tests to the renamed expert_mask
bug
Something isn't working
rocm
Related to AMD ROCm
#51980
opened Aug 12, 2026 by
stefankoncarevic
Contributor
Loading…
4 tasks
[Bugfix] Release worker RPC payload before next dequeue
bug
Something isn't working
#51979
opened Aug 12, 2026 by
shipiyouniao
Loading…
[WIP][Feature] Add layerwise KV cache transfer support with mooncake session API
kv-connector
#51978
opened Aug 12, 2026 by
ClockOfDestiny
•
Draft
[CI][AMD][Disagg] Fix Kimi K2.5/K2.6 MXFP4 MLA backends on ROCm nightly for accuracy eval (disable aiter OPUS FMHA)
ci/build
kimi
rocm
Related to AMD ROCm
#51976
opened Aug 12, 2026 by
avininjamay8
Contributor
Loading…
feat: Add support for profile_prefix payload in HTTP /start_profile endpoint
frontend
#51974
opened Aug 12, 2026 by
rushabh-46
Loading…
3 of 4 tasks
[Bugfix] Fix mismatched logger format args and enable ruff PLE1205/PLE1206
bug
Something isn't working
kv-connector
#51973
opened Aug 12, 2026 by
rajathpi
Loading…
[BugFix] Clamp indexer scheduler-metadata seq_lens to avoid OOB read in DeepGEMM paged-MQA metadata kernel
bug
Something isn't working
#51972
opened Aug 12, 2026 by
fergusfinn
Contributor
Loading…
[Security] Enforce server-side num_frames ceiling in VideoMediaIO merge
multi-modality
Related to multi-modality (#4194)
#51969
opened Aug 12, 2026 by
jperezdealgaba
Contributor
Loading…
4 tasks done
[Perf][DSV4] Optimize global top-k index kernel with compile-time constants
#51967
opened Aug 12, 2026 by
chaunceyjiang
Collaborator
Loading…
4 tasks
[XPU][Linear][MXFP4] enable torch asn mxfp4 linear backend on xpu
intel-gpu
Related to Intel GPU
#51965
opened Aug 12, 2026 by
zufangzhu
Contributor
Loading…
[Misc] Point the unquantized-MoE backend error at the speculative-config fix
quantization
#51960
opened Aug 12, 2026 by
spped2000
Loading…
[Build] DeepGEMM pin has no SM120 kernels: family-12 Blackwell cannot run hyperconnections
ci/build
#51959
opened Aug 12, 2026 by
Mirrdhyn
Loading…
[Feature] Turboquant sliding window attn support
quantization
#51958
opened Aug 12, 2026 by
simondanielsson
Contributor
•
Draft
4 tasks
[XPU] Register KV offload mmap region as pinned host memory
intel-gpu
Related to Intel GPU
#51956
opened Aug 12, 2026 by
chaojun-zhang
Contributor
•
Draft
[CI] Split Quantization job into three directory-based steps
ci/build
cpu
Related to CPU backends
documentation
Improvements or additions to documentation
nvidia
quantization
ready
ONLY add when PR is ready to merge/full CI is needed
#51955
opened Aug 12, 2026 by
khluu
Member
Loading…
[Perf][Qwen3.5] Avoid GDN decode gate copies
qwen
Related to Qwen models
#51954
opened Aug 12, 2026 by
feednetinfra
•
Draft
[Bugfix][MRV2] Fix DiffusionGemma runtime OOM via tiled logits projection
bug
Something isn't working
mrv2
Model Runner V2 specific
#51953
opened Aug 12, 2026 by
guan404ming
Contributor
Loading…
NIXL: Use int32 array for indices to avoid intermediate conversion
kv-connector
#51952
opened Aug 12, 2026 by
iyastreb
Contributor
Loading…
[Bugfix][Model] Fix Inkling NVIDIA sconv cache block-size mismatch
bug
Something isn't working
nvidia
#51951
opened Aug 12, 2026 by
fjosw
Contributor
Loading…
[Model] Enable LoRA support for tower and connector in Cosmos3-Edge
documentation
Improvements or additions to documentation
#51949
opened Aug 12, 2026 by
charitarthchugh
Loading…
11 tasks done
Previous Next
ProTip!
Type g i on any issue or pull request to go back to the issue listing page.