Skip to content

Pull requests: vllm-project/vllm

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

[Bugfix][ROCm][MoE] Update AITER MXFP4 W4A16 tests to the renamed expert_mask bug Something isn't working rocm Related to AMD ROCm
#51980 opened Aug 12, 2026 by stefankoncarevic Contributor Loading…
4 tasks
[Bugfix] Release worker RPC payload before next dequeue bug Something isn't working
#51979 opened Aug 12, 2026 by shipiyouniao Loading…
[Security] Enforce server-side num_frames ceiling in VideoMediaIO merge multi-modality Related to multi-modality (#4194)
#51969 opened Aug 12, 2026 by jperezdealgaba Contributor Loading…
4 tasks done
[Perf][DSV4] Optimize global top-k index kernel with compile-time constants
#51967 opened Aug 12, 2026 by chaunceyjiang Collaborator Loading…
4 tasks
[XPU][Linear][MXFP4] enable torch asn mxfp4 linear backend on xpu intel-gpu Related to Intel GPU
#51965 opened Aug 12, 2026 by zufangzhu Contributor Loading…
[Bugfix][Model] Reuse CUDA segments when loading Inkling expert weights bug Something isn't working nvidia
#51962 opened Aug 12, 2026 by mo-ke-ke Loading…
4 tasks done
[Test] Add Thai coverage to detokenizer tests
#51961 opened Aug 12, 2026 by spped2000 Loading…
[XPU] Register KV offload mmap region as pinned host memory intel-gpu Related to Intel GPU
#51956 opened Aug 12, 2026 by chaojun-zhang Contributor Draft
[CI] Split Quantization job into three directory-based steps ci/build cpu Related to CPU backends documentation Improvements or additions to documentation nvidia quantization ready ONLY add when PR is ready to merge/full CI is needed
#51955 opened Aug 12, 2026 by khluu Member Loading…
[Perf][Qwen3.5] Avoid GDN decode gate copies qwen Related to Qwen models
#51954 opened Aug 12, 2026 by feednetinfra Draft
[Bugfix][MRV2] Fix DiffusionGemma runtime OOM via tiled logits projection bug Something isn't working mrv2 Model Runner V2 specific
#51953 opened Aug 12, 2026 by guan404ming Contributor Loading…
[Bugfix][Model] Fix Inkling NVIDIA sconv cache block-size mismatch bug Something isn't working nvidia
#51951 opened Aug 12, 2026 by fjosw Contributor Loading…
[Model] Enable LoRA support for tower and connector in Cosmos3-Edge documentation Improvements or additions to documentation
#51949 opened Aug 12, 2026 by charitarthchugh Loading…
11 tasks done
ProTip! Type g i on any issue or pull request to go back to the issue listing page.