-
Notifications
You must be signed in to change notification settings - Fork 308
Pull requests: vllm-project/tpu-inference
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Autobump tokamax pin
ready
ONLY add when PR is ready to merge/full CI is needed
#3535
opened Sep 4, 2026 by
vanbasten23
Collaborator
Loading…
Make raiden weight-sync bind failures diagnosable
#3534
opened Sep 4, 2026 by
lokic233
Contributor
Loading…
Route dense Qwen3.5/3.8 to the vLLM implementation
#3533
opened Sep 4, 2026 by
lokic233
Contributor
Loading…
Add hybrid kvcache e2e to buildkite
ready
ONLY add when PR is ready to merge/full CI is needed
#3530
opened Sep 4, 2026 by
pv97
Collaborator
Loading…
[Spec Decode] Support Muse Glimmer DFlash checkpoints
#3528
opened Sep 4, 2026 by
lming2001
Loading…
3 of 4 tasks
build(docker): build and package vllm-rs Rust frontend for TPU
#3523
opened Sep 3, 2026 by
aojea
Loading…
[Perf/Fix] Flax loader: keep kv-cache output sharding instead of forcing one attention spec
#3519
opened Sep 2, 2026 by
wenxindongwork
Collaborator
•
Draft
Decoupled Dual Block Pools for Hybrid Mamba/GDN Prefix Caching in Align Mode
ready
ONLY add when PR is ready to merge/full CI is needed
#3517
opened Sep 2, 2026 by
pv97
Collaborator
Loading…
3 tasks done
[Perf/Fix] Configure heterogeneous KV cache sharding for GDN/Mamba in model_loader
ready
ONLY add when PR is ready to merge/full CI is needed
#3515
opened Sep 2, 2026 by
sierraisland
Collaborator
Loading…
[PCP] Move the host-side preprocessing into runner/pcp_utils.py
#3514
opened Sep 2, 2026 by
bhuvanpkaruturi
Collaborator
•
Draft
[Fix] Preserve layer order in KV cache slots for sequential decoders (MaxText)
ready
ONLY add when PR is ready to merge/full CI is needed
#3512
opened Sep 2, 2026 by
sierraisland
Collaborator
Loading…
[Sampler] Add opt-in vocab-sharded sampling (USE_VOCAB_SHARDED_SAMPLING)
ready
ONLY add when PR is ready to merge/full CI is needed
#3511
opened Sep 2, 2026 by
gutianyu-google
Contributor
Loading…
[Runner] Cap the request ids attached to the execute_model trace annotation
ready
ONLY add when PR is ready to merge/full CI is needed
#3510
opened Sep 2, 2026 by
gutianyu-google
Contributor
Loading…
[continue_decode] Skip the per-step EOS reduction when the EOS check is off
ready
ONLY add when PR is ready to merge/full CI is needed
#3509
opened Sep 2, 2026 by
gutianyu-google
Contributor
Loading…
Dedicated mamba block pools for GDN prefix caching (align mode), no vLLM change
#3508
opened Sep 2, 2026 by
wenxindongwork
Collaborator
•
Draft
[FIX] Fix MoE GMM FP8 accumulator verification error and bump qwix requirement (Maxtext)
ready
ONLY add when PR is ready to merge/full CI is needed
#3507
opened Sep 2, 2026 by
sierraisland
Collaborator
Loading…
[WIP] Debug PCP accuracy gap (gsm8k 0.80 vs 0.96) + fix PCP mamba boot crash
#3502
opened Sep 1, 2026 by
wenxindongwork
Collaborator
•
Draft
1 of 3 tasks
[Bugfix] Tolerate non-resizable unquantized weight storage
#3501
opened Sep 1, 2026 by
GianluigiVitale
Loading…
5 tasks done
feat(pipeline_parallel): support multi-host Ray pipeline parallelism
#3498
opened Aug 31, 2026 by
aashishrampal-lab
Contributor
Loading…
feat(pipeline_parallel): support single-host pipeline parallelism for DeepSeek-V4
#3497
opened Aug 31, 2026 by
aashishrampal-lab
Contributor
Loading…
feat(pipeline_parallel): support single-host pipeline parallelism for Llama
#3496
opened Aug 31, 2026 by
aashishrampal-lab
Contributor
Loading…
feat(loader): add layerwise streaming model loader and memory management
#3495
opened Aug 31, 2026 by
aashishrampal-lab
Contributor
Loading…
feat(quantization): support MOE_REQUANTIZE_WEIGHT_DTYPE and MOE_REQUANTIZE_BLOCK_SIZE in MXFP4 MoE
#3494
opened Aug 31, 2026 by
aashishrampal-lab
Contributor
Loading…
perf(sparsecore): enable dense gather reduce for v6e via FP32 intermediate buffer
#3493
opened Aug 31, 2026 by
prishajain1
Loading…
Previous Next
ProTip!
Type g p on any issue or pull request to go back to the pull request listing page.