-
Notifications
You must be signed in to change notification settings - Fork 60
Pull requests: gittensor-ai-lab/sparkinfer
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
perf(qwen36): routed MoE prefill GEMM reads native quantized experts, no int8 materialize
#621
opened Jul 27, 2026 by
widecloud
Contributor
Loading…
1 task done
perf(qwen36): compacted paged-KV view for GQA-8 sink+window sparse decode @32k (+11.7%)
#619
opened Jul 24, 2026 by
rsnetworkinginc
Loading…
1 task done
perf(qwen36): GQA-fused prefill attention + native mma.sync MoE GEMM (+9% pp @4k)
#618
opened Jul 24, 2026 by
fansilas
Contributor
Loading…
1 task done
qwen36 MoE prefill: fix MOE_GPU device-tilemap correctness (#586) + int8 projection default
area:runtime
subsystem (emission weight 0.26)
needs-benchmark
Box ticked but decode before/after not filled with a real improvement — not evaluated
#601
opened Jul 23, 2026 by
James-CUDA
Contributor
•
Draft
1 task
server: emit GPU ttft/generation/decode_tps in stream usage
area:runtime
subsystem (emission weight 0.26)
hold
Maintainer override: never auto-merge this PR
#570
opened Jul 21, 2026 by
ai-hpc
Contributor
Loading…
1 of 4 tasks
ProTip!
Add no:assignee to see everything that’s not assigned.