forked from ggml-org/llama.cpp
-
-
Notifications
You must be signed in to change notification settings - Fork 406
Pull requests: TheTom/llama-cpp-turboquant
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
cuda: TurboQuant TQ4_1S decode optimisations
CUDA
ggml
#338
opened Aug 31, 2026 by
jasstrong
Loading…
[2/2] cuda: TQ4_1S MMQ prefill on RDNA3 and RDNA2
CUDA
ggml
#337
opened Aug 31, 2026 by
jasstrong
Loading…
[1/2] cuda: native TQ4_1S MMQ prefill on CDNA
CUDA
ggml
#336
opened Aug 31, 2026 by
jasstrong
Loading…
[2/2] test: cover TurboQuant MUL_MAT_ID at real MoE model shapes
CUDA
ggml
testing
#335
opened Aug 31, 2026 by
jasstrong
Loading…
[1/2] cuda: check mmvf applies before the AMD MUL_MAT_ID shortcut
CUDA
ggml
#334
opened Aug 31, 2026 by
jasstrong
Loading…
examples: add TurboAnchorKV evaluation harness
documentation
Improvements or additions to documentation
examples
#333
opened Aug 31, 2026 by
TheTom
Owner
Loading…
hip: VEC flash-attn for D=512 (Gemma 4) on ROCm with quantized KV
ggml
Nvidia GPU
#156
opened May 24, 2026 by
cclecle
Loading…
vulkan: add TurboQuant KV cache support and optimized turbo mat-vec paths
ggml
Vulkan
#140
opened May 10, 2026 by
Fenix46
Loading…
fix(qwen35): support Qwen3.5:9B loading from Ollama GGUF
model
#135
opened May 8, 2026 by
Jordan-HS
Loading…
vendor: bump cpp-httplib to 0.43.2 (openssl 4.0.0 fix)
python
script
#121
opened May 4, 2026 by
TheTom
Owner
Loading…
1 of 3 tasks
HIP mixed TurboQuant vec FA on gfx900/gfx906
build
ggml
Nvidia GPU
#99
opened Apr 21, 2026 by
2bigO
Loading…
fix: HIP/ROCm compatibility — check cudaMemcpyToSymbol errors, guard …
ggml
Nvidia GPU
#41
opened Apr 1, 2026 by
terrysimons
•
Draft
ProTip!
Type g i on any issue or pull request to go back to the issue listing page.