-
Notifications
You must be signed in to change notification settings - Fork 20.9k
Pull requests: ggml-org/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
skill: create Improvements or additions to documentation
add-new-model and code-review
documentation
#26042
opened Jul 23, 2026 by
ngxson
Collaborator
Loading…
ggml: fix backend split scheduler race condition
ggml
changes relating to the ggml tensor library for machine learning
#26040
opened Jul 23, 2026 by
0cc4m
Contributor
Loading…
server: add optional repetition detection
documentation
Improvements or additions to documentation
server
testing
Everything test related
#26039
opened Jul 23, 2026 by
seryogakovalyov
Contributor
•
Draft
Feature: Add p-less sampling
testing
Everything test related
#26035
opened Jul 23, 2026 by
BII-wushuang
Loading…
metal: implement soft max backward operation in Metal backend
Apple Metal
https://en.wikipedia.org/wiki/Metal_(API)
ggml
changes relating to the ggml tensor library for machine learning
#26033
opened Jul 23, 2026 by
kunwar-vikrant
Loading…
hexagon: fix Windows crash when op_poll is enabled
ggml
changes relating to the ggml tensor library for machine learning
Hexagon
#26029
opened Jul 23, 2026 by
adgup
Loading…
CUDA: support non-contiguous rows in L2_NORM
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
#26026
opened Jul 23, 2026 by
MagicalFlames
Loading…
llama : stage mmap uploads on integrated GPUs
#26023
opened Jul 23, 2026 by
liminfei-amd
Contributor
Loading…
Cohere2 MoE template parser: support JSON schema
testing
Everything test related
#26018
opened Jul 22, 2026 by
boondocklabs
Contributor
Loading…
sycl: fuse RMS_NORM + MUL
ggml
changes relating to the ggml tensor library for machine learning
SYCL
https://en.wikipedia.org/wiki/SYCL - GPU programming language
#26015
opened Jul 22, 2026 by
Titaniumtown
Contributor
•
Draft
OAI Responses API json schema support, Cohere2 MoE template parser json schema support, improvements to responses streaming compatibility
server
testing
Everything test related
#26013
opened Jul 22, 2026 by
boondocklabs
Contributor
•
Draft
ui: fix MCP server display name conflicts in tools lists
server/ui
#26011
opened Jul 22, 2026 by
ServeurpersoCom
Contributor
Loading…
hexagon: partial im2col support
ggml
changes relating to the ggml tensor library for machine learning
Hexagon
#26007
opened Jul 22, 2026 by
tboinovski1
Contributor
Loading…
ui: fix system message edit box not expanding to fit content
server/ui
#26006
opened Jul 22, 2026 by
pieroevcc
Loading…
[SYCL] support the missed types in cpy
ggml
changes relating to the ggml tensor library for machine learning
SYCL
https://en.wikipedia.org/wiki/SYCL - GPU programming language
server : preserve context checkpoints across slot save/restore
server
#26004
opened Jul 22, 2026 by
Tough-Respawn
Loading…
llama : add --lazy-experts for MoE models larger than RAM
#26003
opened Jul 22, 2026 by
pwilkin
Member
Loading…
UI: Fix settings precedence, Factory < Admin (--ui-config-file) < Users (Settings panel)
server/ui
#26002
opened Jul 22, 2026 by
ServeurpersoCom
Contributor
Loading…
[Model] Add support for Nanbeige4.2
conversion
model
Model specific
#25994
opened Jul 22, 2026 by
zqlcode
Loading…
add minicpmv46 downsample
mtmd
Related to multimodal functionality (video/image/audio)
server
#25993
opened Jul 22, 2026 by
tc-mb
Contributor
Loading…
ggml-cpu : add LoongArch64 LASX Q4_0 repack path
ggml
changes relating to the ggml tensor library for machine learning
#25991
opened Jul 22, 2026 by
ztsubaki
Loading…
cmake(ppc64le): Add clang toolchain
build
Compilation issues
documentation
Improvements or additions to documentation
#25988
opened Jul 22, 2026 by
JeremyRand
Contributor
Loading…
Previous Next
ProTip!
Type g p on any issue or pull request to go back to the pull request listing page.