Skip to content

Pull requests: ggml-org/llama.cpp

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

skill: create add-new-model and code-review documentation Improvements or additions to documentation
#26042 opened Jul 23, 2026 by ngxson Collaborator Loading…
ggml: fix backend split scheduler race condition ggml changes relating to the ggml tensor library for machine learning
#26040 opened Jul 23, 2026 by 0cc4m Contributor Loading…
server: add optional repetition detection documentation Improvements or additions to documentation server testing Everything test related
#26039 opened Jul 23, 2026 by seryogakovalyov Contributor Draft
Feature: Add p-less sampling testing Everything test related
#26035 opened Jul 23, 2026 by BII-wushuang Loading…
metal: implement soft max backward operation in Metal backend Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning
#26033 opened Jul 23, 2026 by kunwar-vikrant Loading…
hexagon: fix Windows crash when op_poll is enabled ggml changes relating to the ggml tensor library for machine learning Hexagon
#26029 opened Jul 23, 2026 by adgup Loading…
CUDA: support non-contiguous rows in L2_NORM CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#26026 opened Jul 23, 2026 by MagicalFlames Loading…
llama : stage mmap uploads on integrated GPUs
#26023 opened Jul 23, 2026 by liminfei-amd Contributor Loading…
Cohere2 MoE template parser: support JSON schema testing Everything test related
#26018 opened Jul 22, 2026 by boondocklabs Contributor Loading…
sycl: fuse RMS_NORM + MUL ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language
#26015 opened Jul 22, 2026 by Titaniumtown Contributor Draft
Windows unbuffered model load
#26014 opened Jul 22, 2026 by JTischbein Contributor Draft
hexagon: partial im2col support ggml changes relating to the ggml tensor library for machine learning Hexagon
#26007 opened Jul 22, 2026 by tboinovski1 Contributor Loading…
[SYCL] support the missed types in cpy ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language
#26005 opened Jul 22, 2026 by arthw Contributor Draft
llama : add --lazy-experts for MoE models larger than RAM
#26003 opened Jul 22, 2026 by pwilkin Member Loading…
CUDA: Support of GDN chunked kernel for prefill CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning testing Everything test related
#26001 opened Jul 22, 2026 by BLSharda Draft
[Model] Add support for Nanbeige4.2 conversion model Model specific
#25994 opened Jul 22, 2026 by zqlcode Loading…
add minicpmv46 downsample mtmd Related to multimodal functionality (video/image/audio) server
#25993 opened Jul 22, 2026 by tc-mb Contributor Loading…
ggml-cpu : add LoongArch64 LASX Q4_0 repack path ggml changes relating to the ggml tensor library for machine learning
#25991 opened Jul 22, 2026 by ztsubaki Loading…
16393 webui models management server/ui
#25990 opened Jul 22, 2026 by allozaur Contributor Draft
cmake(ppc64le): Add clang toolchain build Compilation issues documentation Improvements or additions to documentation
#25988 opened Jul 22, 2026 by JeremyRand Contributor Loading…
ProTip! Type g p on any issue or pull request to go back to the pull request listing page.