π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-07-31 20:10 UTC
π Assets: 25 files
π What's New:
<details open>
Support rotated kv cache quant (#26180)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10213π Released: 2026-07-31 20:10 UTC
π Assets: 25 files
π What's New:
<details open>
Support rotated kv cache quant (#26180)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-07-31 20:46 UTC
π Assets: 25 files
π What's New:
<details open>
mtmd: add nembdhead (#26342)
Co-authored-by: Daniel Han <unslothai@gmail.com>
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10214π Released: 2026-07-31 20:46 UTC
π Assets: 25 files
π What's New:
<details open>
mtmd: add nembdhead (#26342)
Co-authored-by: Daniel Han <unslothai@gmail.com>
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-07-31 22:04 UTC
π Assets: 25 files
π What's New:
<details open>
vulkan: add POOL1D op (#25431)
* vulkan : add pool1d push constants and pipeline field
Declared data structures needed for POOL1D OP, which are the vkoppool1dpushconstants struct and pipelinepool1df32 field.
* vulkan : add pool1d compute shader
Added pool1d.comp for Vulkan backend mirroring the existing pool2d shader.
* vulkan : add full GGMLOPPOOL1D support
Adde...
π View on GitHub
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10216π Released: 2026-07-31 22:04 UTC
π Assets: 25 files
π What's New:
<details open>
vulkan: add POOL1D op (#25431)
* vulkan : add pool1d push constants and pipeline field
Declared data structures needed for POOL1D OP, which are the vkoppool1dpushconstants struct and pipelinepool1df32 field.
* vulkan : add pool1d compute shader
Added pool1d.comp for Vulkan backend mirroring the existing pool2d shader.
* vulkan : add full GGMLOPPOOL1D support
Adde...
π View on GitHub
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-01 06:50 UTC
π Assets: 25 files
π What's New:
<details open>
chat : enable tool call in thinking for DS4 (#26269)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10217π Released: 2026-08-01 06:50 UTC
π Assets: 25 files
π What's New:
<details open>
chat : enable tool call in thinking for DS4 (#26269)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-01 12:46 UTC
π Assets: 24 files
π What's New:
<details open>
mtmd: add minicpmv46 downsample (#25993)
add minicpmv46 downsample
Signed-off-by: tc-mb <tianchi_cai@icloud.com>
put downsample mode inside gguf.
Signed-off-by: tc-mb <tianchicai@icloud.com>
* build mtmdimagepreprocessorllavauhd
Signed-off-by: tc-mb <tianchicai@icloud.com>
fix code
Signed-off-by: tc-mb <tianchi_cai@icloud.com>
add convert
Signed-off-by: t...
π View on GitHub
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10218π Released: 2026-08-01 12:46 UTC
π Assets: 24 files
π What's New:
<details open>
mtmd: add minicpmv46 downsample (#25993)
add minicpmv46 downsample
Signed-off-by: tc-mb <tianchi_cai@icloud.com>
put downsample mode inside gguf.
Signed-off-by: tc-mb <tianchicai@icloud.com>
* build mtmdimagepreprocessorllavauhd
Signed-off-by: tc-mb <tianchicai@icloud.com>
fix code
Signed-off-by: tc-mb <tianchi_cai@icloud.com>
add convert
Signed-off-by: t...
π View on GitHub
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-01 16:46 UTC
π Assets: 25 files
π What's New:
<details open>
cli : persist reasoningcontent in chat history (#26362)
* cli : persist reasoningcontent in chat history
llama-cli collected reasoning from the stream for display but only
stored assistant content in messages, so --reasoning-preserve could
not re-inject prior thoughts on later turns.
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm6...
π [View on GitHub
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10219π Released: 2026-08-01 16:46 UTC
π Assets: 25 files
π What's New:
<details open>
cli : persist reasoningcontent in chat history (#26362)
* cli : persist reasoningcontent in chat history
llama-cli collected reasoning from the stream for display but only
stored assistant content in messages, so --reasoning-preserve could
not re-inject prior thoughts on later turns.
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm6...
π [View on GitHub
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-01 19:30 UTC
π Assets: 25 files
π What's New:
<details open>
vendor : update BoringSSL to 0.20260730.0 (#26353)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10221π Released: 2026-08-01 19:30 UTC
π Assets: 25 files
π What's New:
<details open>
vendor : update BoringSSL to 0.20260730.0 (#26353)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-01 22:48 UTC
π Assets: 25 files
π What's New:
<details open>
test: fix some CI errors (#26415)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10223π Released: 2026-08-01 22:48 UTC
π Assets: 25 files
π What's New:
<details open>
test: fix some CI errors (#26415)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-02 07:03 UTC
π Assets: 25 files
π What's New:
<details open>
ggml-webgpu: add support for f16 repeat (#26307)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10224π Released: 2026-08-02 07:03 UTC
π Assets: 25 files
π What's New:
<details open>
ggml-webgpu: add support for f16 repeat (#26307)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-02 08:15 UTC
π Assets: 25 files
π What's New:
<details open>
sycl: fix classification of iGPUs (#26105)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10226π Released: 2026-08-02 08:15 UTC
π Assets: 25 files
π What's New:
<details open>
sycl: fix classification of iGPUs (#26105)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-02 08:30 UTC
π Assets: 25 files
π What's New:
<details open>
model : load MiMo V2 MTP tensors only if used (#26412)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)...
π View on GitHub
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10225π Released: 2026-08-02 08:30 UTC
π Assets: 25 files
π What's New:
<details open>
model : load MiMo V2 MTP tensors only if used (#26412)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)...
π View on GitHub
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-02 09:43 UTC
π Assets: 25 files
π What's New:
<details open>
chat : add qwen3 specialized parser (#26252)
Add tagged thinking tool parser
chat : refactor and add permute helper
cont : add support for <tool_call> omission
cont : update tool delimiters
cont : add comment for qwen3-coder
cont : fix trigger pattern for <function
---------
Co-authored-by: Bart de Boer <bart.deboer@gmail.com>
</details>
Website:
- <htt...
π View on GitHub
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10227π Released: 2026-08-02 09:43 UTC
π Assets: 25 files
π What's New:
<details open>
chat : add qwen3 specialized parser (#26252)
Add tagged thinking tool parser
chat : refactor and add permute helper
cont : add support for <tool_call> omission
cont : update tool delimiters
cont : add comment for qwen3-coder
cont : fix trigger pattern for <function
---------
Co-authored-by: Bart de Boer <bart.deboer@gmail.com>
</details>
Website:
- <htt...
π View on GitHub
π₯ Download Release
π New Release Alert!
π¦ thefeed
π· Category: Community Suggested
π Version:
π Released: 2026-08-02 11:01 UTC
π Assets: 27 files
π What's New:
## Install


Android and iOS install...
π View on GitHub
π₯ Download Release
π¦ thefeed
π· Category: Community Suggested
π Version:
v0.38.0π Released: 2026-08-02 11:01 UTC
π Assets: 27 files
π What's New:
## Install


Android and iOS install...
π View on GitHub
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-02 13:28 UTC
π Assets: 25 files
π What's New:
<details open>
DeepseekV4 MTP + DSpark (#25784)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10228π Released: 2026-08-02 13:28 UTC
π Assets: 25 files
π What's New:
<details open>
DeepseekV4 MTP + DSpark (#25784)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-02 14:31 UTC
π Assets: 25 files
π What's New:
<details open>
opencl: bugfix increment refcount in ggmlbackendopenclinit() (#26162)
Incrementing
in the
If we do not increment the
and consequently, the profiling data would not be flushed and written.
( #ifdef GGMLOPENCLPROFILI...
π View on GitHub
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10229π Released: 2026-08-02 14:31 UTC
π Assets: 25 files
π What's New:
<details open>
opencl: bugfix increment refcount in ggmlbackendopenclinit() (#26162)
Incrementing
ref_count at the beginning is important laterin the
free() method of the ggml_backend_opencl_context at program end.If we do not increment the
ref_count, the result would be -1 here,and consequently, the profiling data would not be flushed and written.
( #ifdef GGMLOPENCLPROFILI...
π View on GitHub
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-02 18:14 UTC
π Assets: 25 files
π What's New:
<details open>
common: support the DSpark sidecar resolution (#26458)
The dspark- files resolve like the other speculative sidecars: the
-hfd tag applies to them, a requested sidecar resolves without a full
model at the tag, and an explicit -md selection disables the discovery.
When no type is requested, dspark outranks dflash in the auto-selection
since its sidecar carries the extra Markov h...
π View on GitHub
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10231π Released: 2026-08-02 18:14 UTC
π Assets: 25 files
π What's New:
<details open>
common: support the DSpark sidecar resolution (#26458)
The dspark- files resolve like the other speculative sidecars: the
-hfd tag applies to them, a requested sidecar resolves without a full
model at the tag, and an explicit -md selection disables the discovery.
When no type is requested, dspark outranks dflash in the auto-selection
since its sidecar carries the extra Markov h...
π View on GitHub
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-02 18:57 UTC
π Assets: 25 files
π What's New:
<details open>
metal: implement DeepSeek V4 hyper-connections (#26459)
- Implement GGMLOPDSV4HCCOMB, GGMLOPDSV4HCPRE, and
GGMLOPDSV4HCPOST with SIMDgroup register and shuffle optimized kernels.
- Add Metal dispatch and support plumbing and test the production Sinkhorn
iteration count and embedding width.
Assisted-by: Codex
Co-authored-by: Thiago Padilha <thiago@padilha.cc>
...
π View on GitHub
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10232π Released: 2026-08-02 18:57 UTC
π Assets: 25 files
π What's New:
<details open>
metal: implement DeepSeek V4 hyper-connections (#26459)
- Implement GGMLOPDSV4HCCOMB, GGMLOPDSV4HCPRE, and
GGMLOPDSV4HCPOST with SIMDgroup register and shuffle optimized kernels.
- Add Metal dispatch and support plumbing and test the production Sinkhorn
iteration count and embedding width.
Assisted-by: Codex
Co-authored-by: Thiago Padilha <thiago@padilha.cc>
...
π View on GitHub
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-02 20:17 UTC
π Assets: 25 files
π What's New:
<details open>
metal : add F16 support for bin ops (#26465)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10234π Released: 2026-08-02 20:17 UTC
π Assets: 25 files
π What's New:
<details open>
metal : add F16 support for bin ops (#26465)
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-02 21:02 UTC
π Assets: 25 files
π What's New:
<details open>
metal : add SILUBACK (#25982)
* feat(siluback): implemented siluback op for f32
* fix(siluback): removed redundant asserts in ggml-metal-ops.cpp function ggmlmetalopsiluback.
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
...
π View on GitHub
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10235π Released: 2026-08-02 21:02 UTC
π Assets: 25 files
π What's New:
<details open>
metal : add SILUBACK (#25982)
* feat(siluback): implemented siluback op for f32
* fix(siluback): removed redundant asserts in ggml-metal-ops.cpp function ggmlmetalopsiluback.
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
...
π View on GitHub
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-03 07:50 UTC
π Assets: 25 files
π What's New:
<details open>
llama : MTP support for DeepSeek V3.2 (#26457)
llama : MTP support for DeepSeek V3.2
model : no need to include MTP layers during DeepSeek V3.2 model type discovery
---------
Co-authored-by: StanisΕaw Szymczyk <sszymczy@gmail.com>
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10237π Released: 2026-08-03 07:50 UTC
π Assets: 25 files
π What's New:
<details open>
llama : MTP support for DeepSeek V3.2 (#26457)
llama : MTP support for DeepSeek V3.2
model : no need to include MTP layers during DeepSeek V3.2 model type discovery
---------
Co-authored-by: StanisΕaw Szymczyk <sszymczy@gmail.com>
</details>
Website:
- <https://llama.app>
macOS/iOS:
- macOS Apple Silicon (arm64)
π₯ Download Release
π New Release Alert!
π¦ llama.cpp
π· Category: AI
π Version:
π Released: 2026-08-03 08:56 UTC
π Assets: 25 files
π What's New:
<details open>
model: MTP support for Qwen3-Next (#25589)
mtp for qwen3nex
fix for python type-check
Fix to compute num_mtp from directly mtp layer
define optnummtplayers in QwenMtpMixin and fix some comments
Fix for python type check
Update gguf-py/gguf/constants.py
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
rebase and add load_mtp flags
U...
π View on GitHub
π₯ Download Release
π¦ llama.cpp
π· Category: AI
π Version:
b10238π Released: 2026-08-03 08:56 UTC
π Assets: 25 files
π What's New:
<details open>
model: MTP support for Qwen3-Next (#25589)
mtp for qwen3nex
fix for python type-check
Fix to compute num_mtp from directly mtp layer
define optnummtplayers in QwenMtpMixin and fix some comments
Fix for python type check
Update gguf-py/gguf/constants.py
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
rebase and add load_mtp flags
U...
π View on GitHub
π₯ Download Release