GenAI monitor
552 subscribers
4.4K links
AI frontier model updates & open source LLM releases
Download Telegram
🗓️ Weekly GitHub Activity


🦙 llama.cpp
└ Release: b9437 → b9544
└ 107 commits

- Added support for EXAONE 4.5 #21733
- Added support for Granite4 Vision #23545
- Added support for Step3.7-Flash #23845
- Added support for Mellum architecture #23966
- Added support for Granite Multilingual Embeddings R2 #22716
- Added support for StepFun 3.5 MTP #23274
- Added tokenizer support for jina-embeddings-v2-base-zh #18756
- Initial support for Qwen3 SSM recurrent architectures #24031
- Server: Real-time reasoning interruption via new control endpoint #23971
- Server: Added placeholder bitmap for token counting and input_tokens API #23913
- Web UI: Thinking mode toggle, reasoning effort levels, and single-line preview #23434, #23601
- Web UI: Mermaid diagrams support and interactive preview #24032
- Tensor Parallel: Quantized KV cache support #23792
- Multimodal: Added frame merge support for Qwen-VL models #21858
- Vulkan: Optimized Q3_K/Q6_K performance on Intel Xe2/BMG via block loads 1962000
- CUDA: Improved MTP performance via mul_mat_vec_q_moe enrollment into PDL #24087
- Hexagon: Major optimizations for MUL_MAT, FLASH_ATTN, and GDN #23989
- Metal: Templated GLU kernels to support f16/f32 #23882
- Web UI: Custom CSS injection via configuration #23904
- KV-cache: SWA checkpoints store only non-masked cells #23981
- Deprecated llama_set_warmup #24009
- Fix model parameters not being propagated correctly to backend #23893
- Server: Avoid unnecessary checkpoint restore when new tokens are present #24110
- Fix session state corruption in common_prompt_batch_decode #23468

🔗 All changes | Latest release


🎨 stable-diffusion.cpp
└ Release: master-660-d2797b8 → master-679-f3fd359
└ 19 commits

- Added support for Ideogram 4 models #1609
- Added support for Wan2.2 5B FLF2V #1110
- Implemented PiD support #1585
- Added Adaptive Projected Guidance (APG) and unconditional Skip Layer Guidance (SLG) #593
- Added --stream-layers to stream weights from CPU during generation #1576
- Added img-cfg support for edit models #929
- Optimized performance via pinned host buffer allocation and streaming budget management #1601 #1611
- Fixed Flash Attention KV padding issues #1453

🔗 All changes | Latest release


🤗 Fresh models trending on HuggingFace:

ideogram-ai/ideogram-4-nf4 ♡212
bosonai/higgs-audio-v3-tts-4b ♡153
Hcompany/Holo-3.1-4B ♡57
LiconStudio/LTX-2.3-Multiple-Subject-Reference ♡50
VAST-AI/TripoSplat ♡49
nex-agi/Nex-N2-Pro ♡48
Hcompany/Holo-3.1-35B-A3B ♡36
SupraLabs/Supra-50M-Reasoning ♡30
Aratako/Irodori-TTS-600M-v3-VoiceDesign ♡29
mudler/parakeet-cpp-gguf ♡28
nex-agi/Nex-N2-mini ♡22
Trendyol/Trendyol-TTS ♡22
litert-community/gemma-4-12B-it-litert-lm ♡20
latam-gpt/Llama-3.1-70B-LatamGPT-SFT-1.0 ♡20
Soul-AILab/SoulX-Transcriber ♡17
Hcompany/Holo-3.1-9B ♡17
Hcompany/Holo-3.1-0.8B ♡13
ideogram-ai/ideogram-4-nf4-diffusers ♡13
📰 HuggingFace - Her · हेर — a detective for your Claude Code sessions


https://huggingface.co/blog/build-small-hackathon/her-blog
📰 HuggingFace - Building Pakistan Notice Helper: A Small AI Tool for a Very Local Safety Problem


https://huggingface.co/blog/build-small-hackathon/building-pakistan-notice-helper