GenAI monitor
551 subscribers
4.38K links
AI frontier model updates & open source LLM releases
Download Telegram
🆕 [HF Models] FunAudioLLM - Fun-ASR-Nano-GGUF

https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-GGUF


🆕 [HF Models] FunAudioLLM - Paraformer-GGUF

https://huggingface.co/FunAudioLLM/Paraformer-GGUF


🆕 [HF Models] FunAudioLLM - SenseVoiceSmall-GGUF

https://huggingface.co/FunAudioLLM/SenseVoiceSmall-GGUF


🆕 [HF Models] FunAudioLLM - fsmn-vad-GGUF

https://huggingface.co/FunAudioLLM/fsmn-vad-GGUF
🗓️ Weekly GitHub Activity


🦙 llama.cpp
└ Release: b9627 → b9743
└ 116 commits

- Added support for Cohere2MoE (North Code / Tiny Aya) and GLM-5.2 models. #24615 #24770
- Integrated Eagle3 speculative decoding support for Qwen 3.5 and 3.6. #24593
- Updated OpenVINO backend to 2026.2 with context-shift, Q5_1 weights, and Gemma 4 support. #24503
- Introduced a model management API to the server router for remote model downloads and deletion. #23976
- UI improvements: added HEIC/HEIF image support, SVG/Mermaid rendering with source toggles, and markdown rendering for thinking blocks. #24137 #24080 #24611
- Optimized AMX performance on CPU and improved i-quants prefill speeds for WebGPU. #24806 #24530
- Enhanced Metal backend with concat support for F16/BF16 and rope_back operator. #24724 #24725
- Fixed significant whitespace issues in chat grammar generation and double-escaping in tool-call parsing. #24624 #24667
- Server now includes real-time generation speed metrics and JSONL conversation exports. #24291 #24688
- SYCL backend updates: added Conv2D/Conv3D support, dev-to-dev memcpy, and set F16 as default. #24600 #24476 #23996

🔗 All changes | Latest release


🎨 stable-diffusion.cpp
└ Release: master-694-276025e → master-709-92a3b73
└ 15 commits

- Added RPC support for remote compute execution #1629
- Implemented PuLID-Flux identity-injection support for Flux models #1595
- Added support for cancelling ongoing generations with partial image batch returns #1124
- Introduced disk parameters backend support #1651
- Added backend-specific max-VRAM budgets bb90bfa
- Fixed handling of oversized Vulkan parameter tensors #1662
- Synchronized core library with latest GGML #1656

🔗 All changes | Latest release


🤗 Fresh models trending on HuggingFace:

WeiboAI/VibeThinker-3B ♡511
prefeitura-rio/Rio-3.5-Open-397B ♡327
owensong/Inflect-Nano-v1 | gguf ♡140
Zyphra/ZONOS2 ♡118
datalab-to/lift ♡86
poolside/Laguna-M.1 ♡74
Boogu/Boogu-Image-0.1-Edit ♡67
AlexWortega/SIQ-1-35B ♡59
Boogu/Boogu-Image-0.1-Turbo ♡37
SupraLabs/Supra-1.5-50M-Instruct-exp ♡37
Boogu/Boogu-Image-0.1-Base ♡32
HKUSTAudio/AudioX-Turbo ♡29
FINAL-Bench/Darwin-398B-JGOS ♡28
Multilingual-Multimodal-NLP/LoopCoder-V2 ♡26
YTan2000/Qwen3.6-27B-MTP-TQ3_4S ♡16
Danrisi/UltraReal_FineTune_Anima_base1_v3 ♡14
catnip-ai-tech/MaineCoon ♡14
📰 Google AI Blog - Build Cross-Language Multi-Agent Team with Google’s Agent Development Kit and A2A
How a Python agent and a Go agent collaborate on contract compliance using the Agent2Agent protocolY...

https://developers.googleblog.com/en/build-cross-language-multi-agent-team-with-google-agent-development-kit-and-a2a/
📰 NVIDIA - Enable Real-Time AI for High-Speed Data Acquisition with DAQIRI
When AlphaFold2 revolutionized drug discovery in 2020, its success relied entirely on the roughly 170,000 protein structures collected by scientists since 1971…

https://developer.nvidia.com/blog/enable-real-time-ai-for-high-speed-data-acquisition-with-daqiri/


📰 NVIDIA - Inside NVIDIA Halos for Robotics: A Full-Stack Functional Safety System for Physical AI
Physical AI—robots working autonomously alongside people in factories, warehouses, hospitals, and homes—is arriving faster than most expected.

https://developer.nvidia.com/blog/inside-nvidia-halos-for-robotics-a-full-stack-functional-safety-system-for-physical-ai/
📰 PyTorch - Serving DeepSeek-V4 on GB300 with SGLang: 5x Higher Throughput at the Same Interactivity Since Day-0
TL;DR: DeepSeek-V4 support was live in SGLang on Day-0, but the Day-0 stack was only the starting point. Since launch, we have coordinated a set of kernel, runtime, and hardening...

https://pytorch.org/blog/serving-deepseek-v4-on-gb300-with-sglang-5x-higher-throughput-at-the-same-interactivity-since-day-0/
📰 NVIDIA - Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding
As AI systems move from single-turn interactions to coordinated multiagent workflows, low-latency inference becomes increasingly important.

https://developer.nvidia.com/blog/boost-inference-performance-up-to-15x-on-nvidia-blackwell-using-dflash-speculative-decoding/


📰 NVIDIA - How Telcos Build Autonomous Networks with Agentic AI
Telecom operators are adopting AI across network operations, customer care, and back-office workflows, but most are still early in the journey to autonomy.

https://developer.nvidia.com/blog/how-telcos-build-autonomous-networks-with-agentic-ai/
🆕 [HF Models] Qwen - Qwen-AgentWorld-35B-A3B


https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B