GenAI monitor
552 subscribers
4.39K links
AI frontier model updates & open source LLM releases
Download Telegram
πŸ—“οΈ Weekly GitHub Activity


πŸ¦™ llama.cpp
β”” Release: b9544 β†’ b9627
β”” 83 commits

- Support for EAGLE3 speculative decoding #18039
- Added Gemma 4 Multi-Token Prediction (MTP) and assistant draft-model support #23398 #24282
- New architecture support for Cohere2-MoE #24260
- Multi-Token Multi-Domain (MTMD) models now support video input and a batching API #24269 #24384
- WebUI implemented as a Progressive Web App (PWA) with offline caching #23871
- WebUI added an opt-in sandboxed JavaScript execution tool #24244
- GGML core bumped to version 0.15.1 e08c226
- WebGPU performance improvements for prefill and k-quants #24225
- Vulkan added fast paths for contiguous transfers and dot2 product extension support #23973 #24123
- Fixed CUDA ssm_scan_f32 data-races and CPU rms_norm_back in-place aliasing #24360 #24305
- Server added prompt logging to local directories #22031

πŸ”— All changes | Latest release


🎨 stable-diffusion.cpp
β”” Release: master-679-f3fd359 β†’ master-694-276025e
β”” 15 commits

- Added circular RoPE support for ideogram4 #1627
- Introduced free_sd_images function to manage memory for C API #1633
- Optimized performance by capping planner budget when models exceed streaming limits #1612
- Normalized APG diff_norm calculations by tensor size #1620
- Fixed SD3 conditioning crash when clip_l text encoder is missing #1638
- Corrected mask shape for masked flash attention #1625
- Resolved LoKR application issue by correctly marking w2_a tensors #1650

πŸ”— All changes | Latest release


πŸ€— Fresh models trending on HuggingFace:

bosonai/higgs-audio-v3-tts-4b β™‘414
nex-agi/Nex-N2-mini β™‘193
prefeitura-rio/Rio-3.5-Open-397B β™‘108
RazzzHF/Realism_Engine_Ideogram_4 β™‘90
silx-ai/Quasar-Preview β™‘63
mindlab-research/Macaron-V1-Preview-749B β™‘57
Zyphra/ZONOS2 β™‘56
BennyDaBall/Z-Image-Engineer-V6 β™‘46
PaddlePaddle/pp-ocrv6 β™‘43
MooreThreads/MusaCoder-27B β™‘35
apodex/Apodex-1.0-mini β™‘31
Muhammadreza/alduin-4b-it-base β™‘28
zjunlp/LabVLA β™‘26
fancyfeast/bigasp-3 β™‘20
Photoroom/prxpixel-t2i β™‘20
LatentForce-ai/Cassini-1.0 β™‘20
Gryphe/Pantheon-Reasoning-26B-A4B-1.1 β™‘19
libertywing/FlashMemory-Deepseek-V4 β™‘19
dx8152/Flux2-Klein-9B-Migration β™‘19
apodex/Apodex-1.0-4B-SFT β™‘18
Gryphe/Gemma-4-31B-StyleTune β™‘17
VAGOsolutions/SauerkrautLM-LFM2.5-GLiNER β™‘16
tsolful/zjourney-Ideogram-4-Fantasy-Realism-Refiner β™‘14
πŸ“° OpenAI - New OpenAI Academy courses for the next era of work

https://openai.com/index/academy-courses-applying-ai-at-work


πŸ”“ OpenAI - How an astrophysicist uses Codex to help simulate black holes

https://openai.com/index/using-codex-to-simulate-black-holes


πŸ”“ OpenAI - BBVA puts AI at the core of banking with OpenAI

https://openai.com/index/bbva


πŸ”“ OpenAI - Creating new simulations of black holes with Codex

https://openai.com/index/creating-new-simulations-black-holes


πŸ”“ OpenAI - Ad Tools Terms

https://openai.com/policies/ad-tools-terms


πŸ”“ OpenAI - OpenAI Academy

https://openai.com/academy


πŸ”“ OpenAI - How Preply combines AI and human tutors to personalize learning

https://openai.com/index/preply
πŸ†• [HF Models] microsoft - FastContext-1.0-4B-SFT


https://huggingface.co/microsoft/FastContext-1.0-4B-SFT
πŸ“° NVIDIA - Boosting MoE Training Throughput with Advanced Fusion Kernels
Mixture-of-experts (MoE) models have quickly become a foundational component of modern, large-scale AI systems. They are widely adopted because they enable…

https://developer.nvidia.com/blog/boosting-moe-training-throughput-with-advanced-fusion-kernels/


πŸ“° NVIDIA - Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models
Quick glossary for readers new to VLA/WAM terminology VLA Vision-Language-Action model: a robot policy that starts from a pretrained VLM backbone and adapts it…

https://developer.nvidia.com/blog/pretrained-to-imagine-fine-tuned-to-act-the-rise-of-world-action-models/
πŸ“° Google AI Blog - Unlocking the Power of the TPU Stack: Introducing our new Developer Hub
Google has officially launched the TPU Developer Hub, a centralized educational resource designed to help model builders and developers maximize the performance of Google Cloud TPUs. The hub offers code-first resources, open-source recipes, and deep-dive documentation covering hardware architecture, software optimization, debugging, parallelism, and networking. These materials are tailored for both human developers and AI-assisted tools to streamline everything from large-scale training to low-latency inference workloads.

https://developers.googleblog.com/en/unlocking-the-power-of-the-tpu-stack-introducing-our-new-developer-hub/
πŸ”“ Qwen Research - Qwen-RobotNav: A Scalable Navigation Model Designed for an Agentic Navigation System

https://qwen.ai/blog?id=qwen-robotnav


πŸ“° Qwen Research - Qwen-RobotWorld: Boundless Worlds for Embodied Agents

https://qwen.ai/blog?id=qwen-robotworld


πŸ“° Qwen Research - Qwen-RobotManip: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

https://qwen.ai/blog?id=qwen-robotmanip


πŸ“° Qwen Research - Qwen-Robot Suite: A Foundation Model Suite for Physical World Intelligence

https://qwen.ai/blog?id=qwen-robotsuite