GenAI monitor
551 subscribers
4.35K links
AI frontier model updates & open source LLM releases
Download Telegram
πŸ“° NVIDIA - Deploy a Production-Ready NVIDIA AI-Q Blueprint on Oracle Cloud Infrastructure
AI agents have changed a lot in the last two years. The first could only answer one question at a time. Then came multi-turn chat, where the model could keep…

https://developer.nvidia.com/blog/deploy-a-production-ready-nvidia-ai-q-blueprint-on-oracle-cloud-infrastructure/


πŸ“° NVIDIA - Creating the NVIDIA Nemotron 3 Ultra NVFP4 Checkpoint with NVIDIA Model Optimizer
As context windows grow longer, moving large model weights efficiently becomes critical to performance. A common way to address this is quantization…

https://developer.nvidia.com/blog/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer/
πŸ†• [HF Models] deepseek-ai - DeepSeek-V4-Pro-DSpark

https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark


πŸ†• [HF Models] deepseek-ai - DeepSeek-V4-Flash-DSpark

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-DSpark
πŸ—“οΈ Weekly GitHub Activity


πŸ¦™ llama.cpp
β”” Release: b9743 β†’ b9828
β”” 85 commits

- Added support for Step 3.5/3.7 flash MTP3 speculative decoding #24340,
Granite Speech Plus #24818,
LFM2.5-ColBERT-350M/Embedding-350M #24913,
Eagle3 Qwen3 draft models #24977,
Unlimited-OCR #24969
- Added SSE Replay Buffer to server and UI, allowing text generation to survive HTTP disconnects and resume seamlessly #23226
- Introduced real-time model loading progress tracking via SSE in both server and UI #24828 #24878
- Configured server to create checkpoints before every user message to improve session recovery #24176
- Redesigned the WebUI with a new logo, navigation cleanup, and significant mobile layout improvements #24897
- Enabled dual-GPU tensor parallelism on the SYCL backend via split-mode tensor #24152
- Overhauled Hexagon matrix multiplication kernels with tiled layouts, HVX/HMX microkernels, and graph caching #24954
- Upgraded OpenCL Flash Attention kernels for F16, F32, Q4_0, and Q8_0 #25069
- Moved server model downloading to a dedicated child process #24834
- Added CUDA fast path for strided 2D copies using cudaMemcpy2DAsync #25057
- Reduced synchronization overhead between CPU and CUDA async copies during split compute #20793
- Added 3D convolution support to Vulkan #24612
- Fixed CUDA integer overflows and transposed copy failures #24706 #25000
- Fixed incorrect vector dot computations on SVE-enabled ARM CPUs #24699

πŸ”— All changes | Latest release


🎨 stable-diffusion.cpp
β”” Release: master-709-92a3b73 β†’ master-721-8caa3f9
β”” 12 commits

- Added support for Boogu image generation #1688
- Added support for Krea2 models #1705
- Introduced guidance_schedule support for generation control #1684
- Added logit-normal scheduler #1669
- Added --eager-load flag to pre-load parameters during model initialization #1687
- Added --prompt-file and --negative-prompt-file flags for file-based inputs #1693
- Fixed memory mapping by avoiding writable mmap for read-only weights #1698

πŸ”— All changes | Latest release


πŸ€— Fresh models trending on HuggingFace:

empero-ai/Qwythos-9B-Claude-Mythos-5-1M β™‘488
krea/Krea-2-Turbo β™‘310
krea/Krea-2-Raw β™‘214
deepreinforce-ai/Ornith-1.0-9B β™‘167
deepreinforce-ai/Ornith-1.0-35B β™‘161
deepreinforce-ai/Ornith-1.0-397B β™‘121
Chunjiang-Intelligence/DeepSeek-v4-Fable β™‘112
hustvl/Moebius β™‘51
AutoArk-AI/ARK-ASR-3B β™‘37
paom/texture2albedo-v2 β™‘32
SupraLabs/Supra-A2A-Nano-Exp β™‘30
Gryphe/Gemma-4-26B-A4B-StyleTune-V2 β™‘24
0xSero/GLM-5.2-504B β™‘19
g-astruc/UniverSat β™‘18
allenai/tmax-27b β™‘18
ValiantLabs/Qwen3.6-27B-Esper4 β™‘14
wikeeyang/Flux2-Klein-9B-True-V3 β™‘14
vrgamedevgirl84/Krea2_Enhancer β™‘12
❀2
πŸ†• [HF Models] deepseek-ai - eagle3_gemma4_12b_ttt7

https://huggingface.co/deepseek-ai/eagle3_gemma4_12b_ttt7


πŸ†• [HF Models] deepseek-ai - eagle3_qwen3_14b_ttt7

https://huggingface.co/deepseek-ai/eagle3_qwen3_14b_ttt7


πŸ†• [HF Models] deepseek-ai - eagle3_qwen3_8b_ttt7

https://huggingface.co/deepseek-ai/eagle3_qwen3_8b_ttt7


πŸ†• [HF Models] deepseek-ai - eagle3_qwen3_4b_ttt7

https://huggingface.co/deepseek-ai/eagle3_qwen3_4b_ttt7


πŸ†• [HF Models] deepseek-ai - dflash_gemma4_12b_block7

https://huggingface.co/deepseek-ai/dflash_gemma4_12b_block7


πŸ†• [HF Models] deepseek-ai - dflash_qwen3_14b_block7

https://huggingface.co/deepseek-ai/dflash_qwen3_14b_block7


πŸ†• [HF Models] deepseek-ai - dflash_qwen3_8b_block7

https://huggingface.co/deepseek-ai/dflash_qwen3_8b_block7


πŸ†• [HF Models] deepseek-ai - dflash_qwen3_4b_block7

https://huggingface.co/deepseek-ai/dflash_qwen3_4b_block7


πŸ†• [HF Models] deepseek-ai - dspark_gemma4_12b_block7

https://huggingface.co/deepseek-ai/dspark_gemma4_12b_block7


πŸ†• [HF Models] deepseek-ai - dspark_qwen3_14b_block7

https://huggingface.co/deepseek-ai/dspark_qwen3_14b_block7


πŸ†• [HF Models] deepseek-ai - dspark_qwen3_8b_block7

https://huggingface.co/deepseek-ai/dspark_qwen3_8b_block7


πŸ†• [HF Models] deepseek-ai - dspark_qwen3_4b_block7

https://huggingface.co/deepseek-ai/dspark_qwen3_4b_block7
πŸ“° PyTorch - Introducing Cross-Repository CI Relay: Scalable CI for PyTorch’s Out-of-Tree Backends
TL;DR PyTorch now has a Cross-Repository CI Relay (CRCR) that automatically triggers and tracks CI in downstream repositories whenever a PR is opened or a commit is pushed against pytorch/pytorch....

https://pytorch.org/blog/introducing-cross-repository-ci-relay-scalable-ci-for-pytorchs-out-of-tree-backends/
❀1
πŸ“° NVIDIA - How to Govern Autonomous Agents in Enterprise AI Factories 
AI agents are quickly moving beyond chat. They inspect code, run tests, read documents, search knowledge bases, query internal systems, and operate for hours on…

https://developer.nvidia.com/blog/how-to-govern-autonomous-agents-in-enterprise-ai-factories/