GenAI monitor
551 subscribers
4.33K links
AI frontier model updates & open source LLM releases
Download Telegram
📰 NVIDIA - Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading
Large language model (LLM) training workloads increasingly run into GPU memory limits before compute is fully used. Model weights, gradients, optimizer states…

https://developer.nvidia.com/blog/reducing-high-bandwidth-memory-bottlenecks-in-jax-based-llm-training-with-host-offloading/


📰 NVIDIA - AI Model Co-Design: Hardware-Friendly LLM Design
AI performance comes down to three dimensions: Deployments must balance all three: High accuracy is wasted if responses are slow, and raw throughput means…

https://developer.nvidia.com/blog/ai-model-co-design-hardware-friendly-llm-design/


📰 NVIDIA - Accelerating End-to-End Co-Folding Performance with NVIDIA BioNeMo Agent Toolkit
Biomolecular structure prediction and co-folding with models like OpenFold3 are now mainstream, large-scale workloads powering drug discovery and protein design.

https://developer.nvidia.com/blog/accelerating-end-to-end-co-folding-performance-with-nvidia-bionemo-agent-toolkit/
🆕 [HF Models] openbmb - UltraX-0.6B-Preview


https://huggingface.co/openbmb/UltraX-0.6B-Preview
🗓️ Weekly GitHub Activity


🦙 llama.cpp
└ Release: b9873 → b9966
└ 93 commits

- Added initial support for the ET backend targeting ET-Soc-1 hardware #24179
- Introduced Q2_0 quantization format with CPU backend support for Ternary Bonsai models #24448
- Added multimodal support for DeepSeek-OCR v1 multi-tile dynamic resolution #24717
- Refactored llama-cli into an HTTP-based implementation interacting with the server #24948
- Fused MMVQ post-scale for NVFP4 on CUDA to accelerate FP8 and NVFP4 models #24481
- Optimized OpenCL Flash Attention decoding performance #25366
- Fixed a decode bottleneck in tensor-split mode by compiling regex patterns statically #24710
- Enabled unsafe math optimizations for AMD/HIP builds to match CUDA performance #24668
- Fixed a security vulnerability involving out-of-bounds reads in the UGM tokenizer #18750
- Fixed a crash occurring when using tensor parallelism with CPU-offloaded MoE experts #25028

🔗 All changes | Latest release


🎨 stable-diffusion.cpp
└ Release: master-749-b11c95a → master-775-b5d8120
└ 26 commits

- Added support for Krea2OstrisEdit #1775 and lingbot video #1770
- Support hot-reloading ControlNet to swap models without rebuilding the context #1768
- Added DPM++ 2M SDE and DPM++ 2M SDE Brownian tree samplers #1742 and #1743
- Support loading safetensors index files #1769
- Use denoise strength as the starting noise level #1738
- Moved circular padding from context to per-generation parameters #1748
- Improve generation speed by driving layer splitting from graph-cut segments #1762
- Fixed SDXL ControlNet integration issues regarding diffusers naming and graph size #1752
- Fixed UNet block paths splitting across layers #1741

🔗 All changes | Latest release


🤗 Fresh models trending on HuggingFace:

bottlecapai/ThinkingCap-Qwen3.6-27B ♡235
conradlocke/krea2-identity-edit ♡185
Alissonerdx/LTX-Best-Face-ID ♡99
SupraLabs/Supra-Router-51M ♡98
Patil/Krea-2-depth-controlnet ♡91
robbyant/lingbot-video-moe-30b-a3b ♡84
migtissera/Tess-4-27B ♡84
mgwr/M87 ♡68
ostris/krea2_turbo_style_reference ♡62
ai-sage/GigaChat3.5-432B-A28B ♡60
robbyant/lingbot-world-v2-14b-causal-fast ♡60
empero-ai/Qwythos-9B-v2 ♡43
wikeeyang/Krea2-Turbo-HD-V1 ♡28
MirilAI/Miril-Drone-2B-1 ♡27
robbyant/lingbot-vla-v2-6b ♡27
robbyant/lingbot-video-dense-1.3b ♡23
Ateron/Gemma-4-Novelist-Eclipse-31B ♡23
rzgar/Bernini-R-S2V ♡21
SOLRICKS/ltx-2.3-product-ad-style ♡21
ostris/Krea2OstrisEdit ♡20
ai-sage/GigaChat3.5-432B-A28B-base ♡18
robbyant/lingbot-vision-vit-large ♡18
epfl-neuroai/NEvo ♡18
OrionLLM/GRM-2.6-Plus-0628 ♡17
FrontiersMind/Lumma-0.6B-Base ♡13
sais-org/Polaris_Pro ♡13