GenAI monitor
551 subscribers
4.33K links
AI frontier model updates & open source LLM releases
Download Telegram
🆕 [HF Models] tencent - HiLS-Attention-7B


https://huggingface.co/tencent/HiLS-Attention-7B
📰 Google AI Blog - LiteRT.js, Google's high performance Web AI Inference
We're excited to introduce LiteRT.js, the newest member of the LiteRT family! LiteRT.js is our powerful solution for running machine learning models directly in the browser, extending Google's cross-platform edge AI runtime to the web. Built for JavaScript developers, LiteRT.js delivers state-of-the-art ML model inference performance on WebGPU and upcoming WebNN, with a fallback to WebAssembly for CPU. This post provides a quick tour of LiteRT.js and gives web developers everything they need to get started.

https://developers.googleblog.com/en/litertjs-googles-high-performance-web-ai-inference/
📰 NVIDIA - Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading
Large language model (LLM) training workloads increasingly run into GPU memory limits before compute is fully used. Model weights, gradients, optimizer states…

https://developer.nvidia.com/blog/reducing-high-bandwidth-memory-bottlenecks-in-jax-based-llm-training-with-host-offloading/


📰 NVIDIA - AI Model Co-Design: Hardware-Friendly LLM Design
AI performance comes down to three dimensions: Deployments must balance all three: High accuracy is wasted if responses are slow, and raw throughput means…

https://developer.nvidia.com/blog/ai-model-co-design-hardware-friendly-llm-design/


📰 NVIDIA - Accelerating End-to-End Co-Folding Performance with NVIDIA BioNeMo Agent Toolkit
Biomolecular structure prediction and co-folding with models like OpenFold3 are now mainstream, large-scale workloads powering drug discovery and protein design.

https://developer.nvidia.com/blog/accelerating-end-to-end-co-folding-performance-with-nvidia-bionemo-agent-toolkit/
🆕 [HF Models] openbmb - UltraX-0.6B-Preview


https://huggingface.co/openbmb/UltraX-0.6B-Preview
🗓️ Weekly GitHub Activity


🦙 llama.cpp
└ Release: b9873 → b9966
└ 93 commits

- Added initial support for the ET backend targeting ET-Soc-1 hardware #24179
- Introduced Q2_0 quantization format with CPU backend support for Ternary Bonsai models #24448
- Added multimodal support for DeepSeek-OCR v1 multi-tile dynamic resolution #24717
- Refactored llama-cli into an HTTP-based implementation interacting with the server #24948
- Fused MMVQ post-scale for NVFP4 on CUDA to accelerate FP8 and NVFP4 models #24481
- Optimized OpenCL Flash Attention decoding performance #25366
- Fixed a decode bottleneck in tensor-split mode by compiling regex patterns statically #24710
- Enabled unsafe math optimizations for AMD/HIP builds to match CUDA performance #24668
- Fixed a security vulnerability involving out-of-bounds reads in the UGM tokenizer #18750
- Fixed a crash occurring when using tensor parallelism with CPU-offloaded MoE experts #25028

🔗 All changes | Latest release


🎨 stable-diffusion.cpp
└ Release: master-749-b11c95a → master-775-b5d8120
└ 26 commits

- Added support for Krea2OstrisEdit #1775 and lingbot video #1770
- Support hot-reloading ControlNet to swap models without rebuilding the context #1768
- Added DPM++ 2M SDE and DPM++ 2M SDE Brownian tree samplers #1742 and #1743
- Support loading safetensors index files #1769
- Use denoise strength as the starting noise level #1738
- Moved circular padding from context to per-generation parameters #1748
- Improve generation speed by driving layer splitting from graph-cut segments #1762
- Fixed SDXL ControlNet integration issues regarding diffusers naming and graph size #1752
- Fixed UNet block paths splitting across layers #1741

🔗 All changes | Latest release


🤗 Fresh models trending on HuggingFace:

bottlecapai/ThinkingCap-Qwen3.6-27B ♡235
conradlocke/krea2-identity-edit ♡185
Alissonerdx/LTX-Best-Face-ID ♡99
SupraLabs/Supra-Router-51M ♡98
Patil/Krea-2-depth-controlnet ♡91
robbyant/lingbot-video-moe-30b-a3b ♡84
migtissera/Tess-4-27B ♡84
mgwr/M87 ♡68
ostris/krea2_turbo_style_reference ♡62
ai-sage/GigaChat3.5-432B-A28B ♡60
robbyant/lingbot-world-v2-14b-causal-fast ♡60
empero-ai/Qwythos-9B-v2 ♡43
wikeeyang/Krea2-Turbo-HD-V1 ♡28
MirilAI/Miril-Drone-2B-1 ♡27
robbyant/lingbot-vla-v2-6b ♡27
robbyant/lingbot-video-dense-1.3b ♡23
Ateron/Gemma-4-Novelist-Eclipse-31B ♡23
rzgar/Bernini-R-S2V ♡21
SOLRICKS/ltx-2.3-product-ad-style ♡21
ostris/Krea2OstrisEdit ♡20
ai-sage/GigaChat3.5-432B-A28B-base ♡18
robbyant/lingbot-vision-vit-large ♡18
epfl-neuroai/NEvo ♡18
OrionLLM/GRM-2.6-Plus-0628 ♡17
FrontiersMind/Lumma-0.6B-Base ♡13
sais-org/Polaris_Pro ♡13