GenAI monitor
550 subscribers
4.32K links
AI frontier model updates & open source LLM releases
Download Telegram
📰 Google AI Blog - Run Ray on TPU, Part 2: Ray AI libraries
This second installment explores how Ray’s higher-level libraries—Serve, Data, and Train—abstract the complexities of running AI workloads on Google's TPU slices. Ray Serve uses a simple topology configuration to correctly gang-schedule large multi-host models, while Ray Data eliminates data-loading bottlenecks by feeding accelerators directly with native JAX batches. Finally, JaxTrainer streamlines distributed training across TPUs by automatically handling cross-slice coordination, checkpointing, and fault tolerance.

https://developers.googleblog.com/en/run-ray-on-tpu-part-2-ray-ai-libraries/
📰 Claude Blog - How the product designer who built Claude Design uses it to explore ideas before building them

https://claude.com/blog/how-the-product-designer-who-built-claude-design-uses-it-to-explore-ideas-before-building-them


📰 Claude Blog - The new rules of context engineering for Claude 5 generation models

https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models


📰 Claude Blog - Claude models explained: choosing the best model for your use case

https://claude.com/blog/claude-models-explained-choosing-the-best-model-for-your-use-case


📰 Claude Blog - Four role-based certifications for the people who put Claude to work for customers

https://claude.com/blog/four-role-based-claude-certifications


📰 Claude Blog - Think through hard problems in voice mode

https://claude.com/blog/think-through-hard-problems-in-voice-mode


📰 Claude Blog - Building verification loops in Claude Code with skills

https://claude.com/blog/building-verification-loops-in-claude-code-with-skills


📰 Claude Blog - How Outtake built a cyber investigator on Claude

https://claude.com/blog/how-outtake-built-a-cyber-investigator-on-claude
📰 Anthropic System Cards - Claude Opus 5


https://anthropic.com/claude-opus-5-system-card
🗓️ Weekly GitHub Activity


🦙 llama.cpp
└ Release: b10068 → b10107
└ 39 commits

- Consolidated memory mapping and locking CLI options into a unified --load-mode argument #20834
- Added support for Laguna XS.2 and M.1 models #25165
- Improved DeepSeek-V4 support with softplus CUDA kernels, APE tensor fixes, and chat template updates #25896, #25945, #25414
- Fixed coordinate scaling in Qwen3-VL via align_corners interpolation #25781 and HunyuanVL XD-RoPE conversion #25514
- Enabled automatic speculative decoding type inference and sidecar model resolution from draft repositories #25955, #25989
- Upgraded CUDA backend with device-side GET_ROWS support for k-quants, i-quants, and mxfp4 #25962, alongside vectorized same-type copy optimization #25929
- Refactored Vulkan queue submission to bypass host-side locking via per-instance mutexes and unique handles #23570
- Added depthwise 2D convolution (CONV_2D_DW) kernel for WebGPU #25847
- Updated WebUI with bulk conversation management #25815, symbolic math support in the JS sandbox using Nerdamer #25948, and a default reasoning mode selector #25846

🔗 All changes | Latest release


🎨 stable-diffusion.cpp
└ Release: master-782-b290693 → master-795-87a0177
└ 13 commits

- Added support for Hunyuan Video 1.5 #1795
- Added IP-Adapter support for SD 1.5 and SDXL #1803 with CFG conditioning fixes #1815
- Added Mage-Flow support #1808

🔗 All changes | Latest release


🤗 Fresh models trending on HuggingFace:

Nanbeige/Nanbeige4.2-3B | Nanbeige4.2-3B-Base ♡406 | ♡40
Kwaipilot/KAT-Coder-V2.5-Dev ♡165
fdtn-ai/antares-1b | antares-350m ♡163 | ♡50
badtheorylabs/BTL-3 | BTL-3-Compact ♡60 | ♡26
ProCreations/grug-27b ♡57
PaddlePaddle/HPD-Parsing ♡57
FINAL-Bench/Aether-7B-5Attn | Aether-7B-5Attn-it ♡40 | ♡31
mindlab-research/Macaron-V1-Venti ♡35
AliveAi/Krea-2-Edit-Outfit-Transfer ♡34
Glint-Research/Glint-2 ♡28
amd/Instella-MoE-16B-A3B-Think ♡28
neuphonic/neutts-2e ♡26
Reza2kn/Bina-0.1 ♡24
joeygambino/joyai-echo-ltx23-echoVid-ltxAud-surgical ♡23
Trelis/tiron ♡21
BananaMind/BananaMind-2-Medium ♡14
ai9stars/G9v3-3B ♡13
📰 NVIDIA - NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning
NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they…

https://developer.nvidia.com/blog/nvidia-ising-enables-fully-automated-quantum-computer-calibration-with-enhanced-in-context-learning/


📰 NVIDIA - Six Agent Harness Capabilities for Higher Model Performance
Building a great AI agent isn’t just about choosing the right models. The harness is the architecture surrounding the model. How it renders context…

https://developer.nvidia.com/blog/six-agent-harness-capabilities-for-higher-model-performance/


📰 NVIDIA - NVIDIA Nemotron 3 Ultra Leads Open Models on Accuracy and Efficiency in Agentic RTL Coding
Modern chip design is increasingly limited by engineering time. Register transfer level (RTL) development and verification require specialized hardware…

https://developer.nvidia.com/blog/nvidia-nemotron-3-ultra-leads-open-models-on-accuracy-and-efficiency-in-agentic-rtl-coding/