📰 NVIDIA - How to Self-Host a Validated AI Coding Assistant with NVIDIA NeMo Guardrails
Deploying an AI coding assistant in a regulated, sovereign, or source-sensitive environment, often comes with challenges. Three common issues are: the source…
https://developer.nvidia.com/blog/how-to-self-host-a-validated-ai-coding-assistant-with-nvidia-nemo-guardrails/
Deploying an AI coding assistant in a regulated, sovereign, or source-sensitive environment, often comes with challenges. Three common issues are: the source…
https://developer.nvidia.com/blog/how-to-self-host-a-validated-ai-coding-assistant-with-nvidia-nemo-guardrails/
NVIDIA Technical Blog
How to Self-Host a Validated AI Coding Assistant with NVIDIA NeMo Guardrails
Deploying an AI coding assistant in a regulated, sovereign, or source-sensitive environment, often comes with challenges. Three common issues are: the source cannot leave the network…
📰 LMSys - Towards Blackwell-Native 8-bit and 4-bit RL: End-to-End MXFP8 and NVFP4 RL in Miles
https://lmsys.org/blog/2026-07-29-mxfp8-nvfp4-rl
https://lmsys.org/blog/2026-07-29-mxfp8-nvfp4-rl
www.lmsys.org
Towards Blackwell-Native 8-bit and 4-bit RL: End-to-End MXFP8 and NVFP4 RL in Miles
TL;DR: We implemented two Blackwell-native RL recipes in Miles: end-to-end MXFP8 and per-token NVFP4 for MoE experts. Both are supported by fine-grained precision control across checkpoint conversion,...
📰 HuggingFace - GPU Management: Why Idle GPUs Are the New Grounded Aircraft
https://huggingface.co/blog/Dharma-AI/gpu-management
https://huggingface.co/blog/Dharma-AI/gpu-management
huggingface.co
GPU Management: Why Idle GPUs Are the New Grounded Aircraft
A Blog post by Dharma-AI on Hugging Face
📰 PyTorch - FBTriton Infra: Upstream Ingestion, Hierarchical Validation, Ideals vs Realities
TL:DR Learn how Meta’s FBTriton infrastructure powers custom GPU compiler innovations like TLX and autoWS while staying synced with upstream Triton using agentic ingestion and a stratified L1/L2/L3 validation framework....
https://pytorch.org/blog/fbtriton-infra-upstream-ingestion-hierarchical-validation-ideals-vs-realities/
TL:DR Learn how Meta’s FBTriton infrastructure powers custom GPU compiler innovations like TLX and autoWS while staying synced with upstream Triton using agentic ingestion and a stratified L1/L2/L3 validation framework....
https://pytorch.org/blog/fbtriton-infra-upstream-ingestion-hierarchical-validation-ideals-vs-realities/
🔓 Google Model Cards - Gemini Robotics On-Device 2
https://deepmind.google/models/model-cards/gemini-robotics-on-device-2/
🔓 Google Model Cards - Gemini Robotics-ER 2
https://deepmind.google/models/model-cards/gemini-robotics-er-2/
https://deepmind.google/models/model-cards/gemini-robotics-on-device-2/
🔓 Google Model Cards - Gemini Robotics-ER 2
https://deepmind.google/models/model-cards/gemini-robotics-er-2/
Google DeepMind
Gemini Robotics On-Device 2 - Model Card
📰 Google DeepMind - Introducing Gemini Robotics ER 2
https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-robotics-er-2/
https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-robotics-er-2/
Google
Introducing Gemini Robotics ER 2
Gemini Robotics ER 2 is a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
📰 LMSys - RadixArk Joins Forces with Google to Bring Full SGLang Features to TPUs
https://lmsys.org/blog/2026-07-30-sglang-google-tpu
https://lmsys.org/blog/2026-07-30-sglang-google-tpu
www.lmsys.org
RadixArk Joins Forces with Google to Bring Full SGLang Features to TPUs
RadixArk and Google Cloud are partnering to bring SGLang to TPUs, giving developers ultimate flexibility for running workloads on their choice of hardware.
👏1
📰 Google Labs - Take a look at short films created by our latest group of artists in Google’s Flow Sessions program.
https://blog.google/innovation-and-ai/models-and-research/google-labs/google-flow-sessions-short-films/
https://blog.google/innovation-and-ai/models-and-research/google-labs/google-flow-sessions-short-films/
Google
Take a look at short films created by our latest group of artists in Google’s Flow Sessions program.
We’re sharing a look at the short films created by our latest group of artists in Google’s Flow Sessions program.
📰 NVIDIA - Four Ways to Deploy More Secure AI Agents
Knowledge workers are increasingly integrating AI agents into their workflows. Agents that function as “digital coworkers” offer clear benefits. For example…
https://developer.nvidia.com/blog/four-ways-to-deploy-more-secure-ai-agents/
Knowledge workers are increasingly integrating AI agents into their workflows. Agents that function as “digital coworkers” offer clear benefits. For example…
https://developer.nvidia.com/blog/four-ways-to-deploy-more-secure-ai-agents/
NVIDIA Technical Blog
Four Ways to Deploy More Secure AI Agents
Knowledge workers are increasingly integrating AI agents into their workflows. Agents that function as “digital coworkers” offer clear benefits. For example, they can review a bug report…
🔄 [GitHub Releases] PygmalionAI/aphrodite-engine - v0.23.0
https://github.com/dphnAI/sonar/releases/tag/v0.23.0
https://github.com/dphnAI/sonar/releases/tag/v0.23.0
GitHub
Release v0.23.0 · dphnAI/sonar
What's Changed
fix(release): apply manylinux tag before PyPI upload (795396a3a) by AlpinDale
fix(release): attach all wheels to GitHub releases (#1749) (ba857e9b9) by @AlpinDale
fix(build): pa...
fix(release): apply manylinux tag before PyPI upload (795396a3a) by AlpinDale
fix(release): attach all wheels to GitHub releases (#1749) (ba857e9b9) by @AlpinDale
fix(build): pa...
🆕 [HF Models] meituan-longcat - LongCat-Flash-Lite-Sparse
https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse
https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse
huggingface.co
meituan-longcat/LongCat-Flash-Lite-Sparse · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 NVIDIA - Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference
As agentic and long-context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1).
https://developer.nvidia.com/blog/co-designing-ai-model-attention-for-fast-interactive-long-context-inference/
📰 NVIDIA - Run High-Performance Core Math at Scale with NVIDIA nvmath-python
NVIDIA nvmath-python is a library designed to bridge the gap between the Python scientific community and NVIDIA CUDA-X math libraries. It gives Python users…
https://developer.nvidia.com/blog/run-high-performance-core-math-at-scale-with-nvidia-nvmath-python/
As agentic and long-context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1).
https://developer.nvidia.com/blog/co-designing-ai-model-attention-for-fast-interactive-long-context-inference/
📰 NVIDIA - Run High-Performance Core Math at Scale with NVIDIA nvmath-python
NVIDIA nvmath-python is a library designed to bridge the gap between the Python scientific community and NVIDIA CUDA-X math libraries. It gives Python users…
https://developer.nvidia.com/blog/run-high-performance-core-math-at-scale-with-nvidia-nvmath-python/
NVIDIA Technical Blog
Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference
As agentic and long-context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1). Because attention now dominates that cost…
🔄 [GitHub Releases] invoke-ai/InvokeAI - InvokeAI 6.14.0 (release candidate 1)
https://github.com/invoke-ai/InvokeAI/releases/tag/v6.14.0-rc1
https://github.com/invoke-ai/InvokeAI/releases/tag/v6.14.0-rc1
GitHub
Release InvokeAI 6.14.0 (release candidate 1) · invoke-ai/InvokeAI
This is a big release that adds many new user visible features including:
Video generation support via Wan 2.2.
Krea.2-Turbo and Raw model support
Ernie Turbo model support
Ideogram 4 support
Anim...
Video generation support via Wan 2.2.
Krea.2-Turbo and Raw model support
Ernie Turbo model support
Ideogram 4 support
Anim...
🔄 [GitHub Releases] turboderp-org/exllamav3 - 1.3.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.3.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.3.0
GitHub
Release 1.3.0 · turboderp-org/exllamav3
Preliminary support for DeepseekV3 (validated against JoyAI-LLM-Flash and Moonlight-16B-A3B, no routing groups yet)
Second-tier CPU K/V cache, and more intelligent page and checkpoint eviction poli...
Second-tier CPU K/V cache, and more intelligent page and checkpoint eviction poli...
📰 Google AI Blog - Agent and Model Evaluations in Gemini Enterprise Agent Platform are now GA
Agent Platform's evaluation service is now generally available, providing developers with a unified engine to measure agent quality consistently across local development experiments and live production traffic. You can evaluate agents using over 20 pre-built metrics, DeepMind-backed adaptive rubrics, or custom code-based and LLM-as-a-judge metrics stored in a centralized, versioned registry. The service integrates directly into existing workflows via the Agent Platform SDK, agents-cli, and ADK, offering built-in user and environment simulators to automate complex multi-turn testing and streamline CI pipelines.
https://developers.googleblog.com/en/agent-and-model-evaluations-in-gemini-enterprise-agent-platform-are-now-ga/
Agent Platform's evaluation service is now generally available, providing developers with a unified engine to measure agent quality consistently across local development experiments and live production traffic. You can evaluate agents using over 20 pre-built metrics, DeepMind-backed adaptive rubrics, or custom code-based and LLM-as-a-judge metrics stored in a centralized, versioned registry. The service integrates directly into existing workflows via the Agent Platform SDK, agents-cli, and ADK, offering built-in user and environment simulators to automate complex multi-turn testing and streamline CI pipelines.
https://developers.googleblog.com/en/agent-and-model-evaluations-in-gemini-enterprise-agent-platform-are-now-ga/
Googleblog
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Measure AI agent quality from development to production with consistent metrics. Agent and model evaluations in Agent Platform are now generally available.
📰 Anthropic - Investigating three real-world incidents in our cybersecurity evaluations
https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
Anthropic
Investigating three real-world incidents in our cybersecurity evaluations
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems…
📰 OpenAI - Ten advances in mathematics and theoretical computer science
https://openai.com/index/ten-advances-in-mathematics
📰 OpenAI - Building abundant intelligence
https://openai.com/index/building-abundant-intelligence
📰 OpenAI - Advancing the price-performance frontier with GPT-5.6
https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6
📰 OpenAI - How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores
📰 OpenAI - Accelerating scientific discovery with ChatGPT for Academic Researchers
https://openai.com/index/chatgpt-for-academic-researchers
📰 OpenAI - How GPT-5.6 fuses frontier intelligence with frontier efficiency
https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency
📰 OpenAI - Scientific computing in the age of agentic AI
https://openai.com/index/scientific-computing-agentic-ai
📰 OpenAI - How AI is expanding what people do at work
https://openai.com/index/how-ai-is-expanding-what-people-do-at-work
https://openai.com/index/ten-advances-in-mathematics
📰 OpenAI - Building abundant intelligence
https://openai.com/index/building-abundant-intelligence
📰 OpenAI - Advancing the price-performance frontier with GPT-5.6
https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6
📰 OpenAI - How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores
📰 OpenAI - Accelerating scientific discovery with ChatGPT for Academic Researchers
https://openai.com/index/chatgpt-for-academic-researchers
📰 OpenAI - How GPT-5.6 fuses frontier intelligence with frontier efficiency
https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency
📰 OpenAI - Scientific computing in the age of agentic AI
https://openai.com/index/scientific-computing-agentic-ai
📰 OpenAI - How AI is expanding what people do at work
https://openai.com/index/how-ai-is-expanding-what-people-do-at-work
OpenAI
Ten advances in mathematics and theoretical computer science
OpenAI shares new results on long-standing open problems in mathematics and theoretical computer science, including advances in geometry, cryptography, and complexity.