📰 Google DeepMind - The latest AI news we announced in July 2026
https://blog.google/innovation-and-ai/technology/ai/google-ai-updates-july-2026/
https://blog.google/innovation-and-ai/technology/ai/google-ai-updates-july-2026/
Google
The latest AI news we announced in July 2026
Here are Google’s latest AI updates from July 2026
📰 Google AI Blog - A unified API for AI model routing
Google Cloud API Gateway now offers a model routing feature in Public Preview, allowing developers to dynamically route traffic to models like Gemini, Claude, or OpenAI OSS-GPT without hardcoding endpoints or managing open-source proxies. Developers can easily configure these routing rules directly within their OpenAPI 3.x specifications by mapping virtual model names to specific backend targets on a shared host. Once deployed, the Gateway acts as a serverless ingress layer that accepts standard OpenAI-compatible requests, automatically transcodes the payload to the native schema of the target model, and routes the traffic on the fly.
https://developers.googleblog.com/en/a-unified-api-for-ai-model-routing/
Google Cloud API Gateway now offers a model routing feature in Public Preview, allowing developers to dynamically route traffic to models like Gemini, Claude, or OpenAI OSS-GPT without hardcoding endpoints or managing open-source proxies. Developers can easily configure these routing rules directly within their OpenAPI 3.x specifications by mapping virtual model names to specific backend targets on a shared host. Once deployed, the Gateway acts as a serverless ingress layer that accepts standard OpenAI-compatible requests, automatically transcodes the payload to the native schema of the target model, and routes the traffic on the fly.
https://developers.googleblog.com/en/a-unified-api-for-ai-model-routing/
Googleblog
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Discover how developers can configure Google Cloud API Gateway to dynamically route OpenAI-compatible requests without managing open-source proxies.
📰 LMSys - SpecForge v0.3.0: a Unified Disaggregated and Colocated Speculative Decoding Stack, and New Open SpecBundle Draft Models
https://lmsys.org/blog/2026-08-04-specforge-v0-3
https://lmsys.org/blog/2026-08-04-specforge-v0-3
www.lmsys.org
SpecForge v0.3.0: a Unified Disaggregated and Colocated Speculative Decoding Stack, and New Open SpecBundle Draft Models
When we first released SpecForge, a training job owned both the frozen target model and the draft model being optimized. This made EAGLE3 draft-model training practical and directly compatible with SG...
📰 Anthropic - Mariano-Florentino (Tino) Cuéllar to join Anthropic as Chief Global Affairs Officer
https://www.anthropic.com/news/tino-cuellar
https://www.anthropic.com/news/tino-cuellar
Anthropic
Mariano-Florentino (Tino) Cuéllar to join Anthropic as Chief Global Affairs Officer
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
📰 OpenAI - New ways to learn and teach with ChatGPT Work and Codex
Explore new education plugins for ChatGPT Work and Codex that help K–12 teachers, college educators, and students learn, teach, research, and build.
https://openai.com/index/learn-teach-chatgpt-work-codex
📰 OpenAI - Apple is getting this wrong
OpenAI addresses Apple’s baseless lawsuit, corrects claims about its employees, and shares messages documenting what happened.
https://openai.com/index/apple-is-getting-this-wrong
📰 OpenAI - How we built a realtime system for responsive voice AI in six months
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.
https://openai.com/index/continuous-voice-interaction-with-gpt-live
Explore new education plugins for ChatGPT Work and Codex that help K–12 teachers, college educators, and students learn, teach, research, and build.
https://openai.com/index/learn-teach-chatgpt-work-codex
📰 OpenAI - Apple is getting this wrong
OpenAI addresses Apple’s baseless lawsuit, corrects claims about its employees, and shares messages documenting what happened.
https://openai.com/index/apple-is-getting-this-wrong
📰 OpenAI - How we built a realtime system for responsive voice AI in six months
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.
https://openai.com/index/continuous-voice-interaction-with-gpt-live
OpenAI
New ways to learn and teach with ChatGPT Work and Codex
Explore new education plugins for ChatGPT Work and Codex that help K–12 teachers, college educators, and students learn, teach, research, and build.
📰 NVIDIA - Beyond VLAs: How World Action Models Reshape Robot Manipulation
A central challenge in robotics is building policies that generalize beyond the demonstrations they’re trained on. A policy that succeeds in a training scene…
https://developer.nvidia.com/blog/beyond-vlas-how-world-action-models-reshape-robot-manipulation/
A central challenge in robotics is building policies that generalize beyond the demonstrations they’re trained on. A policy that succeeds in a training scene…
https://developer.nvidia.com/blog/beyond-vlas-how-world-action-models-reshape-robot-manipulation/
NVIDIA Technical Blog
Beyond VLAs: How World Action Models Reshape Robot Manipulation
A central challenge in robotics is building policies that generalize beyond the demonstrations they’re trained on. A policy that succeeds in a training scene often fails when object shapes, positions…
📰 PyTorch - PyTorch by the Sea: The inaugural Santa Cruz PyTorch Meetup
TL;DR The inaugural Santa Cruz PyTorch Meetup brought together 45 local engineers, students, and leaders for GPU/CUDA talks and lightning presentations on chemistry, plant health, and autonomous driving – demonstrating...
https://pytorch.org/blog/pytorch-by-the-sea-the-inaugural-santa-cruz-pytorch-meetup/
TL;DR The inaugural Santa Cruz PyTorch Meetup brought together 45 local engineers, students, and leaders for GPU/CUDA talks and lightning presentations on chemistry, plant health, and autonomous driving – demonstrating...
https://pytorch.org/blog/pytorch-by-the-sea-the-inaugural-santa-cruz-pytorch-meetup/
📰 Google DeepMind - Our WeatherNext 2 AI model demonstrated a massive leap forward in predicting cyclones.
https://blog.google/innovation-and-ai/models-and-research/google-deepmind/weathernext-2-cyclones/
https://blog.google/innovation-and-ai/models-and-research/google-deepmind/weathernext-2-cyclones/
Google
Our WeatherNext 2 AI model demonstrated a massive leap forward in predicting cyclones.
Google DeepMind’s WeatherNext 2 shows state-of-the-art accuracy in cyclone prediction.
📰 Google AI Blog - Agent Plugins package your skills, tools, and more
Agent Plugins 1.0.0 is a new, vendor-neutral directory specification—backed by Google, Amazon, Microsoft, and others—for packaging Agent Skills and MCP servers into a single portable unit. By standardizing the manifest (plugin.json) and utilizing a fixed directory layout, it eliminates the need for developers to maintain separate wrappers or configurations to support different AI coding agents and IDEs. Google has officially joined as a Core Maintainer and already rolled out support in the Agents CLI and Data Agent Kit, allowing developers to start building and distributing interoperable plugins today.
https://developers.googleblog.com/en/agent-plugins-package-your-skills-tools-and-more/
📰 Google AI Blog - Scaling AI Agent Infrastructure with the MCP Stateless updates
The 2026-07-28 Model Context Protocol (MCP) specification replaces legacy stateful constraints with a fully stateless core, enabling cloud-native horizontal scaling, serverless deployments, and standard round-robin load balancing. This architectural shift introduces standardized HTTP headers for efficient routing without deep packet inspection, caching controls, and Multi Round-Trip Requests (MRTR) to handle interactive and long-running tasks without blocking connections. Developers can immediately begin migrating their agentic applications to this highly scalable infrastructure using the newly available beta SDKs for Python, TypeScript, Go, and C#.
https://developers.googleblog.com/en/scaling-ai-agent-infrastructure-with-the-mcp-stateless-updates/
Agent Plugins 1.0.0 is a new, vendor-neutral directory specification—backed by Google, Amazon, Microsoft, and others—for packaging Agent Skills and MCP servers into a single portable unit. By standardizing the manifest (plugin.json) and utilizing a fixed directory layout, it eliminates the need for developers to maintain separate wrappers or configurations to support different AI coding agents and IDEs. Google has officially joined as a Core Maintainer and already rolled out support in the Agents CLI and Data Agent Kit, allowing developers to start building and distributing interoperable plugins today.
https://developers.googleblog.com/en/agent-plugins-package-your-skills-tools-and-more/
📰 Google AI Blog - Scaling AI Agent Infrastructure with the MCP Stateless updates
The 2026-07-28 Model Context Protocol (MCP) specification replaces legacy stateful constraints with a fully stateless core, enabling cloud-native horizontal scaling, serverless deployments, and standard round-robin load balancing. This architectural shift introduces standardized HTTP headers for efficient routing without deep packet inspection, caching controls, and Multi Round-Trip Requests (MRTR) to handle interactive and long-running tasks without blocking connections. Developers can immediately begin migrating their agentic applications to this highly scalable infrastructure using the newly available beta SDKs for Python, TypeScript, Go, and C#.
https://developers.googleblog.com/en/scaling-ai-agent-infrastructure-with-the-mcp-stateless-updates/
Googleblog
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Explore Agent Plugins 1.0.0, a new vendor-neutral standard backed by Google for packaging Agent Skills and MCP servers into portable, interoperable AI tools.
📰 LMSys - Full-Stack Performance Optimization of AR+DiT in SGL-Diffusion
https://lmsys.org/blog/2026-08-05-glmImage-optimization
https://lmsys.org/blog/2026-08-05-glmImage-optimization
www.lmsys.org
Full-Stack Performance Optimization of AR+DiT in SGL-Diffusion
- Replaces the HF backend with SRT to accelerate AR modeling and resolve parallelism conflicts, with dedicated TP for AR and SP for DiT
- Boosts hardware utilization via dynamic batching and enables e...
- Boosts hardware utilization via dynamic batching and enables e...
📰 PyTorch - PyTorch Conference North America Announces 2026 Keynotes
PyTorch Conference North America will be held in San Jose, California, on October 20–21, 2026. The 2026 keynote speakers are: Mark Collier, Executive Director, PyTorch Foundation Mazin Gilbert, Executive Director,...
https://pytorch.org/blog/pytorch-conference-north-america-announces-2026-keynotes/
PyTorch Conference North America will be held in San Jose, California, on October 20–21, 2026. The 2026 keynote speakers are: Mark Collier, Executive Director, PyTorch Foundation Mazin Gilbert, Executive Director,...
https://pytorch.org/blog/pytorch-conference-north-america-announces-2026-keynotes/
🔄 [GitHub Releases] turboderp-org/exllamav3 - 1.4.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.4.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.4.0
GitHub
Release 1.4.0 · turboderp-org/exllamav3
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs - Release 1.4.0 · turboderp-org/exllamav3
🆕 [HF Models] ATH-MaaS - Lumen
https://huggingface.co/ATH-MaaS/Lumen
🆕 [HF Models] ATH-MaaS - Lumen_wmt26
https://huggingface.co/ATH-MaaS/Lumen_wmt26
https://huggingface.co/ATH-MaaS/Lumen
🆕 [HF Models] ATH-MaaS - Lumen_wmt26
https://huggingface.co/ATH-MaaS/Lumen_wmt26
huggingface.co
ATH-MaaS/Lumen · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
🆕 [HF Models] Wan-AI - Wan2.2-Animate-2-14B-Distilled-Diffusers
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B-Distilled-Diffusers
🆕 [HF Models] Wan-AI - Wan2.2-Animate-2-14B-Diffusers
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B-Diffusers
🆕 [HF Models] Wan-AI - Wan2.2-Animate-2-14B
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B-Distilled-Diffusers
🆕 [HF Models] Wan-AI - Wan2.2-Animate-2-14B-Diffusers
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B-Diffusers
🆕 [HF Models] Wan-AI - Wan2.2-Animate-2-14B
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B
huggingface.co
Wan-AI/Wan2.2-Animate-2-14B-Distilled-Diffusers · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Google DeepMind - See what 5 builders are making with Gemini Omni
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-builders/
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-builders/
Google
See what 5 builders are making with Gemini Omni
Gemini Omni makes creating videos as easy as having a conversation. Here’s how five people use it to edit videos and visualize ideas.
📰 LMSys - HPC-Ops × SGLang: High-Performance Attention, Router GEMM, and MoE Kernels from Tencent Hunyuan
https://lmsys.org/blog/2026-08-07-hpc-ops-sglang
https://lmsys.org/blog/2026-08-07-hpc-ops-sglang
www.lmsys.org
HPC-Ops × SGLang: High-Performance Attention, Router GEMM, and MoE Kernels from Tencent Hunyuan
HPC-Ops is an open-source operator library for LLM inference, deployed in Tencent's large-scale production serving. Its core operators, including Dynamic Attention and Fused MoE, play a critical role ...