🔄 [GitHub Releases] vllm-project/vllm - v0.25.1
https://github.com/vllm-project/vllm/releases/tag/v0.25.1
https://github.com/vllm-project/vllm/releases/tag/v0.25.1
GitHub
Release v0.25.1 · vllm-project/vllm
vLLM v0.25.1
Highlights
This release features 2 commits from 2 contributors (1 new)!
v0.25.1 is a patch release containing two targeted bug fixes on top of v0.25.0.
Bug Fixes
Avoid blocking model ...
Highlights
This release features 2 commits from 2 contributors (1 new)!
v0.25.1 is a patch release containing two targeted bug fixes on top of v0.25.0.
Bug Fixes
Avoid blocking model ...
🆕 [HF Models] nvidia - Nemotron-3-Embed-8B-BF16
https://huggingface.co/nvidia/Nemotron-3-Embed-8B-BF16
🆕 [HF Models] nvidia - Nemotron-3-Embed-1B-BF16
https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16
https://huggingface.co/nvidia/Nemotron-3-Embed-8B-BF16
🆕 [HF Models] nvidia - Nemotron-3-Embed-1B-BF16
https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16
huggingface.co
nvidia/Nemotron-3-Embed-8B-BF16 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
🔄 [GitHub Releases] turboderp-org/exllamav3 - 1.0.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.0.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.0.0
GitHub
Release 1.0.0 · turboderp-org/exllamav3
-> Small writeup with charts.
Remove flash-attention-2 and xformers dependencies
New attention kernel with online cache quantization, dual input for SWA layers and attention sinks
New conv1d ke...
Remove flash-attention-2 and xformers dependencies
New attention kernel with online cache quantization, dual input for SWA layers and attention sinks
New conv1d ke...
📰 Google DeepMind - Reconstructing Pelé’s “lost” goal
https://blog.google/innovation-and-ai/models-and-research/google-deepmind/reconstructing-peles-lost-goal/
https://blog.google/innovation-and-ai/models-and-research/google-deepmind/reconstructing-peles-lost-goal/
Google
Reconstructing Pelé’s “lost” goal
See how Google DeepMind AI technology reconstructed Pelé’s legendary 1959 lost goal at Rua Javari in our new mini-documentary.
📰 Google AI Blog - Unlocking the Next Era of On-Device AI with Google Tensor and Pixel
At Google I/O Connect India, Google showcased the future of 100% private, on-device AI powered by the custom Tensor SoC and TPU for the new Pixel 10 family. The event debuted the lightweight Gemma 4 E2B model, which runs natively on the device to enable completely offline multimodal features like AI chat, real-time image recognition, and personal agent tasks. Developers can start building these secure, edge-based applications today by accessing the newly announced Tensor SDK beta and its accompanying open-source resources.
https://developers.googleblog.com/en/unlocking-the-next-era-of-on-device-ai-with-google-tensor-and-pixel/
At Google I/O Connect India, Google showcased the future of 100% private, on-device AI powered by the custom Tensor SoC and TPU for the new Pixel 10 family. The event debuted the lightweight Gemma 4 E2B model, which runs natively on the device to enable completely offline multimodal features like AI chat, real-time image recognition, and personal agent tasks. Developers can start building these secure, edge-based applications today by accessing the newly announced Tensor SDK beta and its accompanying open-source resources.
https://developers.googleblog.com/en/unlocking-the-next-era-of-on-device-ai-with-google-tensor-and-pixel/
Googleblog
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Build the next generation of offline AI apps. Read our Google I/O Connect recap to learn about Gemma 4 for Pixel 10 and access the new Tensor SDK beta.
📰 LMSys - Serving GLM5.2 NVFP4 Agentic Workload with SGLang: Reaching 500 TPS in 2 Weeks
https://lmsys.org/blog/2026-07-13-glm52-optimization
https://lmsys.org/blog/2026-07-13-glm52-optimization
www.lmsys.org
Serving GLM5.2 NVFP4 Agentic Workload with SGLang: Reaching 500 TPS in 2 Weeks
- More than 500 TPS on 8xB300 (bs=1)
- Sync free speculative decoding for GLM 5.2 MTP
- Built-in IndexShare MTP with Spec V2
- 2.33x faster TopK-V2 for ISL 80k
- Indexer prologue fusion
- Gemm kernels...
- Sync free speculative decoding for GLM 5.2 MTP
- Built-in IndexShare MTP with Spec V2
- 2.33x faster TopK-V2 for ISL 80k
- Indexer prologue fusion
- Gemm kernels...
📰 OpenAI - How to manage AI investments in the agentic era
https://openai.com/index/managing-ai-investments-in-agentic-era
📰 OpenAI - How data science teams use ChatGPT Work
https://openai.com/academy/codex-for-work/how-data-science-teams-use-codex
📰 OpenAI - How sales teams use ChatGPT Work
https://openai.com/academy/codex-for-work/how-sales-teams-use-codex
https://openai.com/index/managing-ai-investments-in-agentic-era
📰 OpenAI - How data science teams use ChatGPT Work
https://openai.com/academy/codex-for-work/how-data-science-teams-use-codex
📰 OpenAI - How sales teams use ChatGPT Work
https://openai.com/academy/codex-for-work/how-sales-teams-use-codex
OpenAI
How to manage AI investments in the agentic era
Learn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling high-value workflows.
📰 NVIDIA - Lessons From the Leaderboard: What 5,000+ Kagglers Taught Us About Improving AI Reasoning
The NVIDIA Nemotron Model Reasoning Challenge invited the Kaggle community to explore a focused question: What techniques can improve reasoning accuracy when…
https://developer.nvidia.com/blog/lessons-from-the-leaderboard-what-5000-kagglers-taught-us-about-improving-ai-reasoning/
📰 NVIDIA - How to Run an Autoresearch Workflow with RL Agent Skills and NVIDIA NeMo
Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes…
https://developer.nvidia.com/blog/how-to-run-an-autoresearch-workflow-with-rl-agent-skills-and-nvidia-nemo/
📰 NVIDIA - Post-Train NVIDIA Cosmos 3 in One Day Using Agent Skills
What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning…
https://developer.nvidia.com/blog/post-train-nvidia-cosmos-3-in-one-day-using-agent-skills/
The NVIDIA Nemotron Model Reasoning Challenge invited the Kaggle community to explore a focused question: What techniques can improve reasoning accuracy when…
https://developer.nvidia.com/blog/lessons-from-the-leaderboard-what-5000-kagglers-taught-us-about-improving-ai-reasoning/
📰 NVIDIA - How to Run an Autoresearch Workflow with RL Agent Skills and NVIDIA NeMo
Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes…
https://developer.nvidia.com/blog/how-to-run-an-autoresearch-workflow-with-rl-agent-skills-and-nvidia-nemo/
📰 NVIDIA - Post-Train NVIDIA Cosmos 3 in One Day Using Agent Skills
What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning…
https://developer.nvidia.com/blog/post-train-nvidia-cosmos-3-in-one-day-using-agent-skills/
NVIDIA Technical Blog
Lessons From the Leaderboard: What 5,000+ Kagglers Taught Us About Improving AI Reasoning
The NVIDIA Nemotron Model Reasoning Challenge invited the Kaggle community to explore a focused question: What techniques can improve reasoning accuracy when everyone starts from the same open model…
🆕 [HF Models] microsoft - bitnet-embedding-0.6b
https://huggingface.co/microsoft/bitnet-embedding-0.6b
🆕 [HF Models] microsoft - bitnet-embedding-270m
https://huggingface.co/microsoft/bitnet-embedding-270m
https://huggingface.co/microsoft/bitnet-embedding-0.6b
🆕 [HF Models] microsoft - bitnet-embedding-270m
https://huggingface.co/microsoft/bitnet-embedding-270m
huggingface.co
microsoft/bitnet-embedding-0.6b · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 HuggingFace - Introducing Real World VoiceEQ: Measuring the human quality of voice AI
https://huggingface.co/blog/real-world-voiceeq
https://huggingface.co/blog/real-world-voiceeq
📰 PyTorch - Triton Plugin Extensions: Enabling TLX and Custom Compiler Passes Out of the Box
TLDR The PyTorch-Triton 3.7 release introduces the Triton Plugin Extensions system, a framework for dynamically loading custom compiler passes, dialects (including their ops), and DSL extensions into upstream Triton at...
https://pytorch.org/blog/triton-plugin-extensions-enabling-tlx-and-custom-compiler-passes-out-of-the-box/
TLDR The PyTorch-Triton 3.7 release introduces the Triton Plugin Extensions system, a framework for dynamically loading custom compiler passes, dialects (including their ops), and DSL extensions into upstream Triton at...
https://pytorch.org/blog/triton-plugin-extensions-enabling-tlx-and-custom-compiler-passes-out-of-the-box/
📰 HuggingFace - What building Shippy taught us about building agents
https://huggingface.co/blog/allenai/shippy-tech-blog
📰 HuggingFace - Model Routing Is Simple. Until It Isn’t.
https://huggingface.co/blog/ibm-research/model-routing-is-simple-until-it-isnt
https://huggingface.co/blog/allenai/shippy-tech-blog
📰 HuggingFace - Model Routing Is Simple. Until It Isn’t.
https://huggingface.co/blog/ibm-research/model-routing-is-simple-until-it-isnt
📰 LMSys - SGLang and Miles Add Day-0 Support for Inkling, a Frontier Multimodal Model
https://lmsys.org/blog/2026-07-15-inkling-day0-support
https://lmsys.org/blog/2026-07-15-inkling-day0-support
www.lmsys.org
SGLang and Miles Add Day-0 Support for Inkling, a Frontier Multimodal Model
We're excited to partner with the Thinking Machines team to bring Day-0 support for Inkling to SGLang and Miles, with dedicated optimizations for its new architecture and broad feature coverage, and w...
🔓 HuggingFace - Welcome Inkling by Thinking Machines
https://huggingface.co/blog/thinkingmachines-inkling
https://huggingface.co/blog/thinkingmachines-inkling
huggingface.co
Welcome Inkling by Thinking Machines
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Anthropic Research - How Canada uses Claude: Findings from the Anthropic Economic Index
https://www.anthropic.com/research/how-canada-uses-claude
https://www.anthropic.com/research/how-canada-uses-claude
Anthropic
How Canada uses Claude: Findings from the Anthropic Economic Index
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
📰 Anthropic - Introducing Claude for Teachers
https://www.anthropic.com/news/claude-for-teachers
📰 Anthropic - Anthropic commits $10 million to Canadian AI research
https://www.anthropic.com/news/canadian-ai-research
https://www.anthropic.com/news/claude-for-teachers
📰 Anthropic - Anthropic commits $10 million to Canadian AI research
https://www.anthropic.com/news/canadian-ai-research
Anthropic
Introducing Claude for Teachers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.