📰 PyTorch - Helion on TPU: Towards Hardware Heterogeneous Kernel Authoring
TL;DR Helion is PyTorch’s high-level DSL for writing performance-portable ML kernels. Partnering with Google, we have built a TPU backend that compiles Helion kernels to Pallas, providing a PyTorch-friendly way...
https://pytorch.org/blog/helion-on-tpu-towards-hardware-heterogeneous-kernel-authoring/
TL;DR Helion is PyTorch’s high-level DSL for writing performance-portable ML kernels. Partnering with Google, we have built a TPU backend that compiles Helion kernels to Pallas, providing a PyTorch-friendly way...
https://pytorch.org/blog/helion-on-tpu-towards-hardware-heterogeneous-kernel-authoring/
📰 OpenAI - Launching Health in ChatGPT
https://openai.com/index/health-in-chatgpt
🔓 OpenAI - NTT DATA Group cuts incident analysis to 30 minutes with Codex
https://openai.com/index/ntt-data
https://openai.com/index/health-in-chatgpt
🔓 OpenAI - NTT DATA Group cuts incident analysis to 30 minutes with Codex
https://openai.com/index/ntt-data
OpenAI
Launching Health in ChatGPT
Health in ChatGPT now lets eligible U.S. users securely connect medical records and Apple Health to get more personalized insights and better understand their health.
📰 NVIDIA - Start Customizing NVIDIA Nemotron 3 Nano with Prime Intellect Lab in Minutes
Customization is what enables developers to take a general model and tailor it to use cases, domains, languages, and more. However, customization comes with a…
https://developer.nvidia.com/blog/start-customizing-nvidia-nemotron-3-nano-with-prime-intellect-lab-in-minutes/
Customization is what enables developers to take a general model and tailor it to use cases, domains, languages, and more. However, customization comes with a…
https://developer.nvidia.com/blog/start-customizing-nvidia-nemotron-3-nano-with-prime-intellect-lab-in-minutes/
NVIDIA Technical Blog
Start Customizing NVIDIA Nemotron 3 Nano with Prime Intellect Lab in Minutes
Customization is what enables developers to take a general model and tailor it to use cases, domains, languages, and more. However, customization comes with a few challenges.
🆕 [HF Models] swiss-ai - Apertus-v1.5-70B
https://huggingface.co/swiss-ai/Apertus-v1.5-70B
🆕 [HF Models] swiss-ai - Apertus-v1.5-8B
https://huggingface.co/swiss-ai/Apertus-v1.5-8B
https://huggingface.co/swiss-ai/Apertus-v1.5-70B
🆕 [HF Models] swiss-ai - Apertus-v1.5-8B
https://huggingface.co/swiss-ai/Apertus-v1.5-8B
huggingface.co
swiss-ai/Apertus-v1.5-70B · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Google AI Blog - Run Ray on TPU, Part 2: Ray AI libraries
This second installment explores how Ray’s higher-level libraries—Serve, Data, and Train—abstract the complexities of running AI workloads on Google's TPU slices. Ray Serve uses a simple topology configuration to correctly gang-schedule large multi-host models, while Ray Data eliminates data-loading bottlenecks by feeding accelerators directly with native JAX batches. Finally, JaxTrainer streamlines distributed training across TPUs by automatically handling cross-slice coordination, checkpointing, and fault tolerance.
https://developers.googleblog.com/en/run-ray-on-tpu-part-2-ray-ai-libraries/
This second installment explores how Ray’s higher-level libraries—Serve, Data, and Train—abstract the complexities of running AI workloads on Google's TPU slices. Ray Serve uses a simple topology configuration to correctly gang-schedule large multi-host models, while Ray Data eliminates data-loading bottlenecks by feeding accelerators directly with native JAX batches. Finally, JaxTrainer streamlines distributed training across TPUs by automatically handling cross-slice coordination, checkpointing, and fault tolerance.
https://developers.googleblog.com/en/run-ray-on-tpu-part-2-ray-ai-libraries/
Googleblog
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Learn how to scale AI workloads on TPU slices using Ray Serve for LLM deployment, Ray Data for fast JAX pipelines, and JaxTrainer for distributed training.
🔓 xAI - Bringing Grok 4.5 to iOS, Android, Web, and X
https://x.ai//news/grok-4-5-everywhere
🔓 xAI - Workflows in Grok Build
https://x.ai//news/workflows
🔓 xAI - Grok in Google Workspace
https://x.ai//news/introducing-google-workspace-addon
https://x.ai//news/grok-4-5-everywhere
🔓 xAI - Workflows in Grok Build
https://x.ai//news/workflows
🔓 xAI - Grok in Google Workspace
https://x.ai//news/introducing-google-workspace-addon
x.ai
Bringing Grok 4.5 to iOS, Android, Web, and X
Grok 4.5, our most intelligent model yet, is now on grok.com, X, iOS, and Android.
📰 NVIDIA - ModelExpress: Distributing Model Artifacts at the Speed of Light
Every byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly. To make things even worse…
https://developer.nvidia.com/blog/modelexpress-distributing-model-artifacts-at-the-speed-of-light/
Every byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly. To make things even worse…
https://developer.nvidia.com/blog/modelexpress-distributing-model-artifacts-at-the-speed-of-light/
NVIDIA Technical Blog
ModelExpress: Distributing Model Artifacts at the Speed of Light
Every byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly. To make things even worse, moving these model weights around the cluster…
🔄 [GitHub Releases] sgl-project/sglang - v0.5.16
https://github.com/sgl-project/sglang/releases/tag/v0.5.16
https://github.com/sgl-project/sglang/releases/tag/v0.5.16
GitHub
Release v0.5.16 · sgl-project/sglang
Highlights
574 PRs from 169 contributors.
DSpark: confidence-driven speculative decoding: A new speculative algorithm. It drafts semi-autoregressively in blocks, then sizes each verify window from ...
574 PRs from 169 contributors.
DSpark: confidence-driven speculative decoding: A new speculative algorithm. It drafts semi-autoregressively in blocks, then sizes each verify window from ...
🔄 [GitHub Releases] vllm-project/vllm - v0.26.0
https://github.com/vllm-project/vllm/releases/tag/v0.26.0
https://github.com/vllm-project/vllm/releases/tag/v0.26.0
GitHub
Release v0.26.0 · vllm-project/vllm
vLLM v0.26.0 Release Notes
Highlights
This release features 411 commits from 212 contributors (61 new)!
New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA g...
Highlights
This release features 411 commits from 212 contributors (61 new)!
New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA g...
🔄 [GitHub Releases] turboderp-org/exllamav3 - 1.2.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.2.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.2.0
GitHub
Release 1.2.0 · turboderp-org/exllamav3
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs - Release 1.2.0 · turboderp-org/exllamav3
📰 Claude Blog - How the product designer who built Claude Design uses it to explore ideas before building them
https://claude.com/blog/how-the-product-designer-who-built-claude-design-uses-it-to-explore-ideas-before-building-them
📰 Claude Blog - The new rules of context engineering for Claude 5 generation models
https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models
📰 Claude Blog - Claude models explained: choosing the best model for your use case
https://claude.com/blog/claude-models-explained-choosing-the-best-model-for-your-use-case
📰 Claude Blog - Four role-based certifications for the people who put Claude to work for customers
https://claude.com/blog/four-role-based-claude-certifications
📰 Claude Blog - Think through hard problems in voice mode
https://claude.com/blog/think-through-hard-problems-in-voice-mode
📰 Claude Blog - Building verification loops in Claude Code with skills
https://claude.com/blog/building-verification-loops-in-claude-code-with-skills
📰 Claude Blog - How Outtake built a cyber investigator on Claude
https://claude.com/blog/how-outtake-built-a-cyber-investigator-on-claude
https://claude.com/blog/how-the-product-designer-who-built-claude-design-uses-it-to-explore-ideas-before-building-them
📰 Claude Blog - The new rules of context engineering for Claude 5 generation models
https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models
📰 Claude Blog - Claude models explained: choosing the best model for your use case
https://claude.com/blog/claude-models-explained-choosing-the-best-model-for-your-use-case
📰 Claude Blog - Four role-based certifications for the people who put Claude to work for customers
https://claude.com/blog/four-role-based-claude-certifications
📰 Claude Blog - Think through hard problems in voice mode
https://claude.com/blog/think-through-hard-problems-in-voice-mode
📰 Claude Blog - Building verification loops in Claude Code with skills
https://claude.com/blog/building-verification-loops-in-claude-code-with-skills
📰 Claude Blog - How Outtake built a cyber investigator on Claude
https://claude.com/blog/how-outtake-built-a-cyber-investigator-on-claude
Claude
How the product designer who built Claude Design uses it | Claude by Anthropic
Nate Parrott, a product designer at Anthropic, shares how he uses Claude Design (in beta) to explore, iterate on, and share visual ideas early, from product prototypes to slide decks and animations.
🗓️ Weekly GitHub Activity
🦙 llama.cpp
└ Release: b10068 → b10107
└ 39 commits
- Consolidated memory mapping and locking CLI options into a unified --load-mode argument #20834
- Added support for Laguna XS.2 and M.1 models #25165
- Improved DeepSeek-V4 support with softplus CUDA kernels, APE tensor fixes, and chat template updates #25896, #25945, #25414
- Fixed coordinate scaling in Qwen3-VL via align_corners interpolation #25781 and HunyuanVL XD-RoPE conversion #25514
- Enabled automatic speculative decoding type inference and sidecar model resolution from draft repositories #25955, #25989
- Upgraded CUDA backend with device-side GET_ROWS support for k-quants, i-quants, and mxfp4 #25962, alongside vectorized same-type copy optimization #25929
- Refactored Vulkan queue submission to bypass host-side locking via per-instance mutexes and unique handles #23570
- Added depthwise 2D convolution (CONV_2D_DW) kernel for WebGPU #25847
- Updated WebUI with bulk conversation management #25815, symbolic math support in the JS sandbox using Nerdamer #25948, and a default reasoning mode selector #25846
🔗 All changes | Latest release
🎨 stable-diffusion.cpp
└ Release: master-782-b290693 → master-795-87a0177
└ 13 commits
- Added support for Hunyuan Video 1.5 #1795
- Added IP-Adapter support for SD 1.5 and SDXL #1803 with CFG conditioning fixes #1815
- Added Mage-Flow support #1808
🔗 All changes | Latest release
🤗 Fresh models trending on HuggingFace:
Nanbeige/Nanbeige4.2-3B | Nanbeige4.2-3B-Base ♡406 | ♡40
Kwaipilot/KAT-Coder-V2.5-Dev ♡165
fdtn-ai/antares-1b | antares-350m ♡163 | ♡50
badtheorylabs/BTL-3 | BTL-3-Compact ♡60 | ♡26
ProCreations/grug-27b ♡57
PaddlePaddle/HPD-Parsing ♡57
FINAL-Bench/Aether-7B-5Attn | Aether-7B-5Attn-it ♡40 | ♡31
mindlab-research/Macaron-V1-Venti ♡35
AliveAi/Krea-2-Edit-Outfit-Transfer ♡34
Glint-Research/Glint-2 ♡28
amd/Instella-MoE-16B-A3B-Think ♡28
neuphonic/neutts-2e ♡26
Reza2kn/Bina-0.1 ♡24
joeygambino/joyai-echo-ltx23-echoVid-ltxAud-surgical ♡23
Trelis/tiron ♡21
BananaMind/BananaMind-2-Medium ♡14
ai9stars/G9v3-3B ♡13
🦙 llama.cpp
└ Release: b10068 → b10107
└ 39 commits
- Consolidated memory mapping and locking CLI options into a unified --load-mode argument #20834
- Added support for Laguna XS.2 and M.1 models #25165
- Improved DeepSeek-V4 support with softplus CUDA kernels, APE tensor fixes, and chat template updates #25896, #25945, #25414
- Fixed coordinate scaling in Qwen3-VL via align_corners interpolation #25781 and HunyuanVL XD-RoPE conversion #25514
- Enabled automatic speculative decoding type inference and sidecar model resolution from draft repositories #25955, #25989
- Upgraded CUDA backend with device-side GET_ROWS support for k-quants, i-quants, and mxfp4 #25962, alongside vectorized same-type copy optimization #25929
- Refactored Vulkan queue submission to bypass host-side locking via per-instance mutexes and unique handles #23570
- Added depthwise 2D convolution (CONV_2D_DW) kernel for WebGPU #25847
- Updated WebUI with bulk conversation management #25815, symbolic math support in the JS sandbox using Nerdamer #25948, and a default reasoning mode selector #25846
🔗 All changes | Latest release
🎨 stable-diffusion.cpp
└ Release: master-782-b290693 → master-795-87a0177
└ 13 commits
- Added support for Hunyuan Video 1.5 #1795
- Added IP-Adapter support for SD 1.5 and SDXL #1803 with CFG conditioning fixes #1815
- Added Mage-Flow support #1808
🔗 All changes | Latest release
🤗 Fresh models trending on HuggingFace:
Nanbeige/Nanbeige4.2-3B | Nanbeige4.2-3B-Base ♡406 | ♡40
Kwaipilot/KAT-Coder-V2.5-Dev ♡165
fdtn-ai/antares-1b | antares-350m ♡163 | ♡50
badtheorylabs/BTL-3 | BTL-3-Compact ♡60 | ♡26
ProCreations/grug-27b ♡57
PaddlePaddle/HPD-Parsing ♡57
FINAL-Bench/Aether-7B-5Attn | Aether-7B-5Attn-it ♡40 | ♡31
mindlab-research/Macaron-V1-Venti ♡35
AliveAi/Krea-2-Edit-Outfit-Transfer ♡34
Glint-Research/Glint-2 ♡28
amd/Instella-MoE-16B-A3B-Think ♡28
neuphonic/neutts-2e ♡26
Reza2kn/Bina-0.1 ♡24
joeygambino/joyai-echo-ltx23-echoVid-ltxAud-surgical ♡23
Trelis/tiron ♡21
BananaMind/BananaMind-2-Medium ♡14
ai9stars/G9v3-3B ♡13
GitHub
args: refactor mlock/mmap/directio into load-mode by taronaeo · Pull Request #20834 · ggml-org/llama.cpp
Ref: #20211 (comment)
Obsoletes: #20461
This PR overhauls the three separate loading modes (mlock, mmap, and direct-io) into one single -lm/--load-mode option to simplify the logic. While working o...
Obsoletes: #20461
This PR overhauls the three separate loading modes (mlock, mmap, and direct-io) into one single -lm/--load-mode option to simplify the logic. While working o...
🆕 [HF Models] microsoft - Mage-ViT
https://huggingface.co/microsoft/Mage-ViT
🆕 [HF Models] microsoft - Mage-VL
https://huggingface.co/microsoft/Mage-VL
https://huggingface.co/microsoft/Mage-ViT
🆕 [HF Models] microsoft - Mage-VL
https://huggingface.co/microsoft/Mage-VL
huggingface.co
microsoft/Mage-ViT · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Anthropic Research - Project Pilot: Can AI control a drone?
https://www.anthropic.com/research/project-pilot
https://www.anthropic.com/research/project-pilot
Anthropic
Project Pilot: Can AI control a drone?
We worked with Andon Labs on Drone-Bench, a new benchmark testing whether AI models can autonomously fly a drone to locate and follow a person.
🔄 [GitHub Releases] open-webui/open-webui - v0.11.0
https://github.com/open-webui/open-webui/releases/tag/v0.11.0
https://github.com/open-webui/open-webui/releases/tag/v0.11.0
GitHub
Release v0.11.0 · open-webui/open-webui
Added
🎨 Redesigned interface. Open WebUI has been visually rebuilt from the ground up. All aspects of the User Interface, from the chat view to the admin panel. Now with a narrower conversation co...
🎨 Redesigned interface. Open WebUI has been visually rebuilt from the ground up. All aspects of the User Interface, from the chat view to the admin panel. Now with a narrower conversation co...
📰 HuggingFace - NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics
https://huggingface.co/blog/nvidia/cosmos-h-dreams
https://huggingface.co/blog/nvidia/cosmos-h-dreams
huggingface.co
NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics
A Blog post by NVIDIA on Hugging Face