🆕 [HF Models] swiss-ai - Apertus-v1.5-70B
https://huggingface.co/swiss-ai/Apertus-v1.5-70B
🆕 [HF Models] swiss-ai - Apertus-v1.5-8B
https://huggingface.co/swiss-ai/Apertus-v1.5-8B
https://huggingface.co/swiss-ai/Apertus-v1.5-70B
🆕 [HF Models] swiss-ai - Apertus-v1.5-8B
https://huggingface.co/swiss-ai/Apertus-v1.5-8B
huggingface.co
swiss-ai/Apertus-v1.5-70B · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Google AI Blog - Run Ray on TPU, Part 2: Ray AI libraries
This second installment explores how Ray’s higher-level libraries—Serve, Data, and Train—abstract the complexities of running AI workloads on Google's TPU slices. Ray Serve uses a simple topology configuration to correctly gang-schedule large multi-host models, while Ray Data eliminates data-loading bottlenecks by feeding accelerators directly with native JAX batches. Finally, JaxTrainer streamlines distributed training across TPUs by automatically handling cross-slice coordination, checkpointing, and fault tolerance.
https://developers.googleblog.com/en/run-ray-on-tpu-part-2-ray-ai-libraries/
This second installment explores how Ray’s higher-level libraries—Serve, Data, and Train—abstract the complexities of running AI workloads on Google's TPU slices. Ray Serve uses a simple topology configuration to correctly gang-schedule large multi-host models, while Ray Data eliminates data-loading bottlenecks by feeding accelerators directly with native JAX batches. Finally, JaxTrainer streamlines distributed training across TPUs by automatically handling cross-slice coordination, checkpointing, and fault tolerance.
https://developers.googleblog.com/en/run-ray-on-tpu-part-2-ray-ai-libraries/
Googleblog
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Learn how to scale AI workloads on TPU slices using Ray Serve for LLM deployment, Ray Data for fast JAX pipelines, and JaxTrainer for distributed training.
🔓 xAI - Bringing Grok 4.5 to iOS, Android, Web, and X
https://x.ai//news/grok-4-5-everywhere
🔓 xAI - Workflows in Grok Build
https://x.ai//news/workflows
🔓 xAI - Grok in Google Workspace
https://x.ai//news/introducing-google-workspace-addon
https://x.ai//news/grok-4-5-everywhere
🔓 xAI - Workflows in Grok Build
https://x.ai//news/workflows
🔓 xAI - Grok in Google Workspace
https://x.ai//news/introducing-google-workspace-addon
x.ai
Bringing Grok 4.5 to iOS, Android, Web, and X
Grok 4.5, our most intelligent model yet, is now on grok.com, X, iOS, and Android.
📰 NVIDIA - ModelExpress: Distributing Model Artifacts at the Speed of Light
Every byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly. To make things even worse…
https://developer.nvidia.com/blog/modelexpress-distributing-model-artifacts-at-the-speed-of-light/
Every byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly. To make things even worse…
https://developer.nvidia.com/blog/modelexpress-distributing-model-artifacts-at-the-speed-of-light/
NVIDIA Technical Blog
ModelExpress: Distributing Model Artifacts at the Speed of Light
Every byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly. To make things even worse, moving these model weights around the cluster…
🔄 [GitHub Releases] sgl-project/sglang - v0.5.16
https://github.com/sgl-project/sglang/releases/tag/v0.5.16
https://github.com/sgl-project/sglang/releases/tag/v0.5.16
GitHub
Release v0.5.16 · sgl-project/sglang
Highlights
574 PRs from 169 contributors.
DSpark: confidence-driven speculative decoding: A new speculative algorithm. It drafts semi-autoregressively in blocks, then sizes each verify window from ...
574 PRs from 169 contributors.
DSpark: confidence-driven speculative decoding: A new speculative algorithm. It drafts semi-autoregressively in blocks, then sizes each verify window from ...
🔄 [GitHub Releases] vllm-project/vllm - v0.26.0
https://github.com/vllm-project/vllm/releases/tag/v0.26.0
https://github.com/vllm-project/vllm/releases/tag/v0.26.0
GitHub
Release v0.26.0 · vllm-project/vllm
vLLM v0.26.0 Release Notes
Highlights
This release features 411 commits from 212 contributors (61 new)!
New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA g...
Highlights
This release features 411 commits from 212 contributors (61 new)!
New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA g...
🔄 [GitHub Releases] turboderp-org/exllamav3 - 1.2.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.2.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.2.0
GitHub
Release 1.2.0 · turboderp-org/exllamav3
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs - Release 1.2.0 · turboderp-org/exllamav3
📰 Claude Blog - How the product designer who built Claude Design uses it to explore ideas before building them
https://claude.com/blog/how-the-product-designer-who-built-claude-design-uses-it-to-explore-ideas-before-building-them
📰 Claude Blog - The new rules of context engineering for Claude 5 generation models
https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models
📰 Claude Blog - Claude models explained: choosing the best model for your use case
https://claude.com/blog/claude-models-explained-choosing-the-best-model-for-your-use-case
📰 Claude Blog - Four role-based certifications for the people who put Claude to work for customers
https://claude.com/blog/four-role-based-claude-certifications
📰 Claude Blog - Think through hard problems in voice mode
https://claude.com/blog/think-through-hard-problems-in-voice-mode
📰 Claude Blog - Building verification loops in Claude Code with skills
https://claude.com/blog/building-verification-loops-in-claude-code-with-skills
📰 Claude Blog - How Outtake built a cyber investigator on Claude
https://claude.com/blog/how-outtake-built-a-cyber-investigator-on-claude
https://claude.com/blog/how-the-product-designer-who-built-claude-design-uses-it-to-explore-ideas-before-building-them
📰 Claude Blog - The new rules of context engineering for Claude 5 generation models
https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models
📰 Claude Blog - Claude models explained: choosing the best model for your use case
https://claude.com/blog/claude-models-explained-choosing-the-best-model-for-your-use-case
📰 Claude Blog - Four role-based certifications for the people who put Claude to work for customers
https://claude.com/blog/four-role-based-claude-certifications
📰 Claude Blog - Think through hard problems in voice mode
https://claude.com/blog/think-through-hard-problems-in-voice-mode
📰 Claude Blog - Building verification loops in Claude Code with skills
https://claude.com/blog/building-verification-loops-in-claude-code-with-skills
📰 Claude Blog - How Outtake built a cyber investigator on Claude
https://claude.com/blog/how-outtake-built-a-cyber-investigator-on-claude
Claude
How the product designer who built Claude Design uses it | Claude by Anthropic
Nate Parrott, a product designer at Anthropic, shares how he uses Claude Design (in beta) to explore, iterate on, and share visual ideas early, from product prototypes to slide decks and animations.
🗓️ Weekly GitHub Activity
🦙 llama.cpp
└ Release: b10068 → b10107
└ 39 commits
- Consolidated memory mapping and locking CLI options into a unified --load-mode argument #20834
- Added support for Laguna XS.2 and M.1 models #25165
- Improved DeepSeek-V4 support with softplus CUDA kernels, APE tensor fixes, and chat template updates #25896, #25945, #25414
- Fixed coordinate scaling in Qwen3-VL via align_corners interpolation #25781 and HunyuanVL XD-RoPE conversion #25514
- Enabled automatic speculative decoding type inference and sidecar model resolution from draft repositories #25955, #25989
- Upgraded CUDA backend with device-side GET_ROWS support for k-quants, i-quants, and mxfp4 #25962, alongside vectorized same-type copy optimization #25929
- Refactored Vulkan queue submission to bypass host-side locking via per-instance mutexes and unique handles #23570
- Added depthwise 2D convolution (CONV_2D_DW) kernel for WebGPU #25847
- Updated WebUI with bulk conversation management #25815, symbolic math support in the JS sandbox using Nerdamer #25948, and a default reasoning mode selector #25846
🔗 All changes | Latest release
🎨 stable-diffusion.cpp
└ Release: master-782-b290693 → master-795-87a0177
└ 13 commits
- Added support for Hunyuan Video 1.5 #1795
- Added IP-Adapter support for SD 1.5 and SDXL #1803 with CFG conditioning fixes #1815
- Added Mage-Flow support #1808
🔗 All changes | Latest release
🤗 Fresh models trending on HuggingFace:
Nanbeige/Nanbeige4.2-3B | Nanbeige4.2-3B-Base ♡406 | ♡40
Kwaipilot/KAT-Coder-V2.5-Dev ♡165
fdtn-ai/antares-1b | antares-350m ♡163 | ♡50
badtheorylabs/BTL-3 | BTL-3-Compact ♡60 | ♡26
ProCreations/grug-27b ♡57
PaddlePaddle/HPD-Parsing ♡57
FINAL-Bench/Aether-7B-5Attn | Aether-7B-5Attn-it ♡40 | ♡31
mindlab-research/Macaron-V1-Venti ♡35
AliveAi/Krea-2-Edit-Outfit-Transfer ♡34
Glint-Research/Glint-2 ♡28
amd/Instella-MoE-16B-A3B-Think ♡28
neuphonic/neutts-2e ♡26
Reza2kn/Bina-0.1 ♡24
joeygambino/joyai-echo-ltx23-echoVid-ltxAud-surgical ♡23
Trelis/tiron ♡21
BananaMind/BananaMind-2-Medium ♡14
ai9stars/G9v3-3B ♡13
🦙 llama.cpp
└ Release: b10068 → b10107
└ 39 commits
- Consolidated memory mapping and locking CLI options into a unified --load-mode argument #20834
- Added support for Laguna XS.2 and M.1 models #25165
- Improved DeepSeek-V4 support with softplus CUDA kernels, APE tensor fixes, and chat template updates #25896, #25945, #25414
- Fixed coordinate scaling in Qwen3-VL via align_corners interpolation #25781 and HunyuanVL XD-RoPE conversion #25514
- Enabled automatic speculative decoding type inference and sidecar model resolution from draft repositories #25955, #25989
- Upgraded CUDA backend with device-side GET_ROWS support for k-quants, i-quants, and mxfp4 #25962, alongside vectorized same-type copy optimization #25929
- Refactored Vulkan queue submission to bypass host-side locking via per-instance mutexes and unique handles #23570
- Added depthwise 2D convolution (CONV_2D_DW) kernel for WebGPU #25847
- Updated WebUI with bulk conversation management #25815, symbolic math support in the JS sandbox using Nerdamer #25948, and a default reasoning mode selector #25846
🔗 All changes | Latest release
🎨 stable-diffusion.cpp
└ Release: master-782-b290693 → master-795-87a0177
└ 13 commits
- Added support for Hunyuan Video 1.5 #1795
- Added IP-Adapter support for SD 1.5 and SDXL #1803 with CFG conditioning fixes #1815
- Added Mage-Flow support #1808
🔗 All changes | Latest release
🤗 Fresh models trending on HuggingFace:
Nanbeige/Nanbeige4.2-3B | Nanbeige4.2-3B-Base ♡406 | ♡40
Kwaipilot/KAT-Coder-V2.5-Dev ♡165
fdtn-ai/antares-1b | antares-350m ♡163 | ♡50
badtheorylabs/BTL-3 | BTL-3-Compact ♡60 | ♡26
ProCreations/grug-27b ♡57
PaddlePaddle/HPD-Parsing ♡57
FINAL-Bench/Aether-7B-5Attn | Aether-7B-5Attn-it ♡40 | ♡31
mindlab-research/Macaron-V1-Venti ♡35
AliveAi/Krea-2-Edit-Outfit-Transfer ♡34
Glint-Research/Glint-2 ♡28
amd/Instella-MoE-16B-A3B-Think ♡28
neuphonic/neutts-2e ♡26
Reza2kn/Bina-0.1 ♡24
joeygambino/joyai-echo-ltx23-echoVid-ltxAud-surgical ♡23
Trelis/tiron ♡21
BananaMind/BananaMind-2-Medium ♡14
ai9stars/G9v3-3B ♡13
GitHub
args: refactor mlock/mmap/directio into load-mode by taronaeo · Pull Request #20834 · ggml-org/llama.cpp
Ref: #20211 (comment)
Obsoletes: #20461
This PR overhauls the three separate loading modes (mlock, mmap, and direct-io) into one single -lm/--load-mode option to simplify the logic. While working o...
Obsoletes: #20461
This PR overhauls the three separate loading modes (mlock, mmap, and direct-io) into one single -lm/--load-mode option to simplify the logic. While working o...
🆕 [HF Models] microsoft - Mage-ViT
https://huggingface.co/microsoft/Mage-ViT
🆕 [HF Models] microsoft - Mage-VL
https://huggingface.co/microsoft/Mage-VL
https://huggingface.co/microsoft/Mage-ViT
🆕 [HF Models] microsoft - Mage-VL
https://huggingface.co/microsoft/Mage-VL
huggingface.co
microsoft/Mage-ViT · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Anthropic Research - Project Pilot: Can AI control a drone?
https://www.anthropic.com/research/project-pilot
https://www.anthropic.com/research/project-pilot
Anthropic
Project Pilot: Can AI control a drone?
We worked with Andon Labs on Drone-Bench, a new benchmark testing whether AI models can autonomously fly a drone to locate and follow a person.
🔄 [GitHub Releases] open-webui/open-webui - v0.11.0
https://github.com/open-webui/open-webui/releases/tag/v0.11.0
https://github.com/open-webui/open-webui/releases/tag/v0.11.0
GitHub
Release v0.11.0 · open-webui/open-webui
Added
🎨 Redesigned interface. Open WebUI has been visually rebuilt from the ground up. All aspects of the User Interface, from the chat view to the admin panel. Now with a narrower conversation co...
🎨 Redesigned interface. Open WebUI has been visually rebuilt from the ground up. All aspects of the User Interface, from the chat view to the admin panel. Now with a narrower conversation co...
📰 HuggingFace - NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics
https://huggingface.co/blog/nvidia/cosmos-h-dreams
https://huggingface.co/blog/nvidia/cosmos-h-dreams
huggingface.co
NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics
A Blog post by NVIDIA on Hugging Face
📰 LMSys - SGLang and Miles Add Day-0 Support for Kimi K3
https://lmsys.org/blog/2026-07-27-kimi-k3-day0-support
https://lmsys.org/blog/2026-07-27-kimi-k3-day0-support
www.lmsys.org
SGLang and Miles Add Day-0 Support for Kimi K3
We are excited to announce Day-0 support for Kimi K3
in SGLang and Miles. K3 is the first open-source model in the 3-trillion-parameter class,
and its hybrid architecture departs from convention in al...
in SGLang and Miles. K3 is the first open-source model in the 3-trillion-parameter class,
and its hybrid architecture departs from convention in al...
📰 NVIDIA - NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning
NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they…
https://developer.nvidia.com/blog/nvidia-ising-enables-fully-automated-quantum-computer-calibration-with-enhanced-in-context-learning/
📰 NVIDIA - Six Agent Harness Capabilities for Higher Model Performance
Building a great AI agent isn’t just about choosing the right models. The harness is the architecture surrounding the model. How it renders context…
https://developer.nvidia.com/blog/six-agent-harness-capabilities-for-higher-model-performance/
📰 NVIDIA - NVIDIA Nemotron 3 Ultra Leads Open Models on Accuracy and Efficiency in Agentic RTL Coding
Modern chip design is increasingly limited by engineering time. Register transfer level (RTL) development and verification require specialized hardware…
https://developer.nvidia.com/blog/nvidia-nemotron-3-ultra-leads-open-models-on-accuracy-and-efficiency-in-agentic-rtl-coding/
NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they…
https://developer.nvidia.com/blog/nvidia-ising-enables-fully-automated-quantum-computer-calibration-with-enhanced-in-context-learning/
📰 NVIDIA - Six Agent Harness Capabilities for Higher Model Performance
Building a great AI agent isn’t just about choosing the right models. The harness is the architecture surrounding the model. How it renders context…
https://developer.nvidia.com/blog/six-agent-harness-capabilities-for-higher-model-performance/
📰 NVIDIA - NVIDIA Nemotron 3 Ultra Leads Open Models on Accuracy and Efficiency in Agentic RTL Coding
Modern chip design is increasingly limited by engineering time. Register transfer level (RTL) development and verification require specialized hardware…
https://developer.nvidia.com/blog/nvidia-nemotron-3-ultra-leads-open-models-on-accuracy-and-efficiency-in-agentic-rtl-coding/
NVIDIA Technical Blog
NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning
NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they should be tuned to continue operating.
📰 Google Labs - Google and KDDI are ready to back Japanese startups.
https://blog.google/innovation-and-ai/models-and-research/google-labs/ai-startup-support-program-japan/
https://blog.google/innovation-and-ai/models-and-research/google-labs/ai-startup-support-program-japan/
Google
Google and KDDI are ready to back Japanese startups.
We’re launching the AI Startup Support Program to accelerate innovative AI-native Japanese startups.
🆕 [HF Models] Motif-Technologies - Motif-Audio
https://huggingface.co/Motif-Technologies/Motif-Audio
🆕 [HF Models] Motif-Technologies - Motif-Vision-Encoder
https://huggingface.co/Motif-Technologies/Motif-Vision-Encoder
https://huggingface.co/Motif-Technologies/Motif-Audio
🆕 [HF Models] Motif-Technologies - Motif-Vision-Encoder
https://huggingface.co/Motif-Technologies/Motif-Vision-Encoder
huggingface.co
Motif-Technologies/Motif-Audio · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.