📰 Claude Blog - Building effective human-agent teams
https://claude.com/blog/building-effective-human-agent-teams
📰 Claude Blog - Agent identity in Claude Tag: a new access model for autonomous, team-wide AI
https://claude.com/blog/agent-identity-access-model
📰 Claude Blog - The full Claude Desktop experience on AWS, Google Cloud, and Microsoft Foundry
https://claude.com/blog/the-full-claude-desktop-experience-on-aws-google-cloud-and-microsoft-foundry
https://claude.com/blog/building-effective-human-agent-teams
📰 Claude Blog - Agent identity in Claude Tag: a new access model for autonomous, team-wide AI
https://claude.com/blog/agent-identity-access-model
📰 Claude Blog - The full Claude Desktop experience on AWS, Google Cloud, and Microsoft Foundry
https://claude.com/blog/the-full-claude-desktop-experience-on-aws-google-cloud-and-microsoft-foundry
Claude
Lessons from Anthropic on building effective human-agent teams | Claude by Anthropic
The way we work with AI is evolving from a single-player to a multiplayer experience, where humans and agents work together as a team to achieve shared goals. The Anthropic team shares examples of this new way of working in action.
📰 Qwen Research - Qwen-AgentWorld: Language World Models for General Agents
https://qwen.ai/blog?id=qwen-agentworld
https://qwen.ai/blog?id=qwen-agentworld
qwen.ai
Qwen Studio
Qwen Studio offers comprehensive functionality spanning chatbot, image and video understanding, image generation, document processing, web search integration, tool utilization, and artifacts.
🔄 [GitHub Releases] flashinfer-ai/flashinfer - Release v0.6.13
https://github.com/flashinfer-ai/flashinfer/releases/tag/v0.6.13
https://github.com/flashinfer-ai/flashinfer/releases/tag/v0.6.13
GitHub
Release Release v0.6.13 · flashinfer-ai/flashinfer
What's Changed
Run high-likelihood OOM culprits separately, record memory usage and test duration for analysis by @dierksen in #2961
fix(autotuner): differentiate file cache entries by runner ...
Run high-likelihood OOM culprits separately, record memory usage and test duration for analysis by @dierksen in #2961
fix(autotuner): differentiate file cache entries by runner ...
🆕 [HF Models] LiquidAI - LFM2.5-230M
https://huggingface.co/LiquidAI/LFM2.5-230M
🔓 [HF Models] LiquidAI - LFM2.5-230M-GGUF
https://huggingface.co/LiquidAI/LFM2.5-230M-GGUF
🔓 [HF Models] LiquidAI - LFM2.5-230M-Base
https://huggingface.co/LiquidAI/LFM2.5-230M-Base
https://huggingface.co/LiquidAI/LFM2.5-230M
🔓 [HF Models] LiquidAI - LFM2.5-230M-GGUF
https://huggingface.co/LiquidAI/LFM2.5-230M-GGUF
🔓 [HF Models] LiquidAI - LFM2.5-230M-Base
https://huggingface.co/LiquidAI/LFM2.5-230M-Base
huggingface.co
LiquidAI/LFM2.5-230M · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 PyTorch - TokenSpeed-Kernel: Portable APIs and High-Performance Kernels for Multi-Silicon LLM Inference
TL;DR The TokenSpeed-kernel is a standalone, open-source subsystem designed to solve backend complexity in LLM inference. It introduces a clean, layered API and registry system that decouples the high-level runtime...
https://pytorch.org/blog/lightseek-tokenspeed-kernel/
TL;DR The TokenSpeed-kernel is a standalone, open-source subsystem designed to solve backend complexity in LLM inference. It introduces a clean, layered API and registry system that decouples the high-level runtime...
https://pytorch.org/blog/lightseek-tokenspeed-kernel/
📰 HuggingFace - Which tokens does a hybrid model predict better?
https://huggingface.co/blog/allenai/hybrid-token-prediction
📰 HuggingFace - Run a vLLM Server on HF Jobs in One Command
https://huggingface.co/blog/vllm-jobs
https://huggingface.co/blog/allenai/hybrid-token-prediction
📰 HuggingFace - Run a vLLM Server on HF Jobs in One Command
https://huggingface.co/blog/vllm-jobs
huggingface.co
Which tokens does a hybrid model predict better?
A Blog post by Ai2 on Hugging Face
📰 NVIDIA - Scaling AI Inference Across Multiple GPUs Using NVIDIA TensorRT with Multi-Device Inference Support
Generative AI workloads are rapidly outgrowing the memory and compute budget of single GPUs. For inference developers building media generation pipelines…
https://developer.nvidia.com/blog/scaling-ai-inference-across-multiple-gpus-using-nvidia-tensorrt-with-multi-device-inference-support/
Generative AI workloads are rapidly outgrowing the memory and compute budget of single GPUs. For inference developers building media generation pipelines…
https://developer.nvidia.com/blog/scaling-ai-inference-across-multiple-gpus-using-nvidia-tensorrt-with-multi-device-inference-support/
NVIDIA Technical Blog
Scaling AI Inference Across Multiple GPUs Using NVIDIA TensorRT with Multi-Device Inference Support
Generative AI workloads are rapidly outgrowing the memory and compute budget of single GPUs. For inference developers building media generation pipelines, the challenge is scaling across multiple…
🆕 [HF Models] Qwen - Qwen3-ForcedAligner-0.6B-hf
https://huggingface.co/Qwen/Qwen3-ForcedAligner-0.6B-hf
🆕 [HF Models] Qwen - Qwen3-ASR-0.6B-hf
https://huggingface.co/Qwen/Qwen3-ASR-0.6B-hf
🆕 [HF Models] Qwen - Qwen3-ASR-1.7B-hf
https://huggingface.co/Qwen/Qwen3-ASR-1.7B-hf
https://huggingface.co/Qwen/Qwen3-ForcedAligner-0.6B-hf
🆕 [HF Models] Qwen - Qwen3-ASR-0.6B-hf
https://huggingface.co/Qwen/Qwen3-ASR-0.6B-hf
🆕 [HF Models] Qwen - Qwen3-ASR-1.7B-hf
https://huggingface.co/Qwen/Qwen3-ASR-1.7B-hf
huggingface.co
Qwen/Qwen3-ForcedAligner-0.6B-hf · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 LMSys - Improving DeepEP MoE Load Balance in SGLang with Waterfill and LPLB
https://lmsys.org/blog/2026-06-26-waterfill-lplb
https://lmsys.org/blog/2026-06-26-waterfill-lplb
www.lmsys.org
Improving DeepEP MoE Load Balance in SGLang with Waterfill and LPLB
Mixture-of-Experts (MoE) models rely on Expert Parallelism (EP) to scale inference across multiple GPUs. In SGLang, DeepEP and EPLB provide high-performance serving under EP, but the workload seen by ...
📰 Anthropic Research - Anthropic Economic Index report: Cadences
https://www.anthropic.com/research/economic-index-june-2026-report
https://www.anthropic.com/research/economic-index-june-2026-report
Anthropic
Anthropic Economic Index report: Cadences
In the latest Anthropic Economic Index report, we look at when people come to Claude, what they produce with it, and how they perceive AI’s impact on their work.
📰 NVIDIA - Deploy a Production-Ready NVIDIA AI-Q Blueprint on Oracle Cloud Infrastructure
AI agents have changed a lot in the last two years. The first could only answer one question at a time. Then came multi-turn chat, where the model could keep…
https://developer.nvidia.com/blog/deploy-a-production-ready-nvidia-ai-q-blueprint-on-oracle-cloud-infrastructure/
📰 NVIDIA - Creating the NVIDIA Nemotron 3 Ultra NVFP4 Checkpoint with NVIDIA Model Optimizer
As context windows grow longer, moving large model weights efficiently becomes critical to performance. A common way to address this is quantization…
https://developer.nvidia.com/blog/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer/
AI agents have changed a lot in the last two years. The first could only answer one question at a time. Then came multi-turn chat, where the model could keep…
https://developer.nvidia.com/blog/deploy-a-production-ready-nvidia-ai-q-blueprint-on-oracle-cloud-infrastructure/
📰 NVIDIA - Creating the NVIDIA Nemotron 3 Ultra NVFP4 Checkpoint with NVIDIA Model Optimizer
As context windows grow longer, moving large model weights efficiently becomes critical to performance. A common way to address this is quantization…
https://developer.nvidia.com/blog/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer/
NVIDIA Technical Blog
Deploy a Production-Ready NVIDIA AI-Q Blueprint on Oracle Cloud Infrastructure
AI agents have changed a lot in the last two years. The first could only answer one question at a time. Then came multi-turn chat, where the model could keep some context across a session. Today…
🔄 [GitHub Releases] sgl-project/sglang - v0.5.14
https://github.com/sgl-project/sglang/releases/tag/v0.5.14
https://github.com/sgl-project/sglang/releases/tag/v0.5.14
GitHub
Release v0.5.14 · sgl-project/sglang
Highlights
New Model Support: GLM-5.2, LiquidAI LFM2.5, Kimi-K2.7-Code, Poolside Laguna-M.1, DiffusionGemma, Zyphra ZAYA1, MiMo-V2-ASR
DeepSeek-V4 on GB300 since Day 0: 5x higher throughput at the ...
New Model Support: GLM-5.2, LiquidAI LFM2.5, Kimi-K2.7-Code, Poolside Laguna-M.1, DiffusionGemma, Zyphra ZAYA1, MiMo-V2-ASR
DeepSeek-V4 on GB300 since Day 0: 5x higher throughput at the ...
🆕 [HF Models] deepseek-ai - DeepSeek-V4-Pro-DSpark
https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark
🆕 [HF Models] deepseek-ai - DeepSeek-V4-Flash-DSpark
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-DSpark
https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark
🆕 [HF Models] deepseek-ai - DeepSeek-V4-Flash-DSpark
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-DSpark
🗓️ Weekly GitHub Activity
🦙 llama.cpp
└ Release: b9743 → b9828
└ 85 commits
- Added support for Step 3.5/3.7 flash MTP3 speculative decoding #24340,
Granite Speech Plus #24818,
LFM2.5-ColBERT-350M/Embedding-350M #24913,
Eagle3 Qwen3 draft models #24977,
Unlimited-OCR #24969
- Added SSE Replay Buffer to server and UI, allowing text generation to survive HTTP disconnects and resume seamlessly #23226
- Introduced real-time model loading progress tracking via SSE in both server and UI #24828 #24878
- Configured server to create checkpoints before every user message to improve session recovery #24176
- Redesigned the WebUI with a new logo, navigation cleanup, and significant mobile layout improvements #24897
- Enabled dual-GPU tensor parallelism on the SYCL backend via split-mode tensor #24152
- Overhauled Hexagon matrix multiplication kernels with tiled layouts, HVX/HMX microkernels, and graph caching #24954
- Upgraded OpenCL Flash Attention kernels for F16, F32, Q4_0, and Q8_0 #25069
- Moved server model downloading to a dedicated child process #24834
- Added CUDA fast path for strided 2D copies using cudaMemcpy2DAsync #25057
- Reduced synchronization overhead between CPU and CUDA async copies during split compute #20793
- Added 3D convolution support to Vulkan #24612
- Fixed CUDA integer overflows and transposed copy failures #24706 #25000
- Fixed incorrect vector dot computations on SVE-enabled ARM CPUs #24699
🔗 All changes | Latest release
🎨 stable-diffusion.cpp
└ Release: master-709-92a3b73 → master-721-8caa3f9
└ 12 commits
- Added support for Boogu image generation #1688
- Added support for Krea2 models #1705
- Introduced guidance_schedule support for generation control #1684
- Added logit-normal scheduler #1669
- Added --eager-load flag to pre-load parameters during model initialization #1687
- Added --prompt-file and --negative-prompt-file flags for file-based inputs #1693
- Fixed memory mapping by avoiding writable mmap for read-only weights #1698
🔗 All changes | Latest release
🤗 Fresh models trending on HuggingFace:
empero-ai/Qwythos-9B-Claude-Mythos-5-1M ♡488
krea/Krea-2-Turbo ♡310
krea/Krea-2-Raw ♡214
deepreinforce-ai/Ornith-1.0-9B ♡167
deepreinforce-ai/Ornith-1.0-35B ♡161
deepreinforce-ai/Ornith-1.0-397B ♡121
Chunjiang-Intelligence/DeepSeek-v4-Fable ♡112
hustvl/Moebius ♡51
AutoArk-AI/ARK-ASR-3B ♡37
paom/texture2albedo-v2 ♡32
SupraLabs/Supra-A2A-Nano-Exp ♡30
Gryphe/Gemma-4-26B-A4B-StyleTune-V2 ♡24
0xSero/GLM-5.2-504B ♡19
g-astruc/UniverSat ♡18
allenai/tmax-27b ♡18
ValiantLabs/Qwen3.6-27B-Esper4 ♡14
wikeeyang/Flux2-Klein-9B-True-V3 ♡14
vrgamedevgirl84/Krea2_Enhancer ♡12
🦙 llama.cpp
└ Release: b9743 → b9828
└ 85 commits
- Added support for Step 3.5/3.7 flash MTP3 speculative decoding #24340,
Granite Speech Plus #24818,
LFM2.5-ColBERT-350M/Embedding-350M #24913,
Eagle3 Qwen3 draft models #24977,
Unlimited-OCR #24969
- Added SSE Replay Buffer to server and UI, allowing text generation to survive HTTP disconnects and resume seamlessly #23226
- Introduced real-time model loading progress tracking via SSE in both server and UI #24828 #24878
- Configured server to create checkpoints before every user message to improve session recovery #24176
- Redesigned the WebUI with a new logo, navigation cleanup, and significant mobile layout improvements #24897
- Enabled dual-GPU tensor parallelism on the SYCL backend via split-mode tensor #24152
- Overhauled Hexagon matrix multiplication kernels with tiled layouts, HVX/HMX microkernels, and graph caching #24954
- Upgraded OpenCL Flash Attention kernels for F16, F32, Q4_0, and Q8_0 #25069
- Moved server model downloading to a dedicated child process #24834
- Added CUDA fast path for strided 2D copies using cudaMemcpy2DAsync #25057
- Reduced synchronization overhead between CPU and CUDA async copies during split compute #20793
- Added 3D convolution support to Vulkan #24612
- Fixed CUDA integer overflows and transposed copy failures #24706 #25000
- Fixed incorrect vector dot computations on SVE-enabled ARM CPUs #24699
🔗 All changes | Latest release
🎨 stable-diffusion.cpp
└ Release: master-709-92a3b73 → master-721-8caa3f9
└ 12 commits
- Added support for Boogu image generation #1688
- Added support for Krea2 models #1705
- Introduced guidance_schedule support for generation control #1684
- Added logit-normal scheduler #1669
- Added --eager-load flag to pre-load parameters during model initialization #1687
- Added --prompt-file and --negative-prompt-file flags for file-based inputs #1693
- Fixed memory mapping by avoiding writable mmap for read-only weights #1698
🔗 All changes | Latest release
🤗 Fresh models trending on HuggingFace:
empero-ai/Qwythos-9B-Claude-Mythos-5-1M ♡488
krea/Krea-2-Turbo ♡310
krea/Krea-2-Raw ♡214
deepreinforce-ai/Ornith-1.0-9B ♡167
deepreinforce-ai/Ornith-1.0-35B ♡161
deepreinforce-ai/Ornith-1.0-397B ♡121
Chunjiang-Intelligence/DeepSeek-v4-Fable ♡112
hustvl/Moebius ♡51
AutoArk-AI/ARK-ASR-3B ♡37
paom/texture2albedo-v2 ♡32
SupraLabs/Supra-A2A-Nano-Exp ♡30
Gryphe/Gemma-4-26B-A4B-StyleTune-V2 ♡24
0xSero/GLM-5.2-504B ♡19
g-astruc/UniverSat ♡18
allenai/tmax-27b ♡18
ValiantLabs/Qwen3.6-27B-Esper4 ♡14
wikeeyang/Flux2-Klein-9B-True-V3 ♡14
vrgamedevgirl84/Krea2_Enhancer ♡12
GitHub
Support Step3.5/3.7 flash mtp3 by forforever73 · Pull Request #24340 · ggml-org/llama.cpp
Overview
follow-up to #23274.(cc @pwilkin )
📜 Full data-flow trace — couldn't think of a good way to draw this, so I wrote it all down instead. It's long, but every byte is load-be...
follow-up to #23274.(cc @pwilkin )
📜 Full data-flow trace — couldn't think of a good way to draw this, so I wrote it all down instead. It's long, but every byte is load-be...
❤2