📰 LMSys - DSpark in SGLang: Speculative Decoding with Confidence-Driven, Variable-Length Verification
https://lmsys.org/blog/2026-07-06-dspark-sglang
https://lmsys.org/blog/2026-07-06-dspark-sglang
www.lmsys.org
DSpark in SGLang: Speculative Decoding with Confidence-Driven, Variable-Length Verification
Speculative decoding trades extra compute for fewer decode steps, and the trade
sours as load grows: at batch size B with K speculative tokens the target
verifies B K tokens every step, and past a poi...
sours as load grows: at batch size B with K speculative tokens the target
verifies B K tokens every step, and past a poi...
📰 Anthropic Research - A global workspace in language models
https://www.anthropic.com/research/global-workspace
https://www.anthropic.com/research/global-workspace
Anthropic
A global workspace in language models
Interpretability research on Claude's internal thoughts.
📰 Anthropic - Government of Alberta uses Claude to find and fix cybersecurity vulnerabilities across government systems
https://www.anthropic.com/news/alberta-government-claude-cybersecurity
https://www.anthropic.com/news/alberta-government-claude-cybersecurity
Anthropic
Government of Alberta uses Claude to find and fix cybersecurity vulnerabilities across government systems
The Government of Alberta has been using Claude Code with both Opus and Sonnet models to review its systems, find vulnerabilities, and fix them.
📰 PyTorch - Bringing PyTorch Monarch to AMD GPUs: Single-Controller Distributed Training on ROCm
Training state-of-the-art large language models (LLMs) with billions of parameters requires distributed training across hundreds or thousands of GPUs. At this scale, hardware failures are not exceptional events—they are expected....
https://pytorch.org/blog/bringing-pytorch-monarch-to-amd-gpus-single-controller-distributed-training-on-rocm/
Training state-of-the-art large language models (LLMs) with billions of parameters requires distributed training across hundreds or thousands of GPUs. At this scale, hardware failures are not exceptional events—they are expected....
https://pytorch.org/blog/bringing-pytorch-monarch-to-amd-gpus-single-controller-distributed-training-on-rocm/
🆕 [HF Models] nvidia - Nemotron-Labs-Audex-2B
https://huggingface.co/nvidia/Nemotron-Labs-Audex-2B
🆕 [HF Models] nvidia - Nemotron-Labs-Audex-30B-A3B
https://huggingface.co/nvidia/Nemotron-Labs-Audex-30B-A3B
https://huggingface.co/nvidia/Nemotron-Labs-Audex-2B
🆕 [HF Models] nvidia - Nemotron-Labs-Audex-30B-A3B
https://huggingface.co/nvidia/Nemotron-Labs-Audex-30B-A3B
huggingface.co
nvidia/Nemotron-Labs-Audex-30B-A3B · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
🔓 [HF Models] nvidia - NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16
https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16
https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16
huggingface.co
nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
🆕 [HF Models] CohereLabs - cohere-transcribe-arabic-07-2026
https://huggingface.co/CohereLabs/cohere-transcribe-arabic-07-2026
https://huggingface.co/CohereLabs/cohere-transcribe-arabic-07-2026
huggingface.co
CohereLabs/cohere-transcribe-arabic-07-2026 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 HuggingFace - Hugging Face Models on Foundry Managed Compute
https://huggingface.co/blog/microsoft/foundry-managed-compute
https://huggingface.co/blog/microsoft/foundry-managed-compute
huggingface.co
Hugging Face Models on Foundry Managed Compute
A Blog post by Microsoft on Hugging Face
🔓 HuggingFace - Run AI workloads on any cloud, store on Hugging Face: zero-egress storage with SkyPilot
https://huggingface.co/blog/skypilot-hf-storage
https://huggingface.co/blog/skypilot-hf-storage
huggingface.co
Run AI workloads on any cloud, store on Hugging Face: zero-egress storage with SkyPilot
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Meta AI - Introducing Muse Image and Muse Video
Muse Image follows instructions faithfully, edits with precision, composes from multiple references, and draws on Instagram for social context. Muse Video delivers exceptional visual fidelity with native audio support.
https://ai.meta.com/blog/introducing-muse-image-muse-video-msl/
Muse Image follows instructions faithfully, edits with precision, composes from multiple references, and draws on Instagram for social context. Muse Video delivers exceptional visual fidelity with native audio support.
https://ai.meta.com/blog/introducing-muse-image-muse-video-msl/
Meta AI
Introducing Muse Image and Muse Video
Muse Image follows instructions faithfully, edits with precision, composes from multiple references, and draws on Instagram for social context. Muse Video delivers exceptional visual fidelity with native audio support.
📰 OpenAI - Australian Payments Plus moves faster with ChatGPT and Codex
https://openai.com/index/australian-payments-plus
https://openai.com/index/australian-payments-plus
OpenAI
Australian Payments Plus moves faster with ChatGPT and Codex
See how Australian Payments Plus uses ChatGPT Enterprise and Codex to move faster through payments complexity. AP+ saves time, improves quality, and keeps human judgment central.
📰 HuggingFace - From Hugging Face to Amazon SageMaker Studio in one click
https://huggingface.co/blog/amazon/one-click-to-sagemaker-studio
https://huggingface.co/blog/amazon/one-click-to-sagemaker-studio
📰 NVIDIA - Building an Analysis AI Agent for Industrial Alarm Management with NVIDIA Nemotron
Industrial machinery generates more alarms than technicians can triage. For each important alarm requiring follow-up, the technician pulls historical context…
https://developer.nvidia.com/blog/building-an-analysis-ai-agent-for-industrial-alarm-management-with-nvidia-nemotron/
📰 NVIDIA - NVIDIA Vera CPU Boosts AI Factory Throughput to Accelerate Agentic Workloads
Agentic systems turn model reasoning into action through multi-step workflows that combine inference, tool use, code execution, retrieval, orchestration…
https://developer.nvidia.com/blog/nvidia-vera-cpu-boosts-ai-factory-throughput-to-accelerate-agentic-workloads/
Industrial machinery generates more alarms than technicians can triage. For each important alarm requiring follow-up, the technician pulls historical context…
https://developer.nvidia.com/blog/building-an-analysis-ai-agent-for-industrial-alarm-management-with-nvidia-nemotron/
📰 NVIDIA - NVIDIA Vera CPU Boosts AI Factory Throughput to Accelerate Agentic Workloads
Agentic systems turn model reasoning into action through multi-step workflows that combine inference, tool use, code execution, retrieval, orchestration…
https://developer.nvidia.com/blog/nvidia-vera-cpu-boosts-ai-factory-throughput-to-accelerate-agentic-workloads/
NVIDIA Technical Blog
Building an Analysis AI Agent for Industrial Alarm Management with NVIDIA Nemotron
Industrial machinery generates more alarms than technicians can triage. For each important alarm requiring follow-up, the technician pulls historical context, determines the correct procedure…
🆕 [HF Models] ibm-granite - granite-swash-3b-a600m
https://huggingface.co/ibm-granite/granite-swash-3b-a600m
🆕 [HF Models] ibm-granite - granite-swash-2b
https://huggingface.co/ibm-granite/granite-swash-2b
https://huggingface.co/ibm-granite/granite-swash-3b-a600m
🆕 [HF Models] ibm-granite - granite-swash-2b
https://huggingface.co/ibm-granite/granite-swash-2b
huggingface.co
ibm-granite/granite-swash-3b-a600m · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
🆕 [HF Models] tencent - R3-rerank-0.6b
https://huggingface.co/tencent/R3-rerank-0.6b
🆕 [HF Models] tencent - R3-embedding-0.6b
https://huggingface.co/tencent/R3-embedding-0.6b
https://huggingface.co/tencent/R3-rerank-0.6b
🆕 [HF Models] tencent - R3-embedding-0.6b
https://huggingface.co/tencent/R3-embedding-0.6b
huggingface.co
tencent/R3-rerank-0.6b · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Mistral - Introducing Robostral Navigate
Introducing Robostral Navigate: 8B model achieving 76.6% on R2R-CE with just a single RGB camera. No depth sensors, LiDAR, or multiple cameras needed.
https://mistral.ai/news/robostral-navigate/
Introducing Robostral Navigate: 8B model achieving 76.6% on R2R-CE with just a single RGB camera. No depth sensors, LiDAR, or multiple cameras needed.
https://mistral.ai/news/robostral-navigate/
Mistral AI
Robostral Navigate: single-camera AI navigation | Mistral AI
Introducing Robostral Navigate: 8B model achieving 76.6% on R2R-CE with just a single RGB camera. No depth sensors, LiDAR, or multiple cameras needed.
📰 HuggingFace - Native-speed vLLM transformers modeling backend
https://huggingface.co/blog/native-speed-vllm-transformers-backend
https://huggingface.co/blog/native-speed-vllm-transformers-backend