📰 HuggingFace - Build real agentic apps using CUGA: two dozen working examples on a lightweight harness
https://huggingface.co/blog/ibm-research/cuga-apps
https://huggingface.co/blog/ibm-research/cuga-apps
huggingface.co
Build real agentic apps using CUGA: two dozen working examples on a lightweight harness
A Blog post by IBM Research on Hugging Face
📰 PyTorch - Serving DeepSeek-V4 on GB300 with SGLang: 5x Higher Throughput at the Same Interactivity Since Day-0
TL;DR: DeepSeek-V4 support was live in SGLang on Day-0, but the Day-0 stack was only the starting point. Since launch, we have coordinated a set of kernel, runtime, and hardening...
https://pytorch.org/blog/serving-deepseek-v4-on-gb300-with-sglang-5x-higher-throughput-at-the-same-interactivity-since-day-0/
TL;DR: DeepSeek-V4 support was live in SGLang on Day-0, but the Day-0 stack was only the starting point. Since launch, we have coordinated a set of kernel, runtime, and hardening...
https://pytorch.org/blog/serving-deepseek-v4-on-gb300-with-sglang-5x-higher-throughput-at-the-same-interactivity-since-day-0/
🔓 HuggingFace - Experimenting with the proposed Cross-Origin Storage API in Transformers.js
https://huggingface.co/blog/cross-origin-storage
https://huggingface.co/blog/cross-origin-storage
huggingface.co
Experimenting with the proposed Cross-Origin Storage API in Transformers.js
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 NVIDIA - Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding
As AI systems move from single-turn interactions to coordinated multiagent workflows, low-latency inference becomes increasingly important.
https://developer.nvidia.com/blog/boost-inference-performance-up-to-15x-on-nvidia-blackwell-using-dflash-speculative-decoding/
📰 NVIDIA - How Telcos Build Autonomous Networks with Agentic AI
Telecom operators are adopting AI across network operations, customer care, and back-office workflows, but most are still early in the journey to autonomy.
https://developer.nvidia.com/blog/how-telcos-build-autonomous-networks-with-agentic-ai/
As AI systems move from single-turn interactions to coordinated multiagent workflows, low-latency inference becomes increasingly important.
https://developer.nvidia.com/blog/boost-inference-performance-up-to-15x-on-nvidia-blackwell-using-dflash-speculative-decoding/
📰 NVIDIA - How Telcos Build Autonomous Networks with Agentic AI
Telecom operators are adopting AI across network operations, customer care, and back-office workflows, but most are still early in the journey to autonomy.
https://developer.nvidia.com/blog/how-telcos-build-autonomous-networks-with-agentic-ai/
NVIDIA Technical Blog
Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding
As AI systems move from single-turn interactions to coordinated multiagent workflows, low-latency inference becomes increasingly important. Autoregressive LLMs generate tokens sequentially…
📰 HuggingFace - Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel
https://huggingface.co/blog/nvidia/accelerating-fine-tuning-nvidia-nemo-automodel
https://huggingface.co/blog/nvidia/accelerating-fine-tuning-nvidia-nemo-automodel
huggingface.co
Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel
A Blog post by NVIDIA on Hugging Face
🔓 HuggingFace - Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
https://huggingface.co/blog/ffasr-leaderboard
https://huggingface.co/blog/ffasr-leaderboard
huggingface.co
Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Claude Blog - Building effective human-agent teams
https://claude.com/blog/building-effective-human-agent-teams
📰 Claude Blog - Agent identity in Claude Tag: a new access model for autonomous, team-wide AI
https://claude.com/blog/agent-identity-access-model
📰 Claude Blog - The full Claude Desktop experience on AWS, Google Cloud, and Microsoft Foundry
https://claude.com/blog/the-full-claude-desktop-experience-on-aws-google-cloud-and-microsoft-foundry
https://claude.com/blog/building-effective-human-agent-teams
📰 Claude Blog - Agent identity in Claude Tag: a new access model for autonomous, team-wide AI
https://claude.com/blog/agent-identity-access-model
📰 Claude Blog - The full Claude Desktop experience on AWS, Google Cloud, and Microsoft Foundry
https://claude.com/blog/the-full-claude-desktop-experience-on-aws-google-cloud-and-microsoft-foundry
Claude
Lessons from Anthropic on building effective human-agent teams | Claude by Anthropic
The way we work with AI is evolving from a single-player to a multiplayer experience, where humans and agents work together as a team to achieve shared goals. The Anthropic team shares examples of this new way of working in action.
📰 Qwen Research - Qwen-AgentWorld: Language World Models for General Agents
https://qwen.ai/blog?id=qwen-agentworld
https://qwen.ai/blog?id=qwen-agentworld
qwen.ai
Qwen Studio
Qwen Studio offers comprehensive functionality spanning chatbot, image and video understanding, image generation, document processing, web search integration, tool utilization, and artifacts.
🔄 [GitHub Releases] flashinfer-ai/flashinfer - Release v0.6.13
https://github.com/flashinfer-ai/flashinfer/releases/tag/v0.6.13
https://github.com/flashinfer-ai/flashinfer/releases/tag/v0.6.13
GitHub
Release Release v0.6.13 · flashinfer-ai/flashinfer
What's Changed
Run high-likelihood OOM culprits separately, record memory usage and test duration for analysis by @dierksen in #2961
fix(autotuner): differentiate file cache entries by runner ...
Run high-likelihood OOM culprits separately, record memory usage and test duration for analysis by @dierksen in #2961
fix(autotuner): differentiate file cache entries by runner ...
🆕 [HF Models] LiquidAI - LFM2.5-230M
https://huggingface.co/LiquidAI/LFM2.5-230M
🔓 [HF Models] LiquidAI - LFM2.5-230M-GGUF
https://huggingface.co/LiquidAI/LFM2.5-230M-GGUF
🔓 [HF Models] LiquidAI - LFM2.5-230M-Base
https://huggingface.co/LiquidAI/LFM2.5-230M-Base
https://huggingface.co/LiquidAI/LFM2.5-230M
🔓 [HF Models] LiquidAI - LFM2.5-230M-GGUF
https://huggingface.co/LiquidAI/LFM2.5-230M-GGUF
🔓 [HF Models] LiquidAI - LFM2.5-230M-Base
https://huggingface.co/LiquidAI/LFM2.5-230M-Base
huggingface.co
LiquidAI/LFM2.5-230M · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 PyTorch - TokenSpeed-Kernel: Portable APIs and High-Performance Kernels for Multi-Silicon LLM Inference
TL;DR The TokenSpeed-kernel is a standalone, open-source subsystem designed to solve backend complexity in LLM inference. It introduces a clean, layered API and registry system that decouples the high-level runtime...
https://pytorch.org/blog/lightseek-tokenspeed-kernel/
TL;DR The TokenSpeed-kernel is a standalone, open-source subsystem designed to solve backend complexity in LLM inference. It introduces a clean, layered API and registry system that decouples the high-level runtime...
https://pytorch.org/blog/lightseek-tokenspeed-kernel/
📰 HuggingFace - Which tokens does a hybrid model predict better?
https://huggingface.co/blog/allenai/hybrid-token-prediction
📰 HuggingFace - Run a vLLM Server on HF Jobs in One Command
https://huggingface.co/blog/vllm-jobs
https://huggingface.co/blog/allenai/hybrid-token-prediction
📰 HuggingFace - Run a vLLM Server on HF Jobs in One Command
https://huggingface.co/blog/vllm-jobs
huggingface.co
Which tokens does a hybrid model predict better?
A Blog post by Ai2 on Hugging Face
📰 NVIDIA - Scaling AI Inference Across Multiple GPUs Using NVIDIA TensorRT with Multi-Device Inference Support
Generative AI workloads are rapidly outgrowing the memory and compute budget of single GPUs. For inference developers building media generation pipelines…
https://developer.nvidia.com/blog/scaling-ai-inference-across-multiple-gpus-using-nvidia-tensorrt-with-multi-device-inference-support/
Generative AI workloads are rapidly outgrowing the memory and compute budget of single GPUs. For inference developers building media generation pipelines…
https://developer.nvidia.com/blog/scaling-ai-inference-across-multiple-gpus-using-nvidia-tensorrt-with-multi-device-inference-support/
NVIDIA Technical Blog
Scaling AI Inference Across Multiple GPUs Using NVIDIA TensorRT with Multi-Device Inference Support
Generative AI workloads are rapidly outgrowing the memory and compute budget of single GPUs. For inference developers building media generation pipelines, the challenge is scaling across multiple…
🆕 [HF Models] Qwen - Qwen3-ForcedAligner-0.6B-hf
https://huggingface.co/Qwen/Qwen3-ForcedAligner-0.6B-hf
🆕 [HF Models] Qwen - Qwen3-ASR-0.6B-hf
https://huggingface.co/Qwen/Qwen3-ASR-0.6B-hf
🆕 [HF Models] Qwen - Qwen3-ASR-1.7B-hf
https://huggingface.co/Qwen/Qwen3-ASR-1.7B-hf
https://huggingface.co/Qwen/Qwen3-ForcedAligner-0.6B-hf
🆕 [HF Models] Qwen - Qwen3-ASR-0.6B-hf
https://huggingface.co/Qwen/Qwen3-ASR-0.6B-hf
🆕 [HF Models] Qwen - Qwen3-ASR-1.7B-hf
https://huggingface.co/Qwen/Qwen3-ASR-1.7B-hf
huggingface.co
Qwen/Qwen3-ForcedAligner-0.6B-hf · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.