📰 NVIDIA - Enable Real-Time AI for High-Speed Data Acquisition with DAQIRI
When AlphaFold2 revolutionized drug discovery in 2020, its success relied entirely on the roughly 170,000 protein structures collected by scientists since 1971…
https://developer.nvidia.com/blog/enable-real-time-ai-for-high-speed-data-acquisition-with-daqiri/
📰 NVIDIA - Inside NVIDIA Halos for Robotics: A Full-Stack Functional Safety System for Physical AI
Physical AI—robots working autonomously alongside people in factories, warehouses, hospitals, and homes—is arriving faster than most expected.
https://developer.nvidia.com/blog/inside-nvidia-halos-for-robotics-a-full-stack-functional-safety-system-for-physical-ai/
When AlphaFold2 revolutionized drug discovery in 2020, its success relied entirely on the roughly 170,000 protein structures collected by scientists since 1971…
https://developer.nvidia.com/blog/enable-real-time-ai-for-high-speed-data-acquisition-with-daqiri/
📰 NVIDIA - Inside NVIDIA Halos for Robotics: A Full-Stack Functional Safety System for Physical AI
Physical AI—robots working autonomously alongside people in factories, warehouses, hospitals, and homes—is arriving faster than most expected.
https://developer.nvidia.com/blog/inside-nvidia-halos-for-robotics-a-full-stack-functional-safety-system-for-physical-ai/
NVIDIA Technical Blog
Enable Real-Time AI for High-Speed Data Acquisition with DAQIRI
When AlphaFold2 revolutionized drug discovery in 2020, its success relied entirely on the roughly 170,000 protein structures collected by scientists since 1971 and preserved in the Protein Data Bank.
📰 HuggingFace - Shipping huggingface_hub every week with AI, open tools, and a human in the loop
https://huggingface.co/blog/huggingface-hub-release-ci
🔓 HuggingFace - We got local models to triage the OpenClaw repo for FREE!*
https://huggingface.co/blog/local-models-pr-triage
https://huggingface.co/blog/huggingface-hub-release-ci
🔓 HuggingFace - We got local models to triage the OpenClaw repo for FREE!*
https://huggingface.co/blog/local-models-pr-triage
huggingface.co
Shipping huggingface_hub every week with AI, open tools, and a human in the loop
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 HuggingFace - Build real agentic apps using CUGA: two dozen working examples on a lightweight harness
https://huggingface.co/blog/ibm-research/cuga-apps
https://huggingface.co/blog/ibm-research/cuga-apps
huggingface.co
Build real agentic apps using CUGA: two dozen working examples on a lightweight harness
A Blog post by IBM Research on Hugging Face
📰 PyTorch - Serving DeepSeek-V4 on GB300 with SGLang: 5x Higher Throughput at the Same Interactivity Since Day-0
TL;DR: DeepSeek-V4 support was live in SGLang on Day-0, but the Day-0 stack was only the starting point. Since launch, we have coordinated a set of kernel, runtime, and hardening...
https://pytorch.org/blog/serving-deepseek-v4-on-gb300-with-sglang-5x-higher-throughput-at-the-same-interactivity-since-day-0/
TL;DR: DeepSeek-V4 support was live in SGLang on Day-0, but the Day-0 stack was only the starting point. Since launch, we have coordinated a set of kernel, runtime, and hardening...
https://pytorch.org/blog/serving-deepseek-v4-on-gb300-with-sglang-5x-higher-throughput-at-the-same-interactivity-since-day-0/
🔓 HuggingFace - Experimenting with the proposed Cross-Origin Storage API in Transformers.js
https://huggingface.co/blog/cross-origin-storage
https://huggingface.co/blog/cross-origin-storage
huggingface.co
Experimenting with the proposed Cross-Origin Storage API in Transformers.js
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 NVIDIA - Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding
As AI systems move from single-turn interactions to coordinated multiagent workflows, low-latency inference becomes increasingly important.
https://developer.nvidia.com/blog/boost-inference-performance-up-to-15x-on-nvidia-blackwell-using-dflash-speculative-decoding/
📰 NVIDIA - How Telcos Build Autonomous Networks with Agentic AI
Telecom operators are adopting AI across network operations, customer care, and back-office workflows, but most are still early in the journey to autonomy.
https://developer.nvidia.com/blog/how-telcos-build-autonomous-networks-with-agentic-ai/
As AI systems move from single-turn interactions to coordinated multiagent workflows, low-latency inference becomes increasingly important.
https://developer.nvidia.com/blog/boost-inference-performance-up-to-15x-on-nvidia-blackwell-using-dflash-speculative-decoding/
📰 NVIDIA - How Telcos Build Autonomous Networks with Agentic AI
Telecom operators are adopting AI across network operations, customer care, and back-office workflows, but most are still early in the journey to autonomy.
https://developer.nvidia.com/blog/how-telcos-build-autonomous-networks-with-agentic-ai/
NVIDIA Technical Blog
Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding
As AI systems move from single-turn interactions to coordinated multiagent workflows, low-latency inference becomes increasingly important. Autoregressive LLMs generate tokens sequentially…
📰 HuggingFace - Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel
https://huggingface.co/blog/nvidia/accelerating-fine-tuning-nvidia-nemo-automodel
https://huggingface.co/blog/nvidia/accelerating-fine-tuning-nvidia-nemo-automodel
huggingface.co
Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel
A Blog post by NVIDIA on Hugging Face
🔓 HuggingFace - Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
https://huggingface.co/blog/ffasr-leaderboard
https://huggingface.co/blog/ffasr-leaderboard
huggingface.co
Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Claude Blog - Building effective human-agent teams
https://claude.com/blog/building-effective-human-agent-teams
📰 Claude Blog - Agent identity in Claude Tag: a new access model for autonomous, team-wide AI
https://claude.com/blog/agent-identity-access-model
📰 Claude Blog - The full Claude Desktop experience on AWS, Google Cloud, and Microsoft Foundry
https://claude.com/blog/the-full-claude-desktop-experience-on-aws-google-cloud-and-microsoft-foundry
https://claude.com/blog/building-effective-human-agent-teams
📰 Claude Blog - Agent identity in Claude Tag: a new access model for autonomous, team-wide AI
https://claude.com/blog/agent-identity-access-model
📰 Claude Blog - The full Claude Desktop experience on AWS, Google Cloud, and Microsoft Foundry
https://claude.com/blog/the-full-claude-desktop-experience-on-aws-google-cloud-and-microsoft-foundry
Claude
Lessons from Anthropic on building effective human-agent teams | Claude by Anthropic
The way we work with AI is evolving from a single-player to a multiplayer experience, where humans and agents work together as a team to achieve shared goals. The Anthropic team shares examples of this new way of working in action.
📰 Qwen Research - Qwen-AgentWorld: Language World Models for General Agents
https://qwen.ai/blog?id=qwen-agentworld
https://qwen.ai/blog?id=qwen-agentworld
qwen.ai
Qwen Studio
Qwen Studio offers comprehensive functionality spanning chatbot, image and video understanding, image generation, document processing, web search integration, tool utilization, and artifacts.
🔄 [GitHub Releases] flashinfer-ai/flashinfer - Release v0.6.13
https://github.com/flashinfer-ai/flashinfer/releases/tag/v0.6.13
https://github.com/flashinfer-ai/flashinfer/releases/tag/v0.6.13
GitHub
Release Release v0.6.13 · flashinfer-ai/flashinfer
What's Changed
Run high-likelihood OOM culprits separately, record memory usage and test duration for analysis by @dierksen in #2961
fix(autotuner): differentiate file cache entries by runner ...
Run high-likelihood OOM culprits separately, record memory usage and test duration for analysis by @dierksen in #2961
fix(autotuner): differentiate file cache entries by runner ...
🆕 [HF Models] LiquidAI - LFM2.5-230M
https://huggingface.co/LiquidAI/LFM2.5-230M
🔓 [HF Models] LiquidAI - LFM2.5-230M-GGUF
https://huggingface.co/LiquidAI/LFM2.5-230M-GGUF
🔓 [HF Models] LiquidAI - LFM2.5-230M-Base
https://huggingface.co/LiquidAI/LFM2.5-230M-Base
https://huggingface.co/LiquidAI/LFM2.5-230M
🔓 [HF Models] LiquidAI - LFM2.5-230M-GGUF
https://huggingface.co/LiquidAI/LFM2.5-230M-GGUF
🔓 [HF Models] LiquidAI - LFM2.5-230M-Base
https://huggingface.co/LiquidAI/LFM2.5-230M-Base
huggingface.co
LiquidAI/LFM2.5-230M · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 PyTorch - TokenSpeed-Kernel: Portable APIs and High-Performance Kernels for Multi-Silicon LLM Inference
TL;DR The TokenSpeed-kernel is a standalone, open-source subsystem designed to solve backend complexity in LLM inference. It introduces a clean, layered API and registry system that decouples the high-level runtime...
https://pytorch.org/blog/lightseek-tokenspeed-kernel/
TL;DR The TokenSpeed-kernel is a standalone, open-source subsystem designed to solve backend complexity in LLM inference. It introduces a clean, layered API and registry system that decouples the high-level runtime...
https://pytorch.org/blog/lightseek-tokenspeed-kernel/
📰 HuggingFace - Which tokens does a hybrid model predict better?
https://huggingface.co/blog/allenai/hybrid-token-prediction
📰 HuggingFace - Run a vLLM Server on HF Jobs in One Command
https://huggingface.co/blog/vllm-jobs
https://huggingface.co/blog/allenai/hybrid-token-prediction
📰 HuggingFace - Run a vLLM Server on HF Jobs in One Command
https://huggingface.co/blog/vllm-jobs
huggingface.co
Which tokens does a hybrid model predict better?
A Blog post by Ai2 on Hugging Face
📰 NVIDIA - Scaling AI Inference Across Multiple GPUs Using NVIDIA TensorRT with Multi-Device Inference Support
Generative AI workloads are rapidly outgrowing the memory and compute budget of single GPUs. For inference developers building media generation pipelines…
https://developer.nvidia.com/blog/scaling-ai-inference-across-multiple-gpus-using-nvidia-tensorrt-with-multi-device-inference-support/
Generative AI workloads are rapidly outgrowing the memory and compute budget of single GPUs. For inference developers building media generation pipelines…
https://developer.nvidia.com/blog/scaling-ai-inference-across-multiple-gpus-using-nvidia-tensorrt-with-multi-device-inference-support/
NVIDIA Technical Blog
Scaling AI Inference Across Multiple GPUs Using NVIDIA TensorRT with Multi-Device Inference Support
Generative AI workloads are rapidly outgrowing the memory and compute budget of single GPUs. For inference developers building media generation pipelines, the challenge is scaling across multiple…