NVIDIA PAIR Virtual Inference Router Expands Available Compute on Your Local Network
https://developer.nvidia.com/blog/nvidia-pair-virtual-inference-router-expands-available-compute-on-your-local-network/
https://developer.nvidia.com/blog/nvidia-pair-virtual-inference-router-expands-available-compute-on-your-local-network/
NVIDIA Technical Blog
NVIDIA PAIR Virtual Inference Router Expands Available Compute on Your Local Network
AI agents are learning to do more by working together. A lead agent can break a complex task into smaller jobs and assign those jobs to specialized subagents. Additionally, users are starting to run…
How to Carry User Identity Across Federated Kubernetes and AI Platforms
https://developer.nvidia.com/blog/how-to-carry-user-identity-across-federated-kubernetes-and-ai-platforms/
https://developer.nvidia.com/blog/how-to-carry-user-identity-across-federated-kubernetes-and-ai-platforms/
NVIDIA Technical Blog
How to Carry User Identity Across Federated Kubernetes and AI Platforms
Modern AI platforms are no longer a single application behind one login screen. A user may start in a central portal, open a governed dataset, launch a notebook where that data resides…
Frontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA Jetson
https://developer.nvidia.com/blog/frontier-reasoning-reaches-the-edge-how-to-deploy-and-optimize-models-on-nvidia-jetson/
https://developer.nvidia.com/blog/frontier-reasoning-reaches-the-edge-how-to-deploy-and-optimize-models-on-nvidia-jetson/
NVIDIA Technical Blog
Frontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA Jetson
Running reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run locally on edge hardware.
Building a Memory-Driven Agent with NVIDIA NemoClaw
https://developer.nvidia.com/blog/building-a-memory-driven-agent-with-nvidia-nemoclaw/
https://developer.nvidia.com/blog/building-a-memory-driven-agent-with-nvidia-nemoclaw/
NVIDIA Technical Blog
Building a Memory-Driven Agent with NVIDIA NemoClaw
Enterprise work spans messages, decisions, projects, and obligations that change over time. An AI agent that starts without this context must reconstruct it before contributing. To provide agents with…
Introducing CUDA Rust: Two Tracks for Writing GPU Kernels
https://developer.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernels/
https://developer.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernels/
NVIDIA Technical Blog
Introducing CUDA Rust: Two Tracks for Writing GPU Kernels
In September 2026, NVIDIA announced it is leaning into native GPU programming in Rust. CUDA C++ and CUDA Python are mature, enterprise-grade toolchains, and NVIDIA will be growing and maturing CUDA…
CUDA Toolkit 13.4 Adds Windows on Arm Support and Greater Control over Shared GPUs
https://developer.nvidia.com/blog/cuda-toolkit-13-4-adds-windows-on-arm-support-and-greater-control-over-shared-gpus/
https://developer.nvidia.com/blog/cuda-toolkit-13-4-adds-windows-on-arm-support-and-greater-control-over-shared-gpus/
NVIDIA Technical Blog
CUDA Toolkit 13.4 Adds Windows on Arm Support and Greater Control over Shared GPUs
Every NVIDIA CUDA Toolkit release adds functionality and performance improvements that help developers get more from NVIDIA GPUs and the broader NVIDIA software platform. CUDA Toolkit 13.4…
When to Use Encode-Prefill-Decode Disaggregation to Accelerate Multimodal Model Serving
https://developer.nvidia.com/blog/when-to-use-encode-prefill-decode-disaggregation-to-accelerate-multimodal-model-serving/
https://developer.nvidia.com/blog/when-to-use-encode-prefill-decode-disaggregation-to-accelerate-multimodal-model-serving/
NVIDIA Technical Blog
When to Use Encode-Prefill-Decode Disaggregation to Accelerate Multimodal Model Serving
Encode-prefill-decode (EPD) disaggregation is an inference optimization technique for multimodal models that separates the vision encoder stage from the prefill and decode stages. It is most effective…
NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC
https://blogs.nvidia.com/blog/ibc-news-2026/
https://blogs.nvidia.com/blog/ibc-news-2026/
NVIDIA Blog
NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC
At IBC, NVIDIA is announcing a major expansion to NVIDIA AI for Media to unlock new ways to understand motion, verify and enhance video, localize programming and build AI-powered media applications.
From Wafer-Out to First Token: Codifying Supply Chain Expertise with Nemotron and Palantir Foundry
https://developer.nvidia.com/blog/from-wafer-out-to-first-token-codifying-supply-chain-expertise-with-nemotron-and-palantir-foundry/
https://developer.nvidia.com/blog/from-wafer-out-to-first-token-codifying-supply-chain-expertise-with-nemotron-and-palantir-foundry/
NVIDIA Technical Blog
From Wafer-Out to First Token: Codifying Supply Chain Expertise with Nemotron and Palantir Foundry
NVIDIA has one of the largest and most complex supply chains in the world, and its performance is measured from wafer-out to first token. The interval is in two parts. Time-to-rack runs from silicon…
Boots on the Ground: ‘WARDOGS’ Goes All Out on GeForce NOW at Early-Access Launch
https://blogs.nvidia.com/blog/geforce-now-thursday-wardogs/
https://blogs.nvidia.com/blog/geforce-now-thursday-wardogs/
NVIDIA Blog
Boots on the Ground: ‘WARDOGS’ Goes All Out on GeForce NOW at Early-Access Launch
WARDOGS drops onto the cloud at launch, alongside the Valheim 1.0 Deep North update and Bus Simulator 27 — part of nine new games.
d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment
https://blogs.nvidia.com/blog/d-matrix-nvlink-fusion/
https://blogs.nvidia.com/blog/d-matrix-nvlink-fusion/
NVIDIA Blog
d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment
AI inference chipmaker d-Matrix today announced it will use NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA’s AI infrastructure platform — joining a growing roster of ecosystem partners. By connecting Raptor to NVIDIA NVLink scale…
Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies
https://blogs.nvidia.com/blog/robotaxi-leaders-full-stack-open-platform/
https://blogs.nvidia.com/blog/robotaxi-leaders-full-stack-open-platform/
NVIDIA Blog
Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies
NVIDIA’s modular, full-stack robotaxi pipeline — a three-computer solution spanning AI training, simulation and in-vehicle computing — is being adopted across the robotaxi ecosystem.
Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video
https://blogs.nvidia.com/blog/skild-ai-s1-physical-ai/
https://blogs.nvidia.com/blog/skild-ai-s1-physical-ai/
NVIDIA Blog
Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video
Skild AI’s S1 robotic foundation model harnesses NVIDIA technologies spanning synthetic data generation, model training, simulation and real-world deployment.
High-Throughput Structure Prediction with BioNeMo Inference Runtime
https://developer.nvidia.com/blog/high-throughput-structure-prediction-with-bionemo-inference-runtime/
https://developer.nvidia.com/blog/high-throughput-structure-prediction-with-bionemo-inference-runtime/
NVIDIA Technical Blog
High-Throughput Structure Prediction with BioNeMo Inference Runtime
Biomolecular structure prediction is now often run at proteome scale, where the goal is to move an entire worklist through the pipeline efficiently. NVIDIA BioNeMo Inference Runtime (BioIR) helps…
How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra
https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/
https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/
NVIDIA Technical Blog
How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra
Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while…
Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX
https://blogs.nvidia.com/blog/local-ai-perplexity-windows-pcs/
https://blogs.nvidia.com/blog/local-ai-perplexity-windows-pcs/
NVIDIA Blog
Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX
Perplexity Portable Computer is now available on Windows, powered by NVIDIA RTX GPUs. Run local AI agent workflows without cloud credits.
Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine
https://developer.nvidia.com/blog/accelerating-dropless-moe-training-in-jax-with-nvidia-transformer-engine/
https://developer.nvidia.com/blog/accelerating-dropless-moe-training-in-jax-with-nvidia-transformer-engine/
NVIDIA Technical Blog
Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine
Mixture of experts (MoE) has become one of the defining architectural trends in large-scale AI model training. DeepSeek, Qwen, and Mixtral are examples of MoE models that match or exceed the…
Heart of the Matter: How a Major Children’s Hospital Uses Open Source NVIDIA AI for Cardiac Care
https://blogs.nvidia.com/blog/childrens-hospital-open-source-ai-cardiac-care/
https://blogs.nvidia.com/blog/childrens-hospital-open-source-ai-cardiac-care/
NVIDIA Blog
Heart of the Matter: How a Major Children’s Hospital Uses Open Source NVIDIA AI for Cardiac Care
Children’s Hospital of Philadelphia is using open source AI tools to model children’s hearts in seconds — with the goal of enabling safer, more precise care for kids with congenital heart disease.
AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories
https://blogs.nvidia.com/blog/ai-infra-summit-vera-rubin-dsx-energy-efficiencies-tokens-per-watt-ai-factories/
https://blogs.nvidia.com/blog/ai-infra-summit-vera-rubin-dsx-energy-efficiencies-tokens-per-watt-ai-factories/
NVIDIA Blog
AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for…
Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, Tuesday spoke on AI factory efficiency at the AI Infra Summit, the Santa Clara Convention Center event that has morphed into a Coachella of infrastructure tech. Before a packed…
From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production
https://blogs.nvidia.com/blog/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production/
https://blogs.nvidia.com/blog/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production/
NVIDIA Blog
From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production
On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked, Silicon Valley Power sent a signal to an AI factory to adjust its power consumption. Varun Sivaram was watching on Zoom with about forty others — his team…
Scaling Federated Learning Across Docker, Kubernetes, and Slurm with NVIDIA FLARE
https://developer.nvidia.com/blog/scaling-federated-learning-across-docker-kubernetes-and-slurm-with-nvidia-flare/
https://developer.nvidia.com/blog/scaling-federated-learning-across-docker-kubernetes-and-slurm-with-nvidia-flare/
NVIDIA Technical Blog
Scaling Federated Learning Across Docker, Kubernetes, and Slurm with NVIDIA FLARE
Federated learning (FL) projects often begin with a straightforward setup: one server, a few clients, and one dataset at each site. As those projects grow, the challenge shifts from running an…