Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference
https://developer.nvidia.com/blog/co-designing-ai-models-using-speculative-decoding-for-faster-llm-inference/
https://developer.nvidia.com/blog/co-designing-ai-models-using-speculative-decoding-for-faster-llm-inference/
NVIDIA Technical Blog
Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference
This post is the third in a series on AI model co-design. It explores how to accelerate LLM inference while maintaining accuracy using speculative decoding and offers five guidelines for selecting…
The Modern CUDA Toolbox in Practice: A Step-by-Step Optimization Walkthrough
https://developer.nvidia.com/blog/the-modern-cuda-toolbox-in-practice-a-step-by-step-optimization-walkthrough/
https://developer.nvidia.com/blog/the-modern-cuda-toolbox-in-practice-a-step-by-step-optimization-walkthrough/
NVIDIA Technical Blog
The Modern CUDA Toolbox in Practice: A Step-by-Step Optimization Walkthrough
NVIDIA CUDA remains the foundation of GPU-accelerated computing, powering everything from scientific simulations to large-scale AI training. But writing correct, maintainable, and performant CUDA code…
‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW
https://blogs.nvidia.com/blog/geforce-now-thursday-september-2026-games-list/
https://blogs.nvidia.com/blog/geforce-now-thursday-september-2026-games-list/
NVIDIA Blog
‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW
Stream ‘NBA 2K27’ with DLSS 5 across supported devices — plus, jump into ‘The Blood of Dawnwalker,’ ‘Onimusha: Way of the Sword’ and more.
Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026
https://blogs.nvidia.com/blog/local-ai-ifa-next-gen-agents-nv-pair-rtx-spark/
https://blogs.nvidia.com/blog/local-ai-ifa-next-gen-agents-nv-pair-rtx-spark/
NVIDIA Blog
Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026
At IFA 2026, NVIDIA and partners are providing faster inference and new tools that make agents easier to set up and run locally.
NVIDIA PAIR Virtual Inference Router Expands Available Compute on Your Local Network
https://developer.nvidia.com/blog/nvidia-pair-virtual-inference-router-expands-available-compute-on-your-local-network/
https://developer.nvidia.com/blog/nvidia-pair-virtual-inference-router-expands-available-compute-on-your-local-network/
NVIDIA Technical Blog
NVIDIA PAIR Virtual Inference Router Expands Available Compute on Your Local Network
AI agents are learning to do more by working together. A lead agent can break a complex task into smaller jobs and assign those jobs to specialized subagents. Additionally, users are starting to run…
How to Carry User Identity Across Federated Kubernetes and AI Platforms
https://developer.nvidia.com/blog/how-to-carry-user-identity-across-federated-kubernetes-and-ai-platforms/
https://developer.nvidia.com/blog/how-to-carry-user-identity-across-federated-kubernetes-and-ai-platforms/
NVIDIA Technical Blog
How to Carry User Identity Across Federated Kubernetes and AI Platforms
Modern AI platforms are no longer a single application behind one login screen. A user may start in a central portal, open a governed dataset, launch a notebook where that data resides…
Frontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA Jetson
https://developer.nvidia.com/blog/frontier-reasoning-reaches-the-edge-how-to-deploy-and-optimize-models-on-nvidia-jetson/
https://developer.nvidia.com/blog/frontier-reasoning-reaches-the-edge-how-to-deploy-and-optimize-models-on-nvidia-jetson/
NVIDIA Technical Blog
Frontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA Jetson
Running reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run locally on edge hardware.
Building a Memory-Driven Agent with NVIDIA NemoClaw
https://developer.nvidia.com/blog/building-a-memory-driven-agent-with-nvidia-nemoclaw/
https://developer.nvidia.com/blog/building-a-memory-driven-agent-with-nvidia-nemoclaw/
NVIDIA Technical Blog
Building a Memory-Driven Agent with NVIDIA NemoClaw
Enterprise work spans messages, decisions, projects, and obligations that change over time. An AI agent that starts without this context must reconstruct it before contributing. To provide agents with…
Introducing CUDA Rust: Two Tracks for Writing GPU Kernels
https://developer.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernels/
https://developer.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernels/
NVIDIA Technical Blog
Introducing CUDA Rust: Two Tracks for Writing GPU Kernels
In September 2026, NVIDIA announced it is leaning into native GPU programming in Rust. CUDA C++ and CUDA Python are mature, enterprise-grade toolchains, and NVIDIA will be growing and maturing CUDA…
CUDA Toolkit 13.4 Adds Windows on Arm Support and Greater Control over Shared GPUs
https://developer.nvidia.com/blog/cuda-toolkit-13-4-adds-windows-on-arm-support-and-greater-control-over-shared-gpus/
https://developer.nvidia.com/blog/cuda-toolkit-13-4-adds-windows-on-arm-support-and-greater-control-over-shared-gpus/
NVIDIA Technical Blog
CUDA Toolkit 13.4 Adds Windows on Arm Support and Greater Control over Shared GPUs
Every NVIDIA CUDA Toolkit release adds functionality and performance improvements that help developers get more from NVIDIA GPUs and the broader NVIDIA software platform. CUDA Toolkit 13.4…
When to Use Encode-Prefill-Decode Disaggregation to Accelerate Multimodal Model Serving
https://developer.nvidia.com/blog/when-to-use-encode-prefill-decode-disaggregation-to-accelerate-multimodal-model-serving/
https://developer.nvidia.com/blog/when-to-use-encode-prefill-decode-disaggregation-to-accelerate-multimodal-model-serving/
NVIDIA Technical Blog
When to Use Encode-Prefill-Decode Disaggregation to Accelerate Multimodal Model Serving
Encode-prefill-decode (EPD) disaggregation is an inference optimization technique for multimodal models that separates the vision encoder stage from the prefill and decode stages. It is most effective…
NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC
https://blogs.nvidia.com/blog/ibc-news-2026/
https://blogs.nvidia.com/blog/ibc-news-2026/
NVIDIA Blog
NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC
At IBC, NVIDIA is announcing a major expansion to NVIDIA AI for Media to unlock new ways to understand motion, verify and enhance video, localize programming and build AI-powered media applications.
From Wafer-Out to First Token: Codifying Supply Chain Expertise with Nemotron and Palantir Foundry
https://developer.nvidia.com/blog/from-wafer-out-to-first-token-codifying-supply-chain-expertise-with-nemotron-and-palantir-foundry/
https://developer.nvidia.com/blog/from-wafer-out-to-first-token-codifying-supply-chain-expertise-with-nemotron-and-palantir-foundry/
NVIDIA Technical Blog
From Wafer-Out to First Token: Codifying Supply Chain Expertise with Nemotron and Palantir Foundry
NVIDIA has one of the largest and most complex supply chains in the world, and its performance is measured from wafer-out to first token. The interval is in two parts. Time-to-rack runs from silicon…
Boots on the Ground: ‘WARDOGS’ Goes All Out on GeForce NOW at Early-Access Launch
https://blogs.nvidia.com/blog/geforce-now-thursday-wardogs/
https://blogs.nvidia.com/blog/geforce-now-thursday-wardogs/
NVIDIA Blog
Boots on the Ground: ‘WARDOGS’ Goes All Out on GeForce NOW at Early-Access Launch
WARDOGS drops onto the cloud at launch, alongside the Valheim 1.0 Deep North update and Bus Simulator 27 — part of nine new games.
d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment
https://blogs.nvidia.com/blog/d-matrix-nvlink-fusion/
https://blogs.nvidia.com/blog/d-matrix-nvlink-fusion/
NVIDIA Blog
d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment
AI inference chipmaker d-Matrix today announced it will use NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA’s AI infrastructure platform — joining a growing roster of ecosystem partners. By connecting Raptor to NVIDIA NVLink scale…
Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies
https://blogs.nvidia.com/blog/robotaxi-leaders-full-stack-open-platform/
https://blogs.nvidia.com/blog/robotaxi-leaders-full-stack-open-platform/
NVIDIA Blog
Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies
NVIDIA’s modular, full-stack robotaxi pipeline — a three-computer solution spanning AI training, simulation and in-vehicle computing — is being adopted across the robotaxi ecosystem.
Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video
https://blogs.nvidia.com/blog/skild-ai-s1-physical-ai/
https://blogs.nvidia.com/blog/skild-ai-s1-physical-ai/
NVIDIA Blog
Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video
Skild AI’s S1 robotic foundation model harnesses NVIDIA technologies spanning synthetic data generation, model training, simulation and real-world deployment.
High-Throughput Structure Prediction with BioNeMo Inference Runtime
https://developer.nvidia.com/blog/high-throughput-structure-prediction-with-bionemo-inference-runtime/
https://developer.nvidia.com/blog/high-throughput-structure-prediction-with-bionemo-inference-runtime/
NVIDIA Technical Blog
High-Throughput Structure Prediction with BioNeMo Inference Runtime
Biomolecular structure prediction is now often run at proteome scale, where the goal is to move an entire worklist through the pipeline efficiently. NVIDIA BioNeMo Inference Runtime (BioIR) helps…
How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra
https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/
https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/
NVIDIA Technical Blog
How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra
Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while…
Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX
https://blogs.nvidia.com/blog/local-ai-perplexity-windows-pcs/
https://blogs.nvidia.com/blog/local-ai-perplexity-windows-pcs/
NVIDIA Blog
Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX
Perplexity Portable Computer is now available on Windows, powered by NVIDIA RTX GPUs. Run local AI agent workflows without cloud credits.