How Generative Recommenders Are Redefining RecSys at Scale
https://developer.nvidia.com/blog/how-generative-recommenders-are-redefining-recsys-at-scale/
https://developer.nvidia.com/blog/how-generative-recommenders-are-redefining-recsys-at-scale/
NVIDIA Technical Blog
How Generative Recommenders Are Redefining RecSys at Scale
Recommender systems (RecSys) are one of the most ubiquitous machine learning problems in the consumer internet industry yet notoriously difficult to train and serve at scale. The advent of LLMs has…
NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents
https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/
https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/
NVIDIA Technical Blog
NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents
A frontier language model is only one component of an AI agent. The surrounding agent system—often called a harness—determines how the model receives context, uses tools, maintains state…
Where Security Fits in an AI Agent Stack
https://developer.nvidia.com/blog/where-security-fits-in-an-ai-agent-stack/
https://developer.nvidia.com/blog/where-security-fits-in-an-ai-agent-stack/
NVIDIA Technical Blog
Where Security Fits in an AI Agent Stack
As AI agents become more capable and operate over longer horizons, building security and trust into the applications they power becomes increasingly important. Drawing on work with NVIDIA OpenShell…
Maximizing AI Factory Performance per Watt with NVIDIA DSX MaxLPS
https://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/
https://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/
NVIDIA Technical Blog
Maximizing AI Factory Performance per Watt with NVIDIA DSX MaxLPS
AI factories are power-constrained industrial systems. The question is no longer how many GPUs fit in a data center, but how much AI output each available megawatt can deliver.
GPU-Accelerated Clustering for Financial Instruments at Scale
https://developer.nvidia.com/blog/gpu-accelerated-clustering-for-financial-instruments-at-scale/
https://developer.nvidia.com/blog/gpu-accelerated-clustering-for-financial-instruments-at-scale/
NVIDIA Technical Blog
GPU-Accelerated Clustering for Financial Instruments at Scale
Use AdaptGrow, a GPU-accelerated matrix factorization algorithm, to turn rolling correlation and tail-dependence matrices into hard clusters, soft factor loadings, and structural-break signals at…
👍5
Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
https://blogs.nvidia.com/blog/vera-rubin-nvl72-efficiency-ai-agents/
https://blogs.nvidia.com/blog/vera-rubin-nvl72-efficiency-ai-agents/
NVIDIA Blog
Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
New on-silicon performance data measured by NVIDIA using real-world agentic coding trajectories shows Vera Rubin NVL72 systems deliver 30x higher throughput per megawatt and 35x lower token costs than NVIDIA GB300 NVL72.
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/
https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/
NVIDIA Blog
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems.…
Leading Publishers Bring Blockbuster PC Games and Technology to NVIDIA RTX Spark
https://blogs.nvidia.com/blog/gamescom-rtx-spark-pc-games-technology/
https://blogs.nvidia.com/blog/gamescom-rtx-spark-pc-games-technology/
NVIDIA Blog
Leading Publishers Bring Blockbuster PC Games and Technology to NVIDIA RTX Spark
NVIDIA is bringing the next wave of RTX gaming to the Gamescom conference running this week in Cologne, Germany, with support for new games, anti-cheat technologies and increased visual quality.
CUDA Python 1.0: Stable APIs, One Foundation, Full Platform Access
https://developer.nvidia.com/blog/cuda-python-1-0-stable-apis-one-foundation-full-platform-access/
https://developer.nvidia.com/blog/cuda-python-1-0-stable-apis-one-foundation-full-platform-access/
NVIDIA Technical Blog
CUDA Python 1.0: Stable APIs, One Foundation, Full Platform Access
For years, a Python developer who needed a GPU had two realistic choices: Learn NVIDIA CUDA C++ well enough to write an extension, set up a build toolchain, and maintain bindings back to Python…
Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo
https://developer.nvidia.com/blog/restore-llm-inference-capacity-in-seconds-with-shadow-engine-recovery-in-nvidia-dynamo/
https://developer.nvidia.com/blog/restore-llm-inference-capacity-in-seconds-with-shadow-engine-recovery-in-nvidia-dynamo/
NVIDIA Technical Blog
Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo
When an LLM engine process fails, the standard recovery path involves a cold restart. This requires loading weights into HBM from storage, compiling kernels, and capturing NVIDIA CUDA graphs.
Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Coding
https://developer.nvidia.com/blog/experiment-with-qwen3-8-flash-next-176b-model-on-nvidia-gb300-nvl72-for-agentic-coding/
https://developer.nvidia.com/blog/experiment-with-qwen3-8-flash-next-176b-model-on-nvidia-gb300-nvl72-for-agentic-coding/
NVIDIA Technical Blog
Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Coding
Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate. It’s a multimodal mixture-of-experts (MoE)…
How to Train a Cross-Embodiment Robot Navigation Policy with AI Agents
https://developer.nvidia.com/blog/how-to-train-a-cross-embodiment-robot-navigation-policy-with-ai-agents/
https://developer.nvidia.com/blog/how-to-train-a-cross-embodiment-robot-navigation-policy-with-ai-agents/
NVIDIA Technical Blog
How to Train a Cross-Embodiment Robot Navigation Policy with AI Agents
Navigation enables a robot to turn perception and motion into purposeful autonomy. Unlike locomotion, which produces stable movement, navigation must be used to continuously localize the robot…
NVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure
https://developer.nvidia.com/blog/nvidia-nvlink-fusion-brings-nvhbm-to-next-generation-ai-infrastructure/
https://developer.nvidia.com/blog/nvidia-nvlink-fusion-brings-nvhbm-to-next-generation-ai-infrastructure/
NVIDIA Technical Blog
NVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure
AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads, hyperscalers and AI-native companies are…
NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory
https://blogs.nvidia.com/blog/nvlink-fusion-nvhbm-custom-high-bandwidth-memory/
https://blogs.nvidia.com/blog/nvlink-fusion-nvhbm-custom-high-bandwidth-memory/
NVIDIA Blog
NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory
Amazon’s Annapurna Labs will be the first to collaborate on NVHBM technology alongside NVLink Fusion.
GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026
https://blogs.nvidia.com/blog/geforce-now-thursday-gamescom-2026/
https://blogs.nvidia.com/blog/geforce-now-thursday-gamescom-2026/
NVIDIA Blog
GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026
GeForce NOW unveils DLSS 4.5 controls, expanded device support and launch-day games at Gamescom, plus 13 new games this week.
Deploy an Open Model from Checkpoint to Inference in Two Commands with NVIDIA TensorRT Model Connect
https://developer.nvidia.com/blog/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect/
https://developer.nvidia.com/blog/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect/
NVIDIA Technical Blog
Deploy an Open Model from Checkpoint to Inference in Two Commands with NVIDIA TensorRT Model Connect
Open AI models are evolving faster than ever, but bringing them into native applications can still require model-specific conversion, preprocessing, post-processing, and runtime code.
Scale AV Perception Across Vehicle Platforms with NVIDIA Omniverse NuRec
https://developer.nvidia.com/blog/scale-av-perception-across-vehicle-platforms-with-nvidia-omniverse-nurec/
https://developer.nvidia.com/blog/scale-av-perception-across-vehicle-platforms-with-nvidia-omniverse-nurec/
NVIDIA Technical Blog
Scale AV Perception Across Vehicle Platforms with NVIDIA Omniverse NuRec
A perception stack is shaped by the vehicle that carries it. Move the same software to a new carline—for example, from an SUV to a sedan or another vehicle variant in the portfolio—and its perception…
Run NVIDIA BioNeMo NIM Microservices for Protein Structure Prediction in Claude Science
https://developer.nvidia.com/blog/run-nvidia-bionemo-nim-microservices-for-protein-structure-prediction-in-claude-science/
https://developer.nvidia.com/blog/run-nvidia-bionemo-nim-microservices-for-protein-structure-prediction-in-claude-science/
NVIDIA Technical Blog
Run NVIDIA BioNeMo NIM Microservices for Protein Structure Prediction in Claude Science
Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next. First proving their value in software…
How to Size GPUs for AI Inference and TCO Without Overspending
https://developer.nvidia.com/blog/how-to-size-gpus-for-ai-inference-and-tco-without-overspending/
https://developer.nvidia.com/blog/how-to-size-gpus-for-ai-inference-and-tco-without-overspending/
Building an Adaptive Agentic Cybersecurity System with NVIDIA Nemotron
https://developer.nvidia.com/blog/building-an-adaptive-agentic-cybersecurity-system-with-nvidia-nemotron/
https://developer.nvidia.com/blog/building-an-adaptive-agentic-cybersecurity-system-with-nvidia-nemotron/
NVIDIA Technical Blog
Building an Adaptive Agentic Cybersecurity System with NVIDIA Nemotron
AI is changing the pace of cybersecurity. Agentic systems can coordinate work and pursue complex objectives over long horizons. Security teams are beginning to apply agents across security operations…