Enhancing Goodput in Large-Scale LLM Training with Nonuniform Tensor Parallelism
https://developer.nvidia.com/blog/enhancing-goodput-in-large-scale-llm-training-with-nonuniform-tensor-parallelism/
https://developer.nvidia.com/blog/enhancing-goodput-in-large-scale-llm-training-with-nonuniform-tensor-parallelism/
NVIDIA Technical Blog
Enhancing Goodput in Large-Scale LLM Training with Nonuniform Tensor Parallelism
Training LLMs at massive scale brings unique infrastructure challenges, especially as jobs span thousands of GPUs and run for extended periods. The longer these jobs run, the greater the likelihood of…
NVIDIA and Hugging Face Bring New Models and Frameworks to LeRobot for the Open Robotics Community
https://blogs.nvidia.com/blog/hugging-face-lerobot-models-frameworks-open-robotics/
https://blogs.nvidia.com/blog/hugging-face-lerobot-models-frameworks-open-robotics/
NVIDIA Blog
NVIDIA and Hugging Face Bring New Models and Frameworks to LeRobot for the Open Robotics Community
New LeRobot integrations give developers open access to NVIDIA Isaac GR00T 1.7, Isaac Teleop, datasets and robotics workflows, with NVIDIA Cosmos 3 integration planned to bring frontier world models to open robotics development.
AI Innovators Adopt NVIDIA Vera — Why Max Single-Threaded CPU at Scale Matters
https://blogs.nvidia.com/blog/nvidia-vera-max-single-threaded-cpu-at-scale/
https://blogs.nvidia.com/blog/nvidia-vera-max-single-threaded-cpu-at-scale/
NVIDIA Blog
AI Innovators Adopt NVIDIA Vera — Why Max Single-Threaded CPU at Scale Matters
NVIDIA Vera exemplifies a new class of CPU, architected for the era of agents and being adopted by AI innovators including Perplexity; NVIDIA CPU roadmap continues with the NVIDIA Rosa CPU and its Rigel core.
NVIDIA Vera CPU Boosts AI Factory Throughput to Accelerate Agentic Workloads
https://developer.nvidia.com/blog/nvidia-vera-cpu-boosts-ai-factory-throughput-to-accelerate-agentic-workloads/
https://developer.nvidia.com/blog/nvidia-vera-cpu-boosts-ai-factory-throughput-to-accelerate-agentic-workloads/
NVIDIA Technical Blog
NVIDIA Vera CPU Boosts AI Factory Throughput to Accelerate Agentic Workloads
Agentic systems turn model reasoning into action through multi-step workflows that combine inference, tool use, code execution, retrieval, orchestration, and result handling. As these systems scale…
Maximize Spectral Efficiency with AI-Native RAN and NVIDIA AI Aerial
https://developer.nvidia.com/blog/maximize-spectral-efficiency-with-ai-native-ran-and-nvidia-ai-aerial/
https://developer.nvidia.com/blog/maximize-spectral-efficiency-with-ai-native-ran-and-nvidia-ai-aerial/
NVIDIA Technical Blog
Maximize Spectral Efficiency with AI-Native RAN and NVIDIA AI Aerial
Spectrum is one of the most valuable assets in wireless communications. Over the last 30 years, telecom operators in the US have spent more than $240B to acquire wireless spectrum. A goal of a radio…
Building an Analysis AI Agent for Industrial Alarm Management with NVIDIA Nemotron
https://developer.nvidia.com/blog/building-an-analysis-ai-agent-for-industrial-alarm-management-with-nvidia-nemotron/
https://developer.nvidia.com/blog/building-an-analysis-ai-agent-for-industrial-alarm-management-with-nvidia-nemotron/
NVIDIA Technical Blog
Building an Analysis AI Agent for Industrial Alarm Management with NVIDIA Nemotron
Industrial machinery generates more alarms than technicians can triage. For each important alarm requiring follow-up, the technician pulls historical context, determines the correct procedure…
Develop Humanoid Robot Policies End-to-End with NVIDIA Isaac GR00T
https://developer.nvidia.com/blog/develop-humanoid-robot-policies-end-to-end-with-nvidia-isaac-gr00t/
https://developer.nvidia.com/blog/develop-humanoid-robot-policies-end-to-end-with-nvidia-isaac-gr00t/
NVIDIA Technical Blog
Develop Humanoid Robot Policies End-to-End with NVIDIA Isaac GR00T
As more teams move from humanoid robot bring-up to task-specific skill development, the need for repeatable development workflows is growing. Building humanoids remains complex…
👍31
NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness
https://blogs.nvidia.com/blog/nemotron-langchain-agents-open-stack/
https://blogs.nvidia.com/blog/nemotron-langchain-agents-open-stack/
NVIDIA Blog
NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness
NVIDIA Nemotron 3 Ultra is offering leading performance at lower cost than top closed models with the largest and most widely adopted AI agent orchestration platform. LangChain tuned its Deep Agents harness for NVIDIA Nemotron 3 Ultra, achieving the highest…
Create a LangChain Deep Agents Harness Profile for NVIDIA Nemotron 3 Ultra to Improve Performance
https://developer.nvidia.com/blog/create-a-langchain-deep-agents-harness-profile-for-nvidia-nemotron-3-ultra-to-improve-performance/
https://developer.nvidia.com/blog/create-a-langchain-deep-agents-harness-profile-for-nvidia-nemotron-3-ultra-to-improve-performance/
NVIDIA Technical Blog
Create a LangChain Deep Agents Harness Profile for NVIDIA Nemotron 3 Ultra to Improve Performance
Agentic systems often face a trade-off between accuracy and cost. The highest-performing proprietary frontier models and harnesses provide top accuracy but are expensive. Fine-tuning offers one way to…
👍24
Running Low-Latency Analytical Workloads with GPU-Accelerated Presto on NVIDIA GB200 NVL72
https://developer.nvidia.com/blog/running-low-latency-analytical-workloads-with-gpu-accelerated-presto-on-nvidia-gb200-nvl72/
https://developer.nvidia.com/blog/running-low-latency-analytical-workloads-with-gpu-accelerated-presto-on-nvidia-gb200-nvl72/
NVIDIA Technical Blog
Running Low-Latency Analytical Workloads with GPU-Accelerated Presto on NVIDIA GB200 NVL72
Presto is an open source, distributed SQL engine for running fast, interactive queries on very large datasets. On NVIDIA GPUs, Presto delivers peak performance for analytical query workloads and…
👍25
GeForce NOW Turns Up the Heat With New GeForce RTX 5080-Powered Toronto Server
https://blogs.nvidia.com/blog/geforce-now-thursday-toronto-expansion/
https://blogs.nvidia.com/blog/geforce-now-thursday-toronto-expansion/
NVIDIA Blog
GeForce NOW Turns Up the Heat With New GeForce RTX 5080-Powered Toronto Server
Plus, tap into ‘NTE: Neverness to Everness’ with a new update, alongside three new games joining the cloud this week.
A Practical Guide to GPU-Initiated Communication for Molecular Dynamics at Scale
https://developer.nvidia.com/blog/a-practical-guide-to-gpu-initiated-communication-for-molecular-dynamics-at-scale/
https://developer.nvidia.com/blog/a-practical-guide-to-gpu-initiated-communication-for-molecular-dynamics-at-scale/
NVIDIA Technical Blog
A Practical Guide to GPU-Initiated Communication for Molecular Dynamics at Scale
Molecular dynamics (MD) simulations are among the most demanding workloads in computational science. Using them, researchers can observe atomic behavior in extraordinary detail…
Synthetic Data Generation for Financial AI Research with NVIDIA NeMo
https://developer.nvidia.com/blog/synthetic-data-generation-for-financial-ai-research-with-nvidia-nemo/
https://developer.nvidia.com/blog/synthetic-data-generation-for-financial-ai-research-with-nvidia-nemo/
NVIDIA Technical Blog
Synthetic Data Generation for Financial AI Research with NVIDIA NeMo
Fine-tuning LLMs for financial natural language processing (NLP) is constrained by limited, imbalanced data. Real-world financial news overrepresents earnings and stock movements…
Accelerating End-to-End Co-Folding Performance with NVIDIA BioNeMo Agent Toolkit
https://developer.nvidia.com/blog/accelerating-end-to-end-co-folding-performance-with-nvidia-bionemo-agent-toolkit/
https://developer.nvidia.com/blog/accelerating-end-to-end-co-folding-performance-with-nvidia-bionemo-agent-toolkit/
NVIDIA Technical Blog
Accelerating End-to-End Co-Folding Performance with NVIDIA BioNeMo Agent Toolkit
Biomolecular structure prediction and co-folding with models like OpenFold3 are now mainstream, large-scale workloads powering drug discovery and protein design. Increasingly, they’re driven end-to…
AI Model Co-Design: Hardware-Friendly LLM Design
https://developer.nvidia.com/blog/ai-model-co-design-hardware-friendly-llm-design/
https://developer.nvidia.com/blog/ai-model-co-design-hardware-friendly-llm-design/
NVIDIA Technical Blog
AI Model Co-Design: Hardware-Friendly LLM Design
AI performance comes down to three dimensions: Deployments must balance all three: High accuracy is wasted if responses are slow, and raw throughput means little if each user’s experience is laggy.
Kernel Fusion in NVIDIA CUDA: Optimizing Memory Traffic and Launch Overhead
https://developer.nvidia.com/blog/kernel-fusion-in-nvidia-cuda-optimizing-memory-traffic-and-launch-overhead/
https://developer.nvidia.com/blog/kernel-fusion-in-nvidia-cuda-optimizing-memory-traffic-and-launch-overhead/
NVIDIA Technical Blog
Kernel Fusion in NVIDIA CUDA: Optimizing Memory Traffic and Launch Overhead
There are many ways to optimize code for GPUs. In this post, you’ll learn how kernel fusion can improve memory bandwidth and reduce kernel launch overhead, along with multiple ways to apply it in…
Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading
https://developer.nvidia.com/blog/reducing-high-bandwidth-memory-bottlenecks-in-jax-based-llm-training-with-host-offloading/
https://developer.nvidia.com/blog/reducing-high-bandwidth-memory-bottlenecks-in-jax-based-llm-training-with-host-offloading/
NVIDIA Technical Blog
Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading
Large language model (LLM) training workloads increasingly run into GPU memory limits before compute is fully used. Model weights, gradients, optimizer states, communication buffers…
How to Evaluate General-Purpose Robot Policies for Real-World Deployment
https://developer.nvidia.com/blog/how-to-evaluate-general-purpose-robot-policies-for-real-world-deployment/
https://developer.nvidia.com/blog/how-to-evaluate-general-purpose-robot-policies-for-real-world-deployment/
NVIDIA Technical Blog
How to Evaluate General-Purpose Robot Policies for Real-World Deployment
Robotics foundation models have made remarkable progress. Today’s best systems can follow natural language instructions to pick, place, sort, and manipulate a wide variety of objects.
Extreme Event Likelihoods with Guided Generative Models
https://developer.nvidia.com/blog/extreme-event-likelihoods-with-guided-generative-models/
https://developer.nvidia.com/blog/extreme-event-likelihoods-with-guided-generative-models/
NVIDIA Technical Blog
Extreme Event Likelihoods with Guided Generative Models
Across science, engineering, and finance, many of the most important risks come from low-likelihood, high-impact events. Estimating the probability of these events with brute-force Monte Carlo…
NVIDIA Ising Decoding Cuts Color Code Logical Error Rates by Over 300X
https://developer.nvidia.com/blog/nvidia-ising-decoding-cuts-color-code-logical-error-rates-by-over-300x/
https://developer.nvidia.com/blog/nvidia-ising-decoding-cuts-color-code-logical-error-rates-by-over-300x/
NVIDIA Technical Blog
NVIDIA Ising Decoding Cuts Color Code Logical Error Rates by Over 300x
Useful quantum computers will require fault tolerant logical operations. Researchers are actively exploring many different quantum error correction (QEC) codes to enable this…
Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency
https://blogs.nvidia.com/blog/performance-per-watt-ai-infrastructure-efficiency/
https://blogs.nvidia.com/blog/performance-per-watt-ai-infrastructure-efficiency/
NVIDIA Blog
Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency
From benchmark to production, NVIDIA Blackwell NVL72 delivers the highest performance per watt to maximize revenue and the lowest token cost to maximize profit margins.