A Practical Guide to GPU-Initiated Communication for Molecular Dynamics at Scale
https://developer.nvidia.com/blog/a-practical-guide-to-gpu-initiated-communication-for-molecular-dynamics-at-scale/
https://developer.nvidia.com/blog/a-practical-guide-to-gpu-initiated-communication-for-molecular-dynamics-at-scale/
NVIDIA Technical Blog
A Practical Guide to GPU-Initiated Communication for Molecular Dynamics at Scale
Molecular dynamics (MD) simulations are among the most demanding workloads in computational science. Using them, researchers can observe atomic behavior in extraordinary detail…
Synthetic Data Generation for Financial AI Research with NVIDIA NeMo
https://developer.nvidia.com/blog/synthetic-data-generation-for-financial-ai-research-with-nvidia-nemo/
https://developer.nvidia.com/blog/synthetic-data-generation-for-financial-ai-research-with-nvidia-nemo/
NVIDIA Technical Blog
Synthetic Data Generation for Financial AI Research with NVIDIA NeMo
Fine-tuning LLMs for financial natural language processing (NLP) is constrained by limited, imbalanced data. Real-world financial news overrepresents earnings and stock movements…
Accelerating End-to-End Co-Folding Performance with NVIDIA BioNeMo Agent Toolkit
https://developer.nvidia.com/blog/accelerating-end-to-end-co-folding-performance-with-nvidia-bionemo-agent-toolkit/
https://developer.nvidia.com/blog/accelerating-end-to-end-co-folding-performance-with-nvidia-bionemo-agent-toolkit/
NVIDIA Technical Blog
Accelerating End-to-End Co-Folding Performance with NVIDIA BioNeMo Agent Toolkit
Biomolecular structure prediction and co-folding with models like OpenFold3 are now mainstream, large-scale workloads powering drug discovery and protein design. Increasingly, they’re driven end-to…
AI Model Co-Design: Hardware-Friendly LLM Design
https://developer.nvidia.com/blog/ai-model-co-design-hardware-friendly-llm-design/
https://developer.nvidia.com/blog/ai-model-co-design-hardware-friendly-llm-design/
NVIDIA Technical Blog
AI Model Co-Design: Hardware-Friendly LLM Design
AI performance comes down to three dimensions: Deployments must balance all three: High accuracy is wasted if responses are slow, and raw throughput means little if each user’s experience is laggy.
Kernel Fusion in NVIDIA CUDA: Optimizing Memory Traffic and Launch Overhead
https://developer.nvidia.com/blog/kernel-fusion-in-nvidia-cuda-optimizing-memory-traffic-and-launch-overhead/
https://developer.nvidia.com/blog/kernel-fusion-in-nvidia-cuda-optimizing-memory-traffic-and-launch-overhead/
NVIDIA Technical Blog
Kernel Fusion in NVIDIA CUDA: Optimizing Memory Traffic and Launch Overhead
There are many ways to optimize code for GPUs. In this post, you’ll learn how kernel fusion can improve memory bandwidth and reduce kernel launch overhead, along with multiple ways to apply it in…
Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading
https://developer.nvidia.com/blog/reducing-high-bandwidth-memory-bottlenecks-in-jax-based-llm-training-with-host-offloading/
https://developer.nvidia.com/blog/reducing-high-bandwidth-memory-bottlenecks-in-jax-based-llm-training-with-host-offloading/
NVIDIA Technical Blog
Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading
Large language model (LLM) training workloads increasingly run into GPU memory limits before compute is fully used. Model weights, gradients, optimizer states, communication buffers…
How to Evaluate General-Purpose Robot Policies for Real-World Deployment
https://developer.nvidia.com/blog/how-to-evaluate-general-purpose-robot-policies-for-real-world-deployment/
https://developer.nvidia.com/blog/how-to-evaluate-general-purpose-robot-policies-for-real-world-deployment/
NVIDIA Technical Blog
How to Evaluate General-Purpose Robot Policies for Real-World Deployment
Robotics foundation models have made remarkable progress. Today’s best systems can follow natural language instructions to pick, place, sort, and manipulate a wide variety of objects.
Extreme Event Likelihoods with Guided Generative Models
https://developer.nvidia.com/blog/extreme-event-likelihoods-with-guided-generative-models/
https://developer.nvidia.com/blog/extreme-event-likelihoods-with-guided-generative-models/
NVIDIA Technical Blog
Extreme Event Likelihoods with Guided Generative Models
Across science, engineering, and finance, many of the most important risks come from low-likelihood, high-impact events. Estimating the probability of these events with brute-force Monte Carlo…
NVIDIA Ising Decoding Cuts Color Code Logical Error Rates by Over 300X
https://developer.nvidia.com/blog/nvidia-ising-decoding-cuts-color-code-logical-error-rates-by-over-300x/
https://developer.nvidia.com/blog/nvidia-ising-decoding-cuts-color-code-logical-error-rates-by-over-300x/
NVIDIA Technical Blog
NVIDIA Ising Decoding Cuts Color Code Logical Error Rates by Over 300x
Useful quantum computers will require fault tolerant logical operations. Researchers are actively exploring many different quantum error correction (QEC) codes to enable this…
Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency
https://blogs.nvidia.com/blog/performance-per-watt-ai-infrastructure-efficiency/
https://blogs.nvidia.com/blog/performance-per-watt-ai-infrastructure-efficiency/
NVIDIA Blog
Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency
From benchmark to production, NVIDIA Blackwell NVL72 delivers the highest performance per watt to maximize revenue and the lowest token cost to maximize profit margins.
Nemotron Labs: How Open Models Give Enterprises and Nations AI They Can Trust, Control and Customize
https://blogs.nvidia.com/blog/nemotron-open-models-ai-trust-control-customize/
https://blogs.nvidia.com/blog/nemotron-open-models-ai-trust-control-customize/
NVIDIA Blog
Nemotron Labs: How Open Models Give Enterprises and Nations AI They Can Trust, Control and Customize
Learn how NVIDIA Nemotron open models help enterprises build specialized AI they can trust, control and customize with accuracy, efficiency and flexibility.
How to Run an Autoresearch Workflow with RL Agent Skills and NVIDIA NeMo
https://developer.nvidia.com/blog/how-to-run-an-autoresearch-workflow-with-rl-agent-skills-and-nvidia-nemo/
https://developer.nvidia.com/blog/how-to-run-an-autoresearch-workflow-with-rl-agent-skills-and-nvidia-nemo/
NVIDIA Technical Blog
How to Run an Autoresearch Workflow with RL Agent Skills and NVIDIA NeMo
Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes, resolve build issues, launch experiments…
Post-Train NVIDIA Cosmos 3 in One Day Using Agent Skills
https://developer.nvidia.com/blog/post-train-nvidia-cosmos-3-in-one-day-using-agent-skills/
https://developer.nvidia.com/blog/post-train-nvidia-cosmos-3-in-one-day-using-agent-skills/
NVIDIA Technical Blog
Post-Train NVIDIA Cosmos 3 in One Day Using Agent Skills
What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning models to production video tasks…
Lessons From the Leaderboard: What 5,000+ Kagglers Taught Us About Improving AI Reasoning
https://developer.nvidia.com/blog/lessons-from-the-leaderboard-what-5000-kagglers-taught-us-about-improving-ai-reasoning/
https://developer.nvidia.com/blog/lessons-from-the-leaderboard-what-5000-kagglers-taught-us-about-improving-ai-reasoning/
NVIDIA Technical Blog
Lessons From the Leaderboard: What 5,000+ Kagglers Taught Us About Improving AI Reasoning
The NVIDIA Nemotron Model Reasoning Challenge invited the Kaggle community to explore a focused question: What techniques can improve reasoning accuracy when everyone starts from the same open model…
NVIDIA and Japan Bring Full-Stack AI and Robotics to Every Industry
https://blogs.nvidia.com/blog/japan-ecosystem-2026/
https://blogs.nvidia.com/blog/japan-ecosystem-2026/
NVIDIA Blog
NVIDIA and Japan Bring Full-Stack AI and Robotics to Every Industry
NVIDIA and its partners in Japan are this week showcasing the AI ecosystem's latest advancements. Check back here for updates.
NVIDIA Introduces New Jetson Thor Computers to Advance Mainstream Robotics and Edge AI
https://blogs.nvidia.com/blog/jetson-thor-robotics-edge-ai-agent/
https://blogs.nvidia.com/blog/jetson-thor-robotics-edge-ai-agent/
NVIDIA Blog
NVIDIA Introduces New Jetson Thor Computers to Advance Mainstream Robotics and Edge AI
New NVIDIA Blackwell-powered T3000 and T2000 modules, paired with new NVIDIA Jetson software memory optimization and agent skills, help partners and customers move advanced robotics, visual AI and edge workloads onto compact, power-efficient systems.
Building Faster Cryptography with Carryless Multiplication in NVIDIA CUDA 13.3
https://developer.nvidia.com/blog/building-faster-cryptography-with-carryless-multiplication-in-nvidia-cuda-13-3/
https://developer.nvidia.com/blog/building-faster-cryptography-with-carryless-multiplication-in-nvidia-cuda-13-3/
NVIDIA Technical Blog
Building Faster Cryptography with Carryless Multiplication in NVIDIA CUDA 13.3
For over fifteen years, x86 CPUs have shipped with a dedicated hardware instruction for carryless multiplication. It’s a small but stubborn primitive that sits underneath authenticated encryption…
Develop Lightweight USD Runtimes Faster with AI Agents
https://developer.nvidia.com/blog/develop-lightweight-usd-runtimes-faster-with-ai-agents/
https://developer.nvidia.com/blog/develop-lightweight-usd-runtimes-faster-with-ai-agents/
NVIDIA Technical Blog
Develop Lightweight USD Runtimes Faster with AI Agents
OpenUSD is an open, extensible framework that provides a common scene description language for physical AI. It enables teams to bring CAD data, simulation assets, and real-world telemetry into a…
👍13
Build a Multi-Camera 3D Tracking Application with NVIDIA DeepStream 9.1 Skills
https://developer.nvidia.com/blog/build-a-multi-camera-3d-tracking-application-with-nvidia-deepstream-9-1-skills/
https://developer.nvidia.com/blog/build-a-multi-camera-3d-tracking-application-with-nvidia-deepstream-9-1-skills/
NVIDIA Technical Blog
Build a Multi-Camera 3D Tracking Application with NVIDIA DeepStream 9.1 Skills
Developers building video analytics applications across large spaces must track the same object as it moves between camera views. Single-camera 2D tracking lacks reliable depth information and…
👍7
Sharpen the Sword, Skip the Downloads — ‘Onimusha: Way of the Sword’ Is Coming to GeForce NOW
https://blogs.nvidia.com/blog/geforce-now-thursday-onimusha-coming/
https://blogs.nvidia.com/blog/geforce-now-thursday-onimusha-coming/
NVIDIA Blog
Sharpen the Sword, Skip the Downloads — ‘Onimusha: Way of the Sword’ Is Coming to GeForce NOW
Plus, GeForce NOW launches in India moving from beta to public availability, and delivers ‘Denshattack!’ alongside five new games this week.
👍20
Scaling Agentic AI Factories Through Extreme Co-Design with NVIDIA BlueField
https://developer.nvidia.com/blog/scaling-agentic-ai-factories-through-extreme-co-design-with-nvidia-bluefield/
https://developer.nvidia.com/blog/scaling-agentic-ai-factories-through-extreme-co-design-with-nvidia-bluefield/
NVIDIA Technical Blog
Scaling Agentic AI Factories Through Extreme Co-Design with NVIDIA BlueField
Agentic AI changes the infrastructure pattern for AI factories. One request can trigger many model calls, tool calls, memory lookups, policy checks, storage accesses, and network transfers before a…
👍7