At the computer vision conference CVPR 2026, NVIDIA presented three research papers on physical AI (AI for machines that operate in the real world)[1]. The highlights are a foundation model that lets robots grasp objects even with tools they have never held, a self-driving model that decides quickly on in-car hardware, and an agent AI trained across more than 1,000 games[1]. All three share a single idea: train at scale, and the system will generalize to situations it has never seen[1].

A Shared Theme: Generalization Through Large-Scale Training

CVPR (Computer Vision and Pattern Recognition) is one of the largest international conferences in computer vision, held this year from June 3 to 7 in Denver, USA[1]. The three papers from NVIDIA Research address separate challenges — robotic grasping, autonomous driving, and training virtual agents — yet share the same approach: by learning from as many diverse situations as possible at scale, you can build systems that handle scenarios they have never encountered[1].

GraspGen-X: The First Foundation Model for Grasping with Any Gripper

Most grasping AI systems are specialists[1]. A policy trained for a two-finger gripper (the mechanism that serves as a robot's hand) can only grasp with those two fingers, and every new hand design requires repeating data collection, fine-tuning, and validation[1]. As a result, many robotics companies pick a single hand design and build around it[1].

GraspGen-X is the first foundation model for grasping built to remove this constraint[1]. Just as a large language model can apply its understanding of language to a new task without retraining, GraspGen-X applies its understanding of geometry and contact to grippers it has never encountered[1]. Given the geometry of a new gripper and an object it has never seen, it generates candidate grasp poses for picking the object up reliably[1].

Training it required data that cannot be collected at scale in the real world[1]. The researchers generated 2 billion simulated grasps spanning thousands of object shapes and synthetic gripper configurations, covering the diversity of forms a deployed robot might encounter[1]. For developers, this removes the need to retrain for each hand design, and it works out of the box for several commonly used grippers[1]. Combined with curoboV2, a newly released CUDA-accelerated motion planning library, GraspGen-X can achieve target grasp poses even in unknown environments[1]. A related paper, Grasp-MPC, presented at ICRA 2026, advances the next step from grasp generation to closed-loop grasp execution[1].

LCDrive: A Self-Driving Model That Thinks Faster on In-Car Hardware

Recent research has shown that letting an AI work through its reasoning (the process by which a trained AI arrives at an answer) before committing to an answer reliably improves the quality of its decisions[1]. For autonomous vehicles, the challenge is running that reasoning on the hardware actually inside the car[1]. With text-based reasoning, every word becomes a token (the smallest unit of text an AI processes) and takes time to generate[1]. On a processor inside a car, this token count is a real constraint on response speed[1].

LCDrive addresses this by replacing words with compressed latent representations (internally compressed representations)[1]. Instead of writing out human-readable reasoning steps, it thinks within a compressed space that captures spatial information[1]. The architecture alternates between two kinds of thinking — proposing candidate actions, then predicting how the world will change if those actions are taken — and refines its next move based on the predicted state[1]. It runs the same reasoning loop in a far more computationally efficient form than natural language[1]. As a result, it achieves trajectory quality comparable to text-based reasoning using roughly half the tokens[1]. The model is built on NVIDIA Alpamayo and trained with supervision derived from existing vehicle data[1].

NitroGen: An Embodied AI Agent Trained on More Than 1,000 Games

NVIDIA's open foundation model for humanoid robots, Isaac GR00T, is built on a simple principle: expose a model to enough diverse situations, and it will generalize to ones it has never seen[1]. NitroGen extends that principle to virtual environments, using the GR00T architecture to train a foundation model for embodied AI agents (AI that learns to act in the real world) across a wide range of virtual worlds[1].

Video games provide structured, diverse worlds at scale, complete with clear goals and success conditions[1]. NitroGen treats games as a training ground for agents that will eventually handle new situations in the real world or in other simulations[1]. The idea points toward, for example, building the foundation for a robot that helps with chores based on a broad instruction like "put these items away in the pantry"[1]. Training a GR00T-based model across more than 1,000 games and 40,000 hours of interaction produced agents that generalize across environments[1]. They were evaluated across a wide range of games — action RPGs, platformers, roguelikes, and open-world titles — and demonstrated behaviors such as combat, navigation, and exploration[1].

The same techniques could help enable more adaptive NPCs (non-player characters), AI companions, and gameplay systems within games, as well as broader testing of complex game environments[1]. In low-data conditions, where an agent has seen only a handful of examples of a new environment, starting from NitroGen gives the agent a major advantage, improving performance by up to 52 percent over previous state-of-the-art methods[1]. The model is open source and available on GitHub and Hugging Face[1].

New Physical AI Skills to Accelerate Research

At CVPR, alongside these three papers, NVIDIA also unveiled new physical AI agent skills that accelerate the development of autonomous vehicles, robots, and vision AI systems[1]. The aim is to make it easier for researchers and developers to connect a model's capabilities into end-to-end workflows that can be deployed in the real world[1].

Summary

These three papers stand out for binding together separate challenges — grasping, autonomous driving, and virtual agents — under one shared idea: acquiring generality through large-scale training. GraspGen-X, which works with any gripper; LCDrive, which decides quickly on in-car hardware; and NitroGen, trained on more than 1,000 games, all emphasize applicability that is not tied to a single use case. In particular, NitroGen is open source and published on GitHub and elsewhere, so robotics and game AI developers can try it in their own environments — something likely to broaden the reach of this research[1].

Source: https://blogs.nvidia.com/blog/cvpr-research-grasping-driving-agent-training/