NVIDIA's research division has shared results from papers accepted at the International Conference on Robotics and Automation (ICRA)[1]. Eight of its 28 accepted papers focus on "sim-to-real" transfer, moving skills learned in simulation onto physical robots[1]. The work aims for robots that operate reliably outside the lab in unpredictable environments, and the company says this transfer technique is becoming the shared foundation for that shift[1].
Moving Beyond Scripted Automation
Robotics is shifting from automation that merely follows fixed steps toward systems that judge and act on their own as conditions change[1]. The papers span the full range of challenges robot developers face: controlling multiple arms at once, adapting across differently shaped bodies, grasping objects in clutter, performing precise assembly, and vision-language-action models that reason before they move[1].
The first theme is efficient control of multiple robotic arms. Traditional scheduling handles steps one at a time, but ScheduleStream runs computations on GPUs so multiple arms can plan their motions in parallel[1]. On hardware such as the NVIDIA Jetson edge AI platform, it delivers a 3x speedup in multi-arm planning[1].
For navigation, a gait or route-finding skill learned on one body often fails on a differently shaped robot[1]. COMPASS first builds baseline behavior with imitation learning, then tunes it for diverse bodies through reinforcement learning in the NVIDIA Isaac Lab simulation environment[1]. All training happens in simulation, with no real-world robot data used[1]. Compared with an imitation-learning baseline, success rates improved 4.5x, and real-world navigation succeeded in about 80% of 20 trials[1].
For grasping, the last few centimeters are where small errors matter most[1]. Rather than executing a fixed plan, Grasp-MPC continuously corrects the robot's motion as it closes in on the object[1]. Trained on 2 million simulated trajectories across 8,000 objects, it learned to grasp novel objects on cluttered shelves and tabletops, reaching about 75% success on real robots, compared with 41% for the baseline[1]. Deformable Cluster Manipulation, meanwhile, showed how to clear shapeless bundles, such as tree branches tangled over a power line, using the whole arm rather than just the gripper[1].
Tackling Precise Assembly
Precise assembly, such as threading a nut onto a bolt or inserting a peg into a hole, is hard to reproduce in simulation alone[1]. SPARR uses a two-stage approach: a policy learns the general strategy in Isaac Lab, while on real hardware a second layer corrects for the gap between simulation and reality using the robot's camera[1]. This improved success rates by 38% and cut cycle time by about 30% compared with zero-shot sim-to-real methods[1]. On National Institute of Standards and Technology (NIST) assembly tasks not seen during training, success improved by nearly 75%[1].
Refinery handles assembly with multiple sequential steps[1]. As in assembling furniture, where the result of one step determines whether the next is possible, it learns to finish each step in a state that sets up the next[1]. It achieves 91% success in simulation with comparable real-world results, and its policies can be chained for long sequences[1].
Action Models That Keep Their Word
Much of what a camera captures is noise irrelevant to the task[1]. PEEK has a vision-language-action model (VLA, an AI that derives actions from images and instructions) read the task instruction, then highlights only the objects that matter while fading out the rest, focusing the robot's line of sight where it should look[1]. Added to a policy trained purely in simulation, it improved real-world accuracy 41x; for large VLAs and smaller policies, gains ranged from 2x to 3.5x[1].
SEAL addresses a failure mode where the AI reasons correctly about what to do but then executes something different[1]. It generates several candidate action sequences, considers where each would lead, and picks the one that matches what it said it would do[1]. It can be added at runtime without retraining and delivered up to 15% accuracy gains over prior work[1].
Spreading to Universities and Datasets
Beyond papers, NVIDIA is expanding large open datasets for robotics research[1]. The NVIDIA Physical AI Dataset has surpassed 15 million downloads, and NVIDIA Isaac GR00T X Embodiment Sim is among the most-downloaded as well[1]. Research teams at Carnegie Mellon University (CMU), ETH Zurich, the Massachusetts Institute of Technology (MIT), the University of Texas at Austin, and others published roughly 50 papers using NVIDIA technologies[1].
The real world is full of factors simulation tends to ignore, from uneven surfaces to sensor error[1]. This research lays out concrete ways to close that gap, a step toward robots that work practically outside the lab.
Summary
Through its ICRA-accepted papers, NVIDIA Research showed progress in moving simulation-trained skills onto physical robots. ScheduleStream speeds up multi-arm coordination 3x, COMPASS transfers across robot bodies, Grasp-MPC grasps novel objects, SPARR and Refinery handle precise assembly, and PEEK and SEAL keep perception and action on track. Together with university collaborations and growing open datasets, the bridge from simulation to reality is becoming a shared foundation for robot development.
Source: https://blogs.nvidia.com/blog/icra-research-robotics-simulation-to-real-world/
