NVIDIA has restated its case that open world models belong at the base of physical AI systems such as robots and autonomous vehicles[1]. At the center of that case is Cosmos 3, a model family shipped with open weights, and the expansion of the NVIDIA Cosmos Coalition into Japan is starting to show how developers actually put it to work.
Why World Models Are Needed
What physical AI needs is not recognition of appearances but prediction of consequences[1]. World models learn how physical environments behave, what may happen next and which actions make sense, then generate physically grounded world and action data or simulate future states.
This matters in the field because training data is hard to gather[1]. Collecting data at the scale physical AI requires is expensive and slow, and rare or long-tail events are difficult to reproduce safely and repeatedly. World models make it possible to vary weather, lighting, objects and trajectories across diverse environments, giving teams a starting point they can then adapt to their own robot and sensor setup.
That adaptation step is exactly where openness becomes a technical requirement[1]. A general model has never seen a specific team's robot, sensors or operating environment. Closing that gap requires access to the weights, a license that permits modification and the tooling for post-training. NVIDIA's Cosmos world foundation models are offered under the Linux Foundation's OpenMDW 1.1 license, so teams can post-train them on their own data and hardware.
In July, NVIDIA joined more than 200 companies and organizations in signing an open letter titled "Open Weights and American AI Leadership"[1]. Its argument is that AI leadership will be measured not by any single frontier model but by whether an open ecosystem reaches every sector.
Cosmos 3 Combines Vision Reasoning, World Generation and Action Prediction
Cosmos 3 is an open physical AI foundation model built on a mixture-of-transformers architecture, which combines multiple specialized components[1]. It brings vision reasoning, world generation and action prediction into a single model family, so developers can understand scenes, generate synthetic data, simulate future states and build world action models without assembling and maintaining a separate model for each capability.
The lineup splits into three tiers[1]. Cosmos 3 Super, at 64 billion parameters, targets high-fidelity world modeling. Cosmos 3 Nano, at 16 billion parameters, is aimed at efficient reasoning and post-training. Cosmos 3 Edge, at 4 billion parameters, handles on-device vision reasoning and robot policy deployment. Edge is light enough for edge GPUs and can be deployed across NVIDIA RTX GPUs, NVIDIA DGX systems and NVIDIA Jetson, including Jetson Thor.
NVIDIA also cites benchmark results[1]. Cosmos 3 ranks No. 1 on Artificial Analysis for open weights text-to-image and image-to-video generation, on PAI-Bench for world generation and in the image-to-video category of Physics-IQ. For robot policy it ranks No. 1 on RoboLab, and Cosmos 3 Super is the highest-ranked open model on VANTAGE-Bench for vision understanding.
Surrounding libraries fill in what a model alone cannot[1]. Omniverse libraries, part of the NVIDIA Agent Toolkit, provide prebuilt capabilities for assembling simulation-ready worlds, while OpenUSD supplies the open framework for composing, reusing and exchanging 3D data across digital twins, simulations and synthetic data generation. The goal is to cut the duplicated work that piles up every time assets, sensor configurations or environmental conditions change. Beyond Cosmos, NVIDIA's physical AI stack includes Isaac GR00T for robotics, Alpamayo for autonomous vehicles and Metropolis for vision AI.
Adopters, and the Cosmos Coalition's Move Into Japan
Adoption already spans industries[1]. In robotics, Doosan Robotics, LG Electronics, SAMSUNG and Skild AI are building on Cosmos. In autonomous vehicles, Li Auto, Xiaomi and Afari. For vision AI agents powering industrial AI and smart spaces, Centific, Fogsphere, Linker Vision, Milestone Systems and Yuan.
Japan's side of the story was laid out in a July 15 announcement[2]. AIRoA, Classmethod, Enactic, FANUC, Fujitsu, GROOVE X, Hitachi, Honda R&D, Kawasaki Heavy Industries, Kubota, Mitsui & Co., Mitsubishi Corp., Mujin, NEC, Preferred Networks, SoftBank Corp., Sony Group Corporation, Telexistence, TIER IV, TRON K.K., Turing and Yaskawa Electric intend to join the NVIDIA Cosmos Coalition. Members can use the platform's open models, data curation libraries, datasets and frameworks to test and optimize systems before deployment across factories, logistics networks, farms, construction sites, hospitals, roads and homes, shortening development cycles.
Concrete projects are underway[2]. Fujitsu is exploring physical AI business opportunities with FANUC, Yaskawa Electric and Kawasaki Heavy Industries, aiming to build a collaborative control platform that integrates NVIDIA's physical AI stack. NEC, Hitachi, OMRON and Preferred Networks are using Cosmos for world model and industrial AI R&D, and SoftBank Corp. is developing a physical AI platform built on Cosmos, Omniverse and Isaac Sim. Kubota is exploring applications in autonomous agriculture, and Telexistence in retail automation.
Cosmos 3 Edge is a 4-billion-parameter model built on NVIDIA Nemotron, and NVIDIA says the open Cosmos framework lets developers adapt it to specific robots, vehicles, sensors and environments in about a day[2].
Summary
Open world models are becoming a practical answer to the step physical AI development cannot avoid: fitting a general model to a specific set of hardware. Cosmos 3 pairs benchmark results with a tiered lineup and, together with Omniverse, extends that to the validation environment as well. With Japan's manufacturing and robotics companies signing on as a group, the next thing worth watching is how much development time actually gets cut.
[1] https://blogs.nvidia.com/blog/open-world-models-physical-ai/
