NVIDIA has launched Cosmos 3, described as the world's first fully open omnimodel — a world foundation model for physical AI that can understand and generate text, images, video, ambient sound and actions with leading physics accuracy. Built on a mixture-of-transformers architecture, Cosmos 3 combines vision reasoning, world generation and action prediction in a single unified model designed to power the next generation of robots, autonomous vehicles and vision AI agents.
What Makes Cosmos 3 an Omnimodel
NVIDIA describes Cosmos 3 as the world's first fully open omnimodel — a model capable of understanding and generating across five modalities: text, images, video, ambient sound and actions. This positions it well beyond a conventional language or vision model, giving it the breadth needed to perceive and simulate the physical world end to end.
Its mixture-of-transformers architecture pairs a reasoning transformer with an expert generation transformer. This pairing allows the model to understand object interactions, motion and spatial-temporal relationships before generating video outputs and action trajectories — capabilities critical for training real-world autonomous systems.
Physical AI Capabilities and Platform Expansion
NVIDIA says Cosmos 3 allows robots, autonomous vehicles and vision agents to operate in real-world conditions even with limited training data and fragmented simulation stacks. The platform now includes new datasets covering robotics, physics, human motion, autonomous driving, warehouse safety and spatial reasoning.
New physical AI agent skills have also been added for neural scene reconstruction, defect-image generation and video augmentation — tools aimed at supporting researchers across the full data generation, simulation and evaluation pipeline for autonomous systems development.
Cosmos 3 Super, part of NVIDIA's post-training lineup, is positioned for applications requiring the highest physics accuracy and generation quality — making it the flagship option for robotics and autonomous vehicle developers with the most demanding use cases.
"The big bang of physical AI is just around the corner thanks to breakthroughs in multimodal reasoning language, vision and world models. The Cosmos 3 family of open, frontier omnimodels gives developers a generational leap in ability to build robots, AVs and vision AI that perceive, reason, plan and act in the physical world."— Jensen Huang, Founder & CEO, NVIDIA
How Developers Are Using Cosmos 3
Developers can use Cosmos 3 in multiple ways: as a vision language model, as the backbone for world action models, or as a world and video foundation model that simulates physical environments and predicts future world states for training and evaluation. For robotics tasks, the model can generate synthetic data and scene variations and support post-training with embodiment-specific behaviour — from pick-and-place operations to dexterous manipulation.
Robotics adopters already building on the platform include Agile Robots, Doosan Robotics, LG Electronics, Samsung Electronics and Skild AI. Li Auto is using the platform for autonomous vehicle development, while Centific, Fogsphere, Linker Vision, Milestone Systems and Yuan are applying it to vision AI agents for industrial AI and smart space applications.
