From Interactive Worlds to Humanoids: Odyssey-3 Scales Up
The Odyssey-3 world model simulates interactive environments in real time. Its architecture serves as a physical foundation for robotics and vehicles.
Odyssey introduces Odyssey-3, its most advanced world model to date, capable of generating interactive environments in real time while also serving as a foundation for physical systems such as robotic arms, humanoids, or autonomous vehicles. The company describes it as a learned dynamic system designed to predict how objects move, interact, and evolve over time.
Unlike a traditional video model, Odyssey-3 does not merely seek to produce a visually coherent sequence. Its goal is to predict the future state of an environment based on what has just occurred and the actions performed within it. This distinction is central to the world model category: the system must maintain a sufficiently stable representation of the world to respond when a user, an agent, or a machine alters the situation. A Diffusion Transformer That Predicts the Next State of the World
Technically, Odyssey-3 is based on an autoregressive diffusion transformer. The model generates several successive video steps while conditioning each new prediction on previous observations, recent events, and the actions provided to it.
The training corpus combines several types of data: videos from the internet, event annotations precisely localized in time, gameplay captures associated with keyboard and mouse inputs, as well as simulations of rigid body interactions accompanied by descriptions and metadata.
This diversity serves to link what the model sees to what happens next and, when the information exists, to the action that caused that change.
Odyssey then uses teacher forcing and causal masking to train the model to extend a sequence from its previous observations. Finally, a post-training phase mixes distribution matching and adversarial distillation to produce a variant requiring only a few diffusion steps and fast enough to allow real-time interaction. Worlds Generated While They Are Explored
The version accessible in research preview can generate an environment from a prompt and then continue to evolve it while the user moves around inside or triggers an event.
Odyssey-3 supports first-person and third-person navigation, as well as independent camera movement. The system continuously reuses previous observations and new actions to predict what should appear next.
This continuity is one of the major challenges of world models. Producing a beautiful image from a prompt is relatively different from maintaining the layout of objects, their physical properties, and the consequences of an action over several steps.
This is also what makes Odyssey-3 closer to a learned simulator than a simple video generator. A Score of 66.1 on Physics-IQ Verified
Odyssey particularly highlights the physical performance of its model.
On the Physics-IQ Verified benchmark, Odyssey-3 Pro achieves 66.1 in video-to-video in the configuration published by the company. Physics-IQ evaluates a model's ability to extend real physical experiments covering fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics, among others.
The Pro variant also reaches 54.7 in image-to-video, according to the results shared by Odyssey.
However, the term "state of the art" must be qualified. The current dynamic ranking of Physics-IQ Verified places FLUX 3 [large] ahead of Odyssey-3 Pro on certain configurations incorporating custom prompts and a net gain metric. Nevertheless, Odyssey remains among the best models currently listed on this benchmark.
Another important detail: the cost figures used by Odyssey are partially estimated. For its own models, the company assumes a cost of $1 per MI355X GPU hour, excluding potential prompt rewriting fees. First in Three Categories of WorldMark According to Odyssey
Odyssey-3 is also evaluated on WorldMark, a public benchmark designed for interactive world models.
WorldMark uses standardized environments, the same trajectories, and a common control layer to compare different models on visual quality, control adherence, and world consistency. Its evaluation set includes 500 cases covering first-person and third-person, realistic or stylized scenes, and several difficulty levels.
In evaluations conducted by Odyssey using the benchmark's official captions, Odyssey-3 ranks first in three of the four categories: stylized first-person, realistic third-person, and stylized third-person. It remains third in the realistic first-person category behind Lyra 2.0 and AlayaWorld.
These results relate to specific properties of the generated worlds and are not sufficient, on their own, to measure the quality of a real-world robotic system. Odyssey explicitly points this out: adaptation to a machine must also be evaluated on the behaviors that are actually useful to that machine. The Same Backbone Behind Multiple Machines
The most interesting part of Odyssey-3 probably lies in its reuse as a common foundation for multiple physical systems.
Rather than completely retraining a model for each robot or vehicle, Odyssey keeps the world model backbone and then adds an action decoder or a specialized policy. This layer learns to convert the model's internal representations into commands understandable by the machine in question.
For a robotic arm, Odyssey indicates that it used only a few dozen hours of demonstrations. The system then reportedly showed recovery behaviors absent from the initial examples, such as reorienting a gripper after a missed grasp or retrieving a dropped object in an unusual position.
On humanoids, Flexion built control policies based on Odyssey-3. According to the trials presented by Odyssey, they resisted environmental changes better than several VLA baselines, particularly when the lighting was modified. A Car Trained with Twenty Hours of Driving
Odyssey also says it has adapted the model to driving on real roads in India.
The Odyssey-3 backbone remains frozen, and a driving policy is trained using only 20 hours of data. The system leverages the visual representations learned by the world model to predict waypoints in front of the vehicle and