Atlas reconstructs what it sees and imagines the rest
With Atlas, World Labs combines image and video generation, 3D reconstruction, and spatial simulation in a single World Model capable of following a precise camera trajectory.
A photograph normally shows only one point of view. Atlas attempts to turn it into the starting point for a complete space where a camera can move forward, pull back, rotate, or change height. Visible areas remain guided by the original image, while missing regions are reconstructed or imagined by the model.
Introduced by World Labs as a multimodal World Model, Atlas accepts text, images, videos, camera positions, and depth maps. These inputs are combined into a shared “spatial context” that tells the model where each view is located within the scene.
This camera geometry is one of its main differences from conventional video models. An instruction such as “pan right” describes a general intention. Atlas can instead receive a precise position, orientation, and trajectory, then generate the views corresponding to that path.
The user is therefore not simply requesting a camera effect. They define the path the camera should follow through the space. The model then generates the intermediate frames while attempting to preserve the arrangement of objects, volumes, and environmental continuity.
This method makes it possible to begin with a single image and observe the scene from another angle. Atlas must then invent everything the photograph does not contain: the back of a character, the other side of an object, a neighboring room, or the landscape outside the frame. The consistency of these elements depends on knowledge acquired during training, not on information actually present in the source image.
Additional references reduce the amount of invention required. Each photograph can be positioned within the spatial context to indicate which part of the location it represents. Two unrelated images can also be placed in different areas, with Atlas imagining a hallway, doorway, or another transition that connects them within a single environment.
Reconstructing a real location follows the same principle. With one view, only the visible area can be considered documented. The remaining regions may be plausible, but they are fictional. Adding more photographs provides further information about buildings, objects, and proportions. As World Labs summarizes it, the more the model sees, the less it has to imagine.
The company says two or three images may already be enough for some reconstructions. Atlas can also use more than 100 views when fidelity to a real location matters more than interpretation. This flexibility allows the same model to move from a creative exercise to a much more thoroughly documented capture.
Atlas does not produce only images or videos. It can generate depth maps and combine the resulting information as point clouds or 3D Gaussian splats. These representations provide an explicit 3D space that can be explored and integrated into design, gaming, visual effects, or robotics workflows.
Starting from a single photograph, the system generates multiple views, estimates their geometry, and combines the information into a 3D scene. From a video, it predicts the depth of each frame before merging them. In both cases, regions never observed by a camera are completed through generation. A reconstruction created from a small number of references should therefore not be mistaken for an exact survey of the location.
The same spatial context can support several trajectories through one environment. A filmmaker could create a slow initial movement, then reuse the same references to build an aerial path or a faster sequence. World Labs claims that the overall structure remains consistent as the camera follows these different routes.
The main demonstration follows a camera for one minute through a medieval village generated from several reference images. The video reaches 1440p resolution and follows a manually designed trajectory. Atlas does not independently choose the entire staging. The path is prepared first, and the model calculates what the camera should encounter.
The same approach can change the point of view of a real-world video. Using footage from three to five ordinary cameras placed around an action, Atlas reconstructs the scene and generates intermediate angles. The result resembles a bullet-time effect without an installation of dozens of synchronized cameras.
The examples were recorded with phones and action cameras attached to tripods or portable clamps. The system can freeze a moment, move the virtual viewpoint, and reconstruct the scene from an angle that was never directly recorded. The accuracy of obscured details still depends on the information available across the different recordings.
Robotics is the other central application. An environment can be captured with a short video, reconstructed in 3D, and then observed from the virtual position of a robot’s sensors. Atlas generates the visual and depth data the machine would receive while following a specified route.
In the navigation demonstrations, two large environments were reconstructed from 24 frames extracted from phone footage. Different virtual robots can then move through these locations using different positions and onboard cameras. The process is intended to expand the number of training environments without scanning every site with specialized equipment.
For manipulation tasks, World Labs combines Atlas with its Real-to-Sim workflow. A few recordings are used to reconstruct the environment, objects, and their interactions. Developers can then change object positions, lighting, robot movements, and parts of the setting to produce varied training situations.
Atlas should not, however, be presented as the sole physics engine behind these simulations. World Labs explains that its Real-to-Sim system combines several representations and techniques depending on the task. The World Model contributes to appearance, geometry, and sensor observations, while validating contacts, forces, and physical behavior requires a broader system.
Image generation is also supported. Atlas can create visuals and 360-degree panoramas from text or image