SpAItial AI unveils Echo-2, a world model that generates 3D scenes from an image
SpAItial AI launches Echo-2, a world model generating persistent 3D scenes from a single image using 3D Gaussian Splatting for real-time digital twins.
SpAItial AI launches Echo-2, its new world model capable of producing complete 3D scenes from a single source image. Unlike video models whose outputs collapse as soon as you return to a point already filmed, the system generates a spatially persistent environment: the scene maintains its geometric coherence regardless of the viewing angle or camera movement. Rendering is performed in real time using 3D Gaussian Splatting (3DGS), with interactive camera control and visual quality that the team claims is state-of-the-art according to the world score benchmark. Furthermore, the outputs are physically grounded, which opens up the possibility of extracting downstream representations that can be directly utilized: 3D meshes, point clouds, or scenes in 3DGS format. SpAItial AI positions Echo-2 as a bridge between the real and the virtual, in both directions. From physical to virtual, the tool allows for capturing an existing environment from photos to produce an editable digital twin, with use cases in architecture, interior design, or renovation. From virtual to physical, it serves as a simulation sandbox for embodied AI, where robots can learn and anticipate their actions in realistic scenes before operating in the real world. The demo is available at spaitial.ai.