GWM Worlds 2 generates a world as you control it
Runway has introduced GWM Worlds 2, a research model that generates continuous interactive simulations in 720p at 24 frames per second, with sound and live controls.
A character crosses the desert, thrusts a spear, then stops to drink. Meanwhile, another user intensifies the campfire, turns the sunset into a starry night, and triggers a dust storm. The scene was not entirely generated beforehand: GWM Worlds 2 extends it as new commands arrive.
Presented by Runway as its latest General World Model, the system generates a continuous 720p video stream at 24 frames per second, accompanied by 48 kHz audio. Its purpose is no longer limited to creating a clip from an instruction. It attempts to maintain an audiovisual environment that can change while it is being generated.
The user begins by defining the setting, characters, visual style, atmosphere, and certain general rules. An initial image provides the visual reference. Once the session begins, new commands can be directed at a specific character or the entire scene, while camera movements are transmitted continuously.
A command can make a character walk, ask them to pick up an object, trigger a gesture, introduce a line of dialogue, or change the weather. Several actions can overlap. The model must then incorporate them into the existing stream without returning to the beginning or systematically interrupting playback.
Runway organizes this information in a format called WorldPrompt. It separates what should remain stable from what changes over time.
The persistent section includes an initial instruction, called the “genesis prompt,” and the session’s first frame. The instruction describes the environment’s layout, materials, lighting, ambient sound, available characters, and their attributes. It can also specify a camera perspective or conventions governing gravity, collisions, and character abilities.
These “physical rules” should not be confused with those of a conventional simulation engine. Runway does not describe a rigid-body system or an explicitly measurable geometric representation. The rules are written in natural language and interpreted by the model as it generates the images and sound. They guide the result without guaranteeing deterministic or scientifically accurate behavior.
The second part of WorldPrompt consists of a timestamped event stream. Each action has a start time, an end time, and a target. A line of dialogue is treated as an action, just like movement or an interaction with an object. The camera’s requested rotation and translation are also provided for every frame.
This structure allows the viewpoint and subject to be controlled separately. Someone can move a rider forward while rotating the camera around them, drive a vehicle from a first-person perspective, or fly through an environment without necessarily moving a character.
In the live demonstration, users generally do not type complete text commands while playing. Instructions are prepared beforehand and mapped to keyboard and mouse inputs. The W key might send a prompt stating that the character moves forward, while a mouse click could ask them to throw an object.
The controls therefore resemble those of a game, but each input triggers a description for the model to interpret. This helps explain why the response can be more flexible than a conventional animation while remaining less predictable. The same action will not necessarily produce exactly the same movement twice.
Runway describes three ways to use the system. In the first, every event is written before generation begins, much like a technical script. This mode is closer to filmmaking and directing. In the second, the stream pauses at certain moments so the user can make a decision, following a structure suited to interactive narratives.
The third mode operates in real time. The model continues generating the scene as commands arrive. This is the most demanding scenario because each action must be interpreted quickly enough to affect the next frames without creating a noticeable pause.
Not all the sequences presented by Runway were produced under the same conditions. The company states that some demonstrations were fully authored in advance to provide more detailed descriptions. Most were reportedly produced through its real-time interface, but this distinction means that not every example should be treated as a completely improvised session.
Runway also acknowledges that the pre-authored mode currently delivers better results. When a camera turns around, an instruction written beforehand can describe exactly what should appear in the newly revealed direction. During a live session, the model must invent that portion of the environment using only the persistent prompt and recent frames.
GWM Worlds 2 uses an autoregressive diffusion architecture for video and audio. At every step, it receives the global context, the commands active at that moment, camera instructions, and a portion of the previously generated frames. Its new output then becomes context for what follows.
Its visual memory is not unlimited. Recent frames are retained inside a sliding window, while older ones are progressively removed. The initial prompt and first frame remain available as global references, but the model does not continuously review the entire session.
This structure allows the stream to continue without determining its full duration in advance. It does not mean that the world has perfect memory or that a session can continue indefinitely without deteriorating. An object left behind several minutes earlier could change appearance, disappear, or return in a different location after the relevant visual references have left the recent context window.
Runway reports drift in details, textures, and geometry, particularly during fast camera movements. Long-term consistency remains imperfect, and the model cannot accept additional image references after the first frame. A visual identity introduced during a session therefore does not receive the same persistent grounding.
Sound is generated alongside the video rather than simply added afterward. An action can contain dialogue, a sound effect, or an