The Kinetix team is joining Runway to bring its World Models to life

Kinetix is joining Runway's Paris-based team, a specialist in 3D human motion, to strengthen its World Models for robotics.

Making a character walk without its feet sliding, transferring a dance to another body, or keeping a hand resting on an object during a movement. These animation problems now lead to a broader question: how can a machine learn to act in an environment it has never encountered?

Runway has announced that the Kinetix team is joining its Paris research hub. The French lab brings six years of work on 3D human motion, video generation controlled by geometric data, and datasets describing contact between bodies and their surroundings.

The announcement does not explicitly state that Runway has acquired Kinetix. It only says that the team is joining the company. No purchase price, share exchange, or transfer of intellectual property has been disclosed. The future of the French company, its services, and its existing contracts also remains unclear.

The deal therefore looks more like a team integration than a conventional, fully documented acquisition. Runway is bringing in an established group, its expertise, and presumably some of its research, but has not said whether the Kinetix legal entity will be absorbed or remain active.

Founded in Paris in 2020 by Yassine Tahi and Henri Mirande, Kinetix first became known for a tool that could convert video into 3D animation. A person could record a movement and transfer it to an avatar without a specialized motion-capture studio or manual frame-by-frame work.

The startup raised $11 million in 2022 in a seed round led by Adam Ghobarah, founder of Top Harvest Capital. The Sandbox, Zepeto, Sparkle Ventures, Entrepreneur First, Xavier Niel, and several entrepreneurs also participated.

At the time, its messaging was closely tied to the metaverse and user-generated content. Kinetix wanted to give players a way to create their own gestures, dances, and expressions for virtual characters. The gradual disappearance of the word “metaverse” did not eliminate the technical problem the company was working on: understanding how a body moves through space.

Kinetix later shifted its focus toward human-motion research, 3D animation, and controllable generative video. The company says its technologies have been integrated into Adobe Mixamo, Unity’s Muse services, and Krafton’s Overdare platform, where players can create custom animations.

Its core expertise lies in motion conversion and transfer. An animation recorded from one person or character cannot simply be copied onto a different body. Limbs vary in length, joints differ, and torso volume changes. An inaccurate transfer can create floating feet, arms passing through the body, or hands losing contact with an object.

Kinetix developed a motion retargeting system that accounts for the target character’s shape. Instead of processing joint positions alone, it considers body geometry and important contact points. A hand resting on a hip should remain there even if the new character has broader shoulders or shorter arms.

This capability is directly relevant to animation, but using it in robotics requires another step. Transferring a human gesture to a robot involves more than adapting one body shape to another. The system must account for the number of joints, their range of motion, balance, motor speed, applied force, and the constraints of each machine.

Kinetix had already begun connecting its research to this field. Its team built a synthetic data pipeline using hundreds of hours of motion-capture recordings. More than 75 people with different body types and physical backgrounds took part, performing movements ranging from everyday walking to actions carried out by athletes, dancers, and stunt performers.

The sequences were annotated with action types, ground contacts, object interactions, and body positions. Kinetix then transferred them to more than 1,000 virtual characters and placed them in environments created with Unreal Engine. Lighting, camera position, surroundings, and body type could be varied to produce millions of paired examples.

This type of pipeline creates a known correspondence between an observed video and the 3D movement that produced it. The model no longer has to infer depth, limb trajectories, or camera position entirely from a flat image. That information already exists in the training data.

Kinetix applied this principle to Kamo-1, its 3D-conditioned video-generation model. The system receives a body animation, camera path, depth map, reference image, and text instruction. It then generates a video while attempting to preserve the character’s movement and the camera’s motion independently.

This separation addresses a common weakness in generative video. When a camera moves around a person, the system may confuse the viewpoint change with the subject’s movement. The body can veer off course, slide, or change proportions. Three-dimensional coordinates provide a more stable reference for distinguishing those movements.

Kamo-1 does not yet maintain a complete and persistent representation of the entire scene. Kinetix acknowledges that the character and camera receive more precise 3D guidance than the setting, surfaces, and objects. The model generates video guided by geometry, rather than a fully reconstructed environment that can be inspected from any angle.

The company also states that 3D data alone does not teach the laws of physics. It describes position, depth, and spatial relationships, but does not guarantee that the system understands mass, friction, inertia, or the consequences of a collision. A plausible-looking video is not automatically an accurate simulation.

That distinction matters for the team’s new role. Runway argues that models pretrained on large video collections can help robots generalize their behavior. By observing many environments, objects, and actions, they could learn visual patterns that remain useful when facing situations missing from their robotics data.

Moving from video to action remains difficult. A video model predicts the visible evolution of a scene. A robot must choose a