PROWL-2 teaches agents to detect errors in their own simulations.
Odyssey introduces PROWL-2, a multi-agent learning framework where agents improve inside a World Model while a separate process searches for and repairs errors in the simulation. Across StarCraft II and simulated robot coordination tasks, Odyssey reports relative gains of up to 91% over its baseline.
Learning inside a simulation that learns too
An agent trained inside a World Model can learn from situations that never happened. The problem begins when the simulation gets them wrong: the agent may improve its behavior against an inaccurate imagined future.
PROWL-2 organizes training around two simultaneous curricula. Agents work through scenarios ranked by their learning potential, while a dedicated developer agent explores the real environment to identify situations the World Model predicts incorrectly.
The first curriculum moves toward increasingly difficult problems as the agents improve. The second collects failures that can be used to repair the World Model. A gate between useful difficulty and hallucination
PROWL-2 introduces a mechanism called a fidelity gate to connect the two learning processes.
The system periodically takes the agents' highest-priority situations and compares trajectories imagined by the World Model with recorded real trajectories. When the deviation becomes too large, that scenario is temporarily removed from agent training.
Its corresponding real trajectory is instead used to improve the World Model. The scenario only returns to the agents' curriculum after repeated checks show that the simulation predicts it accurately again.
A weakness discovered by the agents can therefore become new training data for their simulated environment before the problem is presented to them again. Up to 91% relative improvement in StarCraft II
Odyssey evaluated the framework across all nine scenarios in the StarCraft Multi-Agent Challenge v2, where unit types and starting positions change to prevent agents from memorizing a fixed battle.
According to the team's results, PROWL-2 achieves the highest mean performance across all nine configurations and reaches up to a 91% relative gain over its baseline World Model learner.
Some of the qualitative evaluation settings show large differences. In Terran 5v5, PROWL-2 reaches an 87.5% win rate versus 7.5% for the baseline across 40 held-out rollouts. Terran 10v11 moves from 2.5% to 55%. In Zerg 10v10, the reported result is 90% versus 0%.
These figures come from Odyssey's evaluation of its own framework. Quadrupeds have to coordinate too
The second evaluation moves from StarCraft to the Multi-Agent Quadruped Environment, a simulated setting where multiple robots need to coordinate their movements.
In Gate-3, three quadrupeds have to pass through a narrow opening without blocking or colliding with one another. After 1.5 million training steps, Odyssey reports a 70.4% success rate for PROWL-2 versus 29.8% for the baseline.
Shepherd-Hard asks two quadrupeds to cooperatively herd a flock of nine sheep toward a target. Success rises from 7.3% to 28.2% in the reported evaluation. Improving the agent and its environment together
PROWL-2 does not replace the policy-learning algorithm or the World Model architecture. Instead, the framework determines which experiences should train the agents and which should first be used to repair the simulation.
Odyssey connects this mechanism to broader research into systems capable of accumulating increasingly difficult experience over time. The company sees this co-evolution between agents and World Models as a path toward open-ended learning and, eventually, superintelligence. The latter remains Odyssey's research thesis rather than a conclusion demonstrated by the PROWL-2 experiments.