Eighteen years of screen recordings, and not a single annotated click

Induction Labs' Photon-1 model learns to use a computer by predicting the next frames of screen recordings. A simple and efficient approach.