Isaac 0.5's developers report that scaling general video from 1,000 to one million hours reduced the teleoperation data needed to reach a held-out action loss target from about 5,900 hours to 28. The 36-billion-parameter sparse model processes images, video, language, robot state, and actions, and was trained on data from over 35 robot systems and 3T multimodal tokens. Checkpoints, training code, and inference code are released via LeRobot.
No score is assigned. Sources and their independence are shown in the citation chain below.