← Back to the wire

PRX Part 3 — Training a Text-to-Image Model in 24h!

AchievementResearchMar 4, 2026

Photoroom researchers David Bertoin, Roman Frigg, and Jon Almazán trained a text-to-image model in 24 hours by combining multiple optimization techniques. The team used x-prediction in pixel space to eliminate the need for a VAE, applied TREAD token routing to reduce per-step compute, and employed REPA with DINOv3 for representation alignment. They open-sourced the code for reproduction. The experiment demonstrates how far careful engineering can advance performance under strict compute budgets.

Receipt № 5461 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

PhotoroomCompanyDavid BertoinPersonRoman FriggPersonJon AlmazánPerson
Canonical: https://huggingface.co/blog/Photoroom/prx-part3