Photoroom researchers David Bertoin, Roman Frigg, and Jon Almazán published findings on training text-to-image models from scratch, evaluating techniques including REPA, iREPA, Contrastive Flow Matching, JiT, Muon Optimizer, and Alchemist against a baseline Flow Matching setup. Published on Hugging Face, the article reports which interventions improved convergence and training efficiency. The team plans to release full training code and conduct a public "speedrun" in their next post.
No score is assigned. Sources and their independence are shown in the citation chain below.