← Back to the wire

Training mRNA Language Models Across 25 Species for $165

AchievementModelMar 31, 2026

OpenMed built an end-to-end protein AI pipeline covering structure prediction, sequence design, and codon optimization, training four production models across 25 species in 55 GPU-hours for $165. The pipeline uses ESMFold for protein folding and ProteinMPNN for sequence design. For codon optimization, CodonRoBERTa-large-v2 achieved a perplexity of 4.10 and a Spearman CAI correlation of 0.40, outperforming ModernBERT. Maziyar Panahi contributed to the work, with CodonJEPA listed as an upcoming model on the project roadmap.

Receipt № 5251 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

OpenMedCompanyESMFoldModelProteinMPNNModelCodonRoBERTaModelCodonJEPAModelMaziyar PanahiPersonHugging FaceCompany
Canonical: https://huggingface.co/blog/OpenMed/training-mrna-models-25-species