← Back to the wire

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

AnnouncementProductOct 1, 2026

Olmo-core 3 is an open training framework upgrade designed to scale mixture-of-experts models into the trillion-parameter range. In one benchmark, expanding the expert pool from 8 to 128 grew total parameters from 4.6B to 47B while throughput fell under 5%. On eight NVIDIA B300 GPUs, a 47-billion-parameter MoE reached 52,000 tokens per second per GPU, roughly 2.7× the earlier implementation. It has been benchmarked at over one trillion parameters.

Receipt № 21761 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01medPRIMARY
Olmo-core 3ModelOlmoModel
Canonical: https://huggingface.co/blog/allenai/olmocore3