Olmo-core 3 is an open training framework upgrade designed to scale mixture-of-experts models into the trillion-parameter range. In one benchmark, expanding the expert pool from 8 to 128 grew total parameters from 4.6B to 47B while throughput fell under 5%. On eight NVIDIA B300 GPUs, a 47-billion-parameter MoE reached 52,000 tokens per second per GPU, roughly 2.7× the earlier implementation. It has been benchmarked at over one trillion parameters.
No score is assigned. Sources and their independence are shown in the citation chain below.