← Back to the wire

NeoMME: an efficient Multimodal-native and Multilingual Encoder

AnnouncementModelSep 3, 2026

NeoMME is a family of 260M and 800M multilingual multimodal encoders trained from scratch with a masked discrete-diffusion objective, using a single bidirectional Transformer for both text tokens and raw image patches. Fine-tuned as NeoMME-Retriever with ColPali's page-image approach, both sizes lie on the ViDoRe v3 Pareto frontier for nDCG@10 and model size. The 260M model encodes about 51 pages per second on an L40S GPU, roughly twice ColModernVBERT's throughput. Checkpoints are released under Apache 2.0.

Receipt № 17121 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01medPRIMARY
NeoMMEModelNeoMME-RetrieverModelColPaliModelColModernVBERTModel
Canonical: https://huggingface.co/blog/Hcompany/neomme