NeoMME is a family of 260M and 800M multilingual multimodal encoders trained from scratch with a masked discrete-diffusion objective, using a single bidirectional Transformer for both text tokens and raw image patches. Fine-tuned as NeoMME-Retriever with ColPali's page-image approach, both sizes lie on the ViDoRe v3 Pareto frontier for nDCG@10 and model size. The 260M model encodes about 51 pages per second on an L40S GPU, roughly twice ColModernVBERT's throughput. Checkpoints are released under Apache 2.0.
No score is assigned. Sources and their independence are shown in the citation chain below.