← Back to the wire

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

AnnouncementModelSep 6, 2026

H Company released NeoMME, a family of 260M and 800M bidirectional encoders that processes text and raw 32×32 image patches through a single Transformer, with no vision tower or decoder. NeoMME-Retriever-260M scores 0.523 nDCG@10 on ViDoRe v3, within 0.002 of 3.75B-parameter ColQwen2.5. Checkpoints ship under Apache 2.0 with Hugging Face support; the 260M model indexes 51.3 pages per second on an NVIDIA L40S. Text retrieval on BEIR-15 trails LateOn.

Receipt № 17781 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

Hugging FaceCompanyNVIDIACompanyH CompanyCompanyNeoMMEModelNeoMME-RetrieverModelColPaliModelALBERTModelSigLIP2Model
Canonical: https://www.marktechpost.com/2026/09/06/h-company-releases-neomme-a-family-of-260m-and-800m-single-tower-multimodal-encoders-that-drop-the-vision-tower-and-causal-decoder/