A MultiVectorEncoder model finetuned for 14.5 hours on a single RTX 3090, released as multi-vector-encoder/mLateOn-medical, outperformed general-purpose dense, sparse, lexical, and multi-vector retrieval models on a medical evaluation, according to the article. The post introduces MultiVectorEncoder in sentence-transformers for ColBERT-style late interaction retrieval, covering models, datasets, losses, training arguments, evaluators, and trainers. It also reports truncation in released checkpoints cost up to 0.24 NDCG@10 on passages averaging 941 tokens.
No score is assigned. Sources and their independence are shown in the citation chain below.