LFM2.5-VL-3B pairs a SigLIP2 400M NaFlex vision encoder with the LFM2.5-2.6B text backbone, pre-trained on roughly 34T tokens. The model improves screen understanding, grounding, multi-image reasoning, and function calling over prior releases. It leads its size class on real-world image tasks and matches peers on tool use. Available on Hugging Face, it ships with support across llama.cpp, vLLM, and ONNX.
No score is assigned. Sources and their independence are shown in the citation chain below.