← Back to the wire

Up to 3.2x Faster Inference with LFM2.5-DSpark

AchievementModelAug 20, 2026

LFM2.5-DSpark achieves up to 3.18x throughput improvement on GPU and up to 2.87x on-device using speculative decoding. The DSpark method combines a DFlash-style parallel backbone, a sequential Markov chain head, and a confidence-scheduled verifier. For LFM2.5-2.6B, function-calling latency drops by 57% on average. Draft model checkpoints ship with day-one support for llama.cpp and SGLang, and greedy decoding output remains identical to baseline.

Receipt № 14861 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

LFM2.5-2.6BModelDFlashModelLFM2.5-DSparkModelEAGLE-3ModelDSparkModel
Canonical: https://huggingface.co/blog/LiquidAI/lfm25-dspark