LFM2.5-DSpark achieves up to 3.18x throughput improvement on GPU and up to 2.87x on-device using speculative decoding. The DSpark method combines a DFlash-style parallel backbone, a sequential Markov chain head, and a confidence-scheduled verifier. For LFM2.5-2.6B, function-calling latency drops by 57% on average. Draft model checkpoints ship with day-one support for llama.cpp and SGLang, and greedy decoding output remains identical to baseline.
No score is assigned. Sources and their independence are shown in the citation chain below.