Up to 3.2x Faster Inference with LFM2.5-DSpark
| Source: Hugging Face Blog
Tags: Liquid AI, LFM2.5, DSpark, speculative decoding, llama.cpp, SGLang, inference optimization
Liquid AI officially releases DSpark draft models for LFM2.5-1.2B, 2.6B, and 8B-A1B — delivering up to 3.18x GPU throughput and 2.87x on-device speedup via speculative decoding, with 57% lower function-calling latency, no change to output quality, and day-one support in llama.cpp and SGLang.
Details
Liquid AI's official blog post on Hugging Face details the architecture and benchmarks for DSpark, its speculative decoding technique for the LFM2.5 family. Each draft model is approximately 300M parameters (295.7M for the 1.2B target, 327.7M for the 2.6B and 8B variants) and adds three components: a DFlash-style parallel backbone conditioned on the target's context features, a lightweight Markov chain head that adds inter-token dependency to raise acceptance at later block positions, and a confidence-scheduled verifier that prunes low-confidence token suffixes before verification.\n\nBenchmark results, measured on a single H100 in BF16 via SGLang and on an M4 Max MacBook Pro via llama.cpp with Metal, show up to 3.18x throughput on GPU and up to 2.87x on-device. Function-calling latency drops 57% on average for LFM2.5-2.6B — meaningful for agentic workloads that chain many tool calls. Under greedy decoding, output is mathematically identical to the target model alone, so all benchmark accuracy scores carry over unchanged.\n\nThe draft model code is open-sourced upstream in both llama.cpp and SGLang. Weights ship as Safetensors and GGUF from Hugging Face. Training ran for 15 epochs on a mix of SFT, chat, code, and function-calling data, with the epoch selected by highest acceptance rate rather than lowest loss.