Liquid AI has published a set of draft models called DSpark for its LFM2.5 language model family. Using speculative decoding, they deliver up to 3.18x higher throughput on a GPU and up to 2.87x on a MacBook without altering the text the model produces. The weights are available on Hugging Face, and both llama.cpp and SGLang supported them from day one.
What speculative decoding actually speeds up
When a large language model emits text one token at a time, the bottleneck is not arithmetic but the time spent pulling model weights out of memory. That holds on a high-end GPU such as the NVIDIA H100 and equally on everyday hardware like a MacBook or a phone. Every single token forces another pass over an enormous set of weights, so the compute units spend most of their time waiting.
Speculative decoding spreads that waiting around. A small, cheap draft model proposes several upcoming tokens, and the main model checks all of them in a single forward pass. Whatever matches is accepted in bulk, and the main model takes over from the first mismatch. Since more tokens get finalized per weight read, the perceived speed goes up.
The important part is that greedy decoding produces the same output as before, by construction. A drafted token is accepted only when it agrees with the target model's distribution, and a rejected one is replaced by whatever the target model would have emitted. Benchmark accuracy therefore does not move. It is easiest to think of this as buying speed and nothing else.
The three pieces DSpark adds
DSpark itself is a method introduced in 2026, sitting in the same lineage as EAGLE-3 and DFlash. It combines three components.
The first is a parallel backbone that takes context features from the target model and emits hidden states and base logits for a whole block of candidate tokens at once. The second is a lightweight sequential head that treats neighboring tokens as a Markov chain, nudging each position toward continuations consistent with the token sampled just before it. Purely parallel drafting loses accuracy at later positions, so this restores the dependency and raises the acceptance rate.
The third is a confidence-scheduled verifier. A separate head predicts the acceptance probability of each drafted token and prunes trailing tokens that would cost more to verify than they are worth. Being able to vary the verification window according to available hardware capacity is the practical engineering trick here.
The published draft models are small: 5 layers, a block size of 9, and roughly 300 million parameters each. Liquid AI notes that training and all ablation studies ran exclusively on AMD hardware using its own training framework.
How much faster depends on how often the draft is right
The numbers swing quite a bit depending on the target model and the workload. Measurements were taken on a single H100 80GB with SGLang and on an M4 Max MacBook Pro with llama.cpp and Metal, both at batch size 1 and temperature 0. The evaluation used MATH500, GSM8K, HumanEval, MBPP and MT-Bench.
The flagship 2.6B model improves across the board on both the GPU and the device, reaching roughly 140 tokens per second on a MacBook. That is faster than many commercial cloud models respond, which makes a clear case for keeping the workload local.
The 1.2B model, a non-reasoning model, shows much more variation in acceptance rate between datasets, and the speedup swings by as much as 52 percent. The 8B-A1B model is more complicated still: it averages 2.54x on a GPU, but only 18 percent on device. The current MoE implementation in the Metal backend of llama.cpp is part of the reason, and verifying several tokens at once activates more experts and therefore moves more weight traffic. Liquid AI published these figures openly and treats the gap as future work.
The payoff shows up where you are kept waiting
The use case highlighted as benefiting most is agentic work that runs through tool calls. In that pattern the model reasons before every call, and the user waits through all of it.
Measured on BFCL, a function-calling evaluation dataset, latency dropped by an average of 57 percent across a range of multi-tool scenarios. Whether the wait is cut by more than half is often the difference between an on-device agent being usable day to day or not. Liquid AI's stated goal of making the 2.6B model the first practical on-device agentic model makes sense in that light.
Running it yourself
The draft models are on Hugging Face in Safetensors format and as GGUF files. GGUF targets on-device inference through llama.cpp, while SGLang covers GPU-backed production serving. Both integrations are merged upstream, so there is no separate fork to track. One caveat: the Metal figures were measured with experimental kernels, so reproducing them exactly will depend on your environment.
The weights are open, free to download, fine-tune and redeploy without restrictions, and the three sizes of 1.2B, 2.6B and 8B-A1B let you trade accuracy against footprint.
Summary
DSpark is a set of draft models that raise inference speed while leaving output untouched. The headline figures are up to 3.18x on a GPU and up to 2.87x on a MacBook, but the gain tracks how often the draft is accepted, and the MoE-based 8B model still sees limited benefit on device. Even so, cutting function-calling latency by an average of 57 percent moves locally running agents a meaningful step closer to being practical.
