Liquid AI releases DSpark draft models for LFM2.5, up to 3.2x faster inference
Original titleLFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook
AISummary
Liquid AI released DSpark speculative decoding draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B on Hugging Face. The draft models reach up to 3.18x throughput improvement on an H100 GPU and up to 2.87x on-device, and the outputs match baseline greedy decoding by construction. Support is available in llama.cpp and SGLang, with the speedup varying by model and dataset.
AIWhy it matters
The release reports measured speedups on both H100 and MacBook hardware, with per-dataset results and acceptance rates that show where speculative decoding helps most.
Source: Liquid AI Blog · liquid.aiPublished · added here