LFM2.5 DSpark Draft Model Release: Inference Speed Up to 3.18 Times Faster
Liquid AI and Hugging Face release LFM2.5 DSpark draft model checkpoints: 1.2B Instruct, 2.6B, and 8B-A1B. A new speculative decoding path boosts inference throughput up to 3.18x on GPU without changing output quality, with significant speedups on edge devices.....
24.7K