Skip to content
StableTechnologyReported 2026-09-24 22:08

LFM2.5-VL-DSpark Announced with Speculative Decoding for Vision-Language Models

The company has released an experimental DSpark draft model for their vision-language model (VLM) LFM2.5-VL-3B, featuring speculative decoding to enhance inference speedup on both CPU and GPU without affecting output quality.

01

Evidence

  • HHugging Face BlogCompany2026-09-24 22:08
    How does speculative decoding work for VLMs Training and Architecture Inference Speedup on CPU and GPU Limitations of speculation for vision workloads How to use LFM2.
    View source
  • HHugging Face BlogCompany2026-09-24 22:08
    How does speculative decoding work for VLMs Training and Architecture Inference Speedup on CPU and GPU Limitations of speculation for vision workloads How to use LFM2.5-VL-DSpark Get Started Citation Today, we release an experimental DSpark draft model for our vision-language mo…
    View source
  • HHugging Face BlogCompany2026-09-24 22:08
    The vision drafter uses the same architecture as our text LFM2.5-DSpark drafters: it captures the target model's hidden states at a fixed set of tapped layers and conditions on them to draft a block of k candidate tokens. Image patches and text tokens are projected into a shared…
    View source