LFM2.5-VL-DSparkを発表 ビジョン・ランゲージモデルにおける推測的デコードが導入されました
会社は、ビジョン・ランゲージモデル(VLM)LFM2.5-VL-3B用の実験的なDSparkドラフトモデルをリリースしました。このモデルは、CPUとGPUの両方で推論速度を向上させるために、speculative decodingを導入していますが、出力品質には影響を与えません。
会社は、ビジョン・ランゲージモデル(VLM)LFM2.5-VL-3B用の実験的なDSparkドラフトモデルをリリースしました。このモデルは、CPUとGPUの両方で推論速度を向上させるために、speculative decodingを導入していますが、出力品質には影響を与えません。
How does speculative decoding work for VLMs Training and Architecture Inference Speedup on CPU and GPU Limitations of speculation for vision workloads How to use LFM2.
How does speculative decoding work for VLMs Training and Architecture Inference Speedup on CPU and GPU Limitations of speculation for vision workloads How to use LFM2.5-VL-DSpark Get Started Citation Today, we release an experimental DSpark draft model for our vision-language mo…
The vision drafter uses the same architecture as our text LFM2.5-DSpark drafters: it captures the target model's hidden states at a fixed set of tapped layers and conditions on them to draft a block of k candidate tokens. Image patches and text tokens are projected into a shared…