Skip to content

Technology

Inference cost and pricing

What it costs to serve models, and how token prices and efficiency techniques move it.

//

This week

  • LFM2.5-VL-DSpark Announced with Speculative Decoding for Vision-Language ModelsThe company has released an experimental DSpark draft model for their vision-language model (VLM) LFM2.5-VL-3B, featuring speculative decoding to enhance inference speedup on both CPU and GPU without affecting output quality.StableTechnology1 sources · 1 primary · 3 evidence2026-09-24 22:08
  • Meta Announces Muse GlimmerMeta has released Muse Glimmer, a local, agentic, multimodal, and open-source AI model with advanced features like speculative decoding and support for various inference endpoints.StableTechnology1 sources · 1 primary · 3 evidence2026-08-10

2 verified changes this week.

Daily activity

Stable
Oct 31Oct 7