Technology
Inference cost and pricing
What it costs to serve models, and how token prices and efficiency techniques move it.
//
This week
- LFM2.5-VL-DSpark Announced with Speculative Decoding for Vision-Language ModelsThe company has released an experimental DSpark draft model for their vision-language model (VLM) LFM2.5-VL-3B, featuring speculative decoding to enhance inference speedup on both CPU and GPU without affecting output quality.StableTechnology1 sources · 1 primary · 3 evidence2026-09-24 22:08
- Meta Announces Muse GlimmerMeta has released Muse Glimmer, a local, agentic, multimodal, and open-source AI model with advanced features like speculative decoding and support for various inference endpoints.StableTechnology1 sources · 1 primary · 3 evidence2026-08-10
2 verified changes this week.
Daily activity
Stable
Signals in this topic
- 01ShamAN-Q: Shampoo Augmented NanoQuant for Sub-1-bit LLM WeightsResearchers have developed ShamAN-Q, a new sub-1-bit post-training quantization method for large language models, which improves upon NanoQuant by incorporating a dense curvature metric and Mahalanobis reconstruction loss.Technology announcement2026-10-02 12:00
- 02LFM2.5-VL-DSpark Announced with Speculative Decoding for Vision-Language ModelsThe company has released an experimental DSpark draft model for their vision-language model (VLM) LFM2.5-VL-3B, featuring speculative decoding to enhance inference speedup on both CPU and GPU without affecting output quality.Technology announcement2026-09-24 22:08
- 03Meta Announces Muse GlimmerMeta has released Muse Glimmer, a local, agentic, multimodal, and open-source AI model with advanced features like speculative decoding and support for various inference endpoints.Technology announcement2026-08-10