Skip to content

Technology

Multimodal and video AI

Models that see, hear and generate video, and the compute they need.

//

This week

  • Nemotron 3.5 Content Safety UpdateNemotron 3.5 introduces unified multimodal evaluation, global language coverage, custom policy enforcement, and reasoning traces, expanding NVIDIA's content safety technology.StableTechnology1 sources · 1 primary · 3 evidence2026-06-05 02:57
  • LFM2.5-VL-DSpark Announced with Speculative Decoding for Vision-Language ModelsThe company has released an experimental DSpark draft model for their vision-language model (VLM) LFM2.5-VL-3B, featuring speculative decoding to enhance inference speedup on both CPU and GPU without affecting output quality.StableTechnology1 sources · 1 primary · 3 evidence2026-09-24 22:08
  • Meta Announces Muse GlimmerMeta has released Muse Glimmer, a local, agentic, multimodal, and open-source AI model with advanced features like speculative decoding and support for various inference endpoints.StableTechnology1 sources · 1 primary · 3 evidence2026-08-10
  • Gemini Omni announced as a multimodal AI modelGemini Omni, a new AI model from Gemini, can create anything from any input, starting with video. It builds on the success of the previous Gemini model, Nano Banana, which helped millions with image generation and editing.Stable1 sources · 1 primary · 4 evidence2026-05-18 03:50

4 verified changes this week.

Daily activity

Developing
Oct 36Oct 7
Signals in this topic