Technology
Multimodal and video AI
Models that see, hear and generate video, and the compute they need.
//
This week
- Nemotron 3.5 Content Safety UpdateNemotron 3.5 introduces unified multimodal evaluation, global language coverage, custom policy enforcement, and reasoning traces, expanding NVIDIA's content safety technology.StableTechnology1 sources · 1 primary · 3 evidence2026-06-05 02:57
- LFM2.5-VL-DSpark Announced with Speculative Decoding for Vision-Language ModelsThe company has released an experimental DSpark draft model for their vision-language model (VLM) LFM2.5-VL-3B, featuring speculative decoding to enhance inference speedup on both CPU and GPU without affecting output quality.StableTechnology1 sources · 1 primary · 3 evidence2026-09-24 22:08
- Meta Announces Muse GlimmerMeta has released Muse Glimmer, a local, agentic, multimodal, and open-source AI model with advanced features like speculative decoding and support for various inference endpoints.StableTechnology1 sources · 1 primary · 3 evidence2026-08-10
- Gemini Omni announced as a multimodal AI modelGemini Omni, a new AI model from Gemini, can create anything from any input, starting with video. It builds on the success of the previous Gemini model, Nano Banana, which helped millions with image generation and editing.Stable1 sources · 1 primary · 4 evidence2026-05-18 03:50
4 verified changes this week.
Daily activity
Developing
Signals in this topic
- 01Google DeepMind releases EmbeddingGemma 2, a lightweight multimodal embedding model for on-device useGoogle DeepMind has launched EmbeddingGemma 2, an open-source multimodal embedding model with 740 million parameters that unifies text, images, audio, and video in a shared embedding space. Built on Gemma 4 architecture and released under Apache 2.0 license, it supports on-device inference with as little as 191MB RAM for text-only workloads and features an 8K token context window.Technology announcement2026-10-07 03:57
- 02SmallRig releases 2025-2026 Top 12 Global Imaging Scenarios Report at 3rd Visionary Storytellers Industry ForumSmallRig hosted the 3rd Visionary Storytellers Industry Forum in Shenzhen and released the '2025-2026 Top 12 Global Imaging Scenarios Report,' identifying 12 core imaging scenarios and four key industry trends including mobile cinematography, decentralization of professional resources, creator democratization, and AI-powered video generation. The report projects 9.2% growth in peripheral equipment over the next five years, outpacing host devices.Technology announcement2026-10-08 16:20
- 03Amazon Announces Improved Voice Agent TechnologyAmazon has released Amazon Nova 2.5 Sonic, an updated speech-to-speech model for voice agents with enhanced reasoning and lower latency, improving real-time voice interactions.Technology announcement2026-10-05 16:10
- 04Google launches EmbeddingGemma 2: supports multimodal, optimizes on-device performanceGoogle has launched EmbeddingGemma 2, which supports multimodal content, with 740 million parameters, suitable for running on mobile devices, requiring only 191MB of memory.Technology announcement2026-10-07 06:41
- 05AiSearch: Interactive Multi-Modal Search with VLMsAiSearch introduces a flexible multimodal retrieval framework using Vision Language Models for natural language search over images and videos, supporting interactive search refinement and visual benchmarking.Technology announcement2026-10-02 12:00
- 06Nemotron 3.5 Content Safety UpdateNemotron 3.5 introduces unified multimodal evaluation, global language coverage, custom policy enforcement, and reasoning traces, expanding NVIDIA's content safety technology.Technology announcement2026-06-05 02:57
- 07Anduril Industries Plans $3.7 Billion Submarine Component Plant at Baltimore PortAnduril Industries is investing $3.7 billion in a submarine-component manufacturing facility at Tradepoint Atlantic, a multimodal logistics hub in Baltimore, creating 3,100 permanent jobs and supporting over 14,000 total jobs.New primary claim2026-10-07 03:00
- 08LFM2.5-VL-DSpark Announced with Speculative Decoding for Vision-Language ModelsThe company has released an experimental DSpark draft model for their vision-language model (VLM) LFM2.5-VL-3B, featuring speculative decoding to enhance inference speedup on both CPU and GPU without affecting output quality.Technology announcement2026-09-24 22:08
- 09Meta Announces Muse GlimmerMeta has released Muse Glimmer, a local, agentic, multimodal, and open-source AI model with advanced features like speculative decoding and support for various inference endpoints.Technology announcement2026-08-10
- 10Gemini Omni announced as a multimodal AI modelGemini Omni, a new AI model from Gemini, can create anything from any input, starting with video. It builds on the success of the previous Gemini model, Nano Banana, which helped millions with image generation and editing.Technology announcement2026-05-18 03:50