Google DeepMind releases EmbeddingGemma 2, a lightweight multimodal embedding model for on-device use
Google DeepMind has launched EmbeddingGemma 2, an open-source multimodal embedding model with 740 million parameters that unifies text, images, audio, and video in a shared embedding space. Built on Gemma 4 architecture and released under Apache 2.0 license, it supports on-device inference with as little as 191MB RAM for text-only workloads and features an 8K token context window.
Why it matters. This release advances on-device AI capabilities by enabling privacy-preserving, low-latency multimodal search and retrieval directly on consumer hardware, benefiting developers building offline RAG pipelines and edge applications.