Skip to content
Google DeepMind releases EmbeddingGemma 2, a lightweight multimodal embedding model for on-device useAI illustration
NewTechnologyReported 2026-10-07 03:57

Google DeepMind releases EmbeddingGemma 2, a lightweight multimodal embedding model for on-device use

Google DeepMind has launched EmbeddingGemma 2, an open-source multimodal embedding model with 740 million parameters that unifies text, images, audio, and video in a shared embedding space. Built on Gemma 4 architecture and released under Apache 2.0 license, it supports on-device inference with as little as 191MB RAM for text-only workloads and features an 8K token context window.

In depthRead the in-depth story

Why it matters. This release advances on-device AI capabilities by enabling privacy-preserving, low-latency multimodal search and retrieval directly on consumer hardware, benefiting developers building offline RAG pipelines and edge applications.

01

Who it touches

  1. 1EmbeddingGemma 2
  2. uses →Fact
    2Gemma 4Technology
  3. enables →Fact
  4. uses →Fact
  5. uses →Fact
    5LiteRTTechnology
  6. uses →Fact
    6llama.cppTechnology
02

Evidence

  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    Create real-time decision engines leveraging multimodal context for classification, routing, and predictive capabilities via the MediaPipe Decision Task API .
    View source
  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    delivering a significant 9.92-point improvement on code performance (in MTEB Code, from 68.76 to 78.68)
    View source
  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    Use your favorite development tools : Serve the model efficiently using transformers, sentence-transformers, MLX , vLLM, llama.cpp , SGLang, Ollama , and LMStudio
    View source
  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    When paired with generative models such as Gemma 4, EmbeddingGemma 2 enables on-device RAG pipelines that understand complex multimodal data.
    View source
  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    EmbeddingGemma 2 has 740 million parameters, making it optimal for on-device inference.
    View source
  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    Because EmbeddingGemma 2 is built on Gemma 4 and shares its text tokenizer and audio encoder, developers can run both models together in a unified pipeline with a lower combined total memory footprint.
    View source
  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    Develop cross-platform apps with Google AI Edge MediaPipe for turnkey embedding, retrieval & decision tasks or LiteRT for custom model integration
    View source
  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    Store your embedding vectors with Qdrant
    View source
  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    Develop cross-platform apps with Google AI Edge MediaPipe for turnkey embedding, retrieval & decision tasks
    View source
  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    Build for the browser with transformers.js or WebGPU
    View source
  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    in MTEB Code, from 68.76 to 78.68
    View source
  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    With more than 20 million downloads, builders have used it to power smarter on-device search tools and privacy-first retrieval augmented generation (RAG) pipelines.
    View source
  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    To learn how to build on-device search and RAG systems with LiteRT, read the Google AI Edge blog post .
    View source
  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    Today, we’re launching EmbeddingGemma 2 , expanding beyond text to unify code, images, video, and audio in a shared embedding space. Built on the Gemma 4 architecture and released under a commercially permissive Apache 2.0 license, EmbeddingGemma 2 has 740 million parameters, ma…
    View source
  • GGoogle DeepMind BlogCompany2026-10-07 03:57
    EmbeddingGemma 2 is the most capable model for on-device multimodal embeddings, natively mapping combinations of text, images, audio, and video into a unified embedding space.
    View source