Skip to content
In depthStableTechnologyFirst published 2026-10-08 04:00

Google DeepMind releases EmbeddingGemma 2 with 740 million parameters for on-device multimodal embeddings

The open-weights architecture maps text, code, audio, and visual inputs into a single vector representation under an Apache 2.0 license. It introduces a modular parameter footprint designed to operate within memory limits on consumer hardware.

Google DeepMind releases EmbeddingGemma 2, a lightweight multimodal embedding model for on-device use
AI illustration
740M

EmbeddingGemma 2 comprises 740 million total parameters released under an Apache 2.0 open-source license.1

8K tokens

The model features an 8K token context window, representing a fourfold increase over EmbeddingGemma 1.1

567MB

Running the full multimodal model quantized on a Google Pixel 11 Pro uses roughly 567MB of operational RAM.1

Story

A Modular 740-Million-Parameter System for Compact Multimodal Embedding

1 1 1 1

1 3 13 1

1 1 1 1

1 1 1 1

1 1 1 1

1 1 1 1

2 2 2 2

13 1 1 3

12 2 2 1

2 2 2 2

2 2 2 12

Structure

Who is connected to whom
  1. 1EmbeddingGemma 2
  2. uses →Fact
    2Gemma 4Technology
  3. enables →Fact
  4. uses →Fact
  5. uses →Fact
    5LiteRTTechnology
  6. uses →Fact
    6llama.cppTechnology

History

How it came to this
  1. Last yearLaunch of EmbeddingGemmaGoogle introduced the initial EmbeddingGemma model, which went on to accumulate more than 20 million downloads.
  2. Now78.68 EmbeddingGemma 2 score on MTEB Code benchmarks
  3. UpcomingEmbeddingGemma 2 is planned for enterprise distribution through the Gemini Enterprise Agent Platform Model Garden.

Impact

Spreading outward, level by level
  1. Level 1Modular multimodal architecture

    EmbeddingGemma 2 maps text, image, video, and audio data into a shared embedding space using a 270M text base and optional 170M vision and 300M audio modules, with vectors truncatable from 768 down to 128 dimensions via Matryoshka Representation Learning.13

    Fact
  2. Level 2On-device and runtime resource footprint

    Quantized execution on mobile hardware such as a Google Pixel 11 Pro consumes approximately 191MB for text and 567MB for all modalities, while shared components with Gemma 4 reduce memory overhead when co-located.12

    Fact
  3. Level 3Local cross-modal retrieval workflows

    Developers can implement private, low-latency multimodal search and vector indexing directly inside web browsers or on client devices without depending on server-side embedding APIs.

    Analysis

Ahead

Checked automatically when due; the result goes to the track record
Upcoming

EmbeddingGemma 2 is planned for enterprise distribution through the Gemini Enterprise Agent Platform Model Garden.

WatchingTrack record

Sources

What each source supports
1
Google DeepMind Blog
2026-10-06 19:57
Original ↗
Supports 10 points
Citing: Google DeepMind
2
Hacker News Front Page
2026-10-06 16:03
Original ↗
Supports 5 points
Citing: Unsloth, Qdrant
3
IT之家
2026-10-06 22:41
Original ↗
Supports 2 points
Citing: IT之家, 谷歌

Written by AI from the sources listed below: every fact was checked word for word against its source, and inference is marked apart. How we write

2026-10-07 19:17 First published