Google DeepMind releases EmbeddingGemma 2 with 740 million parameters for on-device multimodal embeddings
The open-weights architecture maps text, code, audio, and visual inputs into a single vector representation under an Apache 2.0 license. It introduces a modular parameter footprint designed to operate within memory limits on consumer hardware.
EmbeddingGemma 2 comprises 740 million total parameters released under an Apache 2.0 open-source license.1
The model features an 8K token context window, representing a fourfold increase over EmbeddingGemma 1.1
Running the full multimodal model quantized on a Google Pixel 11 Pro uses roughly 567MB of operational RAM.1
Story
A Modular 740-Million-Parameter System for Compact Multimodal EmbeddingGoogle DeepMind has introduced EmbeddingGemma 2, a compact embedding system comprising 740 million parameters made available under an Apache 2.0 license.1 The release builds directly upon the foundational engineering that powers Google's Gemini Embedding family.1 The system translates arbitrary combinations of written passages, still pictures, sound recordings, and video sequences into an integrated vector space.1 By projecting disparate forms of media into a shared mathematical coordinate system, the architecture enables cross-modal retrieval directly within a single index.1
The release follows the widespread uptake of the original EmbeddingGemma release, which accumulated more than 20 million downloads after launching last year.1 Whereas early iterations targeted textual inputs, IT之家 reported that EmbeddingGemma 2 broadens these functional capabilities to span programming code, pictures, video, and acoustic content within that same vector space.3 This operational breadth allows downstream search pipelines to process both written language and multimedia without needing disconnected indexing pipelines.13 The substantial download history of the initial release established an existing base of developers seeking small-footprint embedding components.1
To keep computation manageable across diverse edge platforms, EmbeddingGemma 2 relies on an explicitly segmented design.1 Standard language tasks require only a base text engine configured with 270M parameters.1 Systems handling visual or auditory data can attach an optional 170M vision encoder alongside a 300M audio encoder according to deployment demands.1 The complete multimodal assembly is reported to total 740 million parameters, an amount that matches the combined sum of the individual text, visual, and acoustic building blocks.1
Vector storage demands are further reduced through the integration of Matryoshka Representation Learning within the model.1 This methodology allows operators to slice full 768-dimensional output embeddings down to narrower lengths without retraining the weights.1 Depending on downstream latency targets and storage limits, outputs can be restricted to 512 dimensions, 256 dimensions, or 128 dimensions.1 Smaller dimensional cuts provide substantial bandwidth and memory savings while preserving the directional geometry of the higher-dimensional vectors.1
In addition to dimensional flexibility, EmbeddingGemma 2 expands contextual ingestion boundaries for lengthy content.1 The framework features an 8K token context window for reading and encoding lengthy inputs.1 This capacity represents a fourfold expansion over the context limits present in EmbeddingGemma 1.1 The enlarged context span allows systems to process extended code files or multi-page documents in one pass without premature truncation.1
Because the model is targeted for local execution, its active memory consumption has been calibrated for mobile processors.1 When evaluated under quantization on a Google Pixel 11 Pro, the model's text weights demand approximately 191MB of active operational RAM.1 Activating the full multimodal suite on the same Google Pixel 11 Pro hardware requires about 567MB of active RAM.1 These hardware requirements ensure that even the fully loaded multimodal configuration can run alongside user tasks on handheld devices.1
System co-location benefits from shared algorithmic components across Google's lightweight model line.2 EmbeddingGemma 2 shares its text tokenizer directly with Gemma 4.2 It also shares an identical audio encoder with Gemma 4.2 Running the two models side by side on one machine lowers overall memory consumption by deduplicating these shared assets.2
Evaluations across standard benchmark suites demonstrate measurable accuracy improvements across technical and multilingual workloads.13 On the MTEB Code evaluation suite, EmbeddingGemma 2 achieved a score of 78.68, climbing from the 68.76 recorded by its direct predecessor.1 This progression reflects an overall margin of improvement of 9.92 points over the previous generation.1 IT之家 noted that across sub-1B parameter multimodal embedding benchmarks including MTEB Code and MAEB, EmbeddingGemma 2 delivers leading marks that equal or exceed those of larger models.3
Distribution channels have been opened across public open-weights hubs to facilitate adoption.12 Engineers can retrieve model checkpoints directly via Hugging Face as well as Kaggle.2 Google also plans to place the model into the Gemini Enterprise Agent Platform Model Garden for enterprise environments.2 The Apache 2.0 licensing terms permit broad commercial integration and adaptation without proprietary restrictions.1
To support widespread serving, the model integrates with an extensive selection of open-source deployment engines.2 Implementations can be loaded through transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, and LMStudio.2 For managing generated vector indexes, Qdrant is officially supported as a vector storage destination.2 These serving pathways allow teams to deploy the model on both local Apple hardware via MLX and distributed backends using vLLM or SGLang.2
Edge accessibility extends into web browsers without requiring dedicated backend infrastructure.2 Developers can run client-side applications inside the browser using transformers.js or through WebGPU acceleration.2 For engineers looking to customize representations on specific datasets, Unsloth provides guidance for fine-tuning EmbeddingGemma 2.2 Combined with flexible serving runtimes, these tools allow the 740M parameter model to be adapted and executed entirely on client endpoints.12
Structure
Who is connected to whom- 1EmbeddingGemma 2
- uses →Fact
- enables →Fact
- uses →Fact
- uses →Fact
- uses →Fact
History
How it came to this- Last yearLaunch of EmbeddingGemmaGoogle introduced the initial EmbeddingGemma model, which went on to accumulate more than 20 million downloads.
- Now78.68 EmbeddingGemma 2 score on MTEB Code benchmarks
- UpcomingEmbeddingGemma 2 is planned for enterprise distribution through the Gemini Enterprise Agent Platform Model Garden.
Impact
Spreading outward, level by level- Level 1Modular multimodal architecture
EmbeddingGemma 2 maps text, image, video, and audio data into a shared embedding space using a 270M text base and optional 170M vision and 300M audio modules, with vectors truncatable from 768 down to 128 dimensions via Matryoshka Representation Learning.13
Fact - Level 2On-device and runtime resource footprint
Quantized execution on mobile hardware such as a Google Pixel 11 Pro consumes approximately 191MB for text and 567MB for all modalities, while shared components with Gemma 4 reduce memory overhead when co-located.12
Fact - Level 3Local cross-modal retrieval workflows
Developers can implement private, low-latency multimodal search and vector indexing directly inside web browsers or on client devices without depending on server-side embedding APIs.
Analysis
Ahead
Checked automatically when due; the result goes to the track recordEmbeddingGemma 2 is planned for enterprise distribution through the Gemini Enterprise Agent Platform Model Garden.
Sources
What each source supportsWritten by AI from the sources listed below: every fact was checked word for word against its source, and inference is marked apart. How we write