Google DeepMindは、デバイス上で使用できる軽量なマルチモーダル埋め込みモデルであるEmbeddingGemma 2をリリースしました。
Google DeepMindは、テキスト、画像、音声、および動画を共有された埋め込み空間で統合したオープンソースのマルチモーダル埋め込みモデルであるEmbeddingGemma 2をリリースしました。このモデルは740億のパラメータを持ち、Gemma 4アーキテクチャに基づいて構築され、Apache 2.0ライセンスで公開されています。テキストのみのワークロードでは191MBのRAMでデバイス上での推論をサポートしており、8Kトークンのコンテキストウィンドウを備えています。
Create real-time decision engines leveraging multimodal context for classification, routing, and predictive capabilities via the MediaPipe Decision Task API .
Use your favorite development tools : Serve the model efficiently using transformers, sentence-transformers, MLX , vLLM, llama.cpp , SGLang, Ollama , and LMStudio
Because EmbeddingGemma 2 is built on Gemma 4 and shares its text tokenizer and audio encoder, developers can run both models together in a unified pipeline with a lower combined total memory footprint.
With more than 20 million downloads, builders have used it to power smarter on-device search tools and privacy-first retrieval augmented generation (RAG) pipelines.
Today, we’re launching EmbeddingGemma 2 , expanding beyond text to unify code, images, video, and audio in a shared embedding space. Built on the Gemma 4 architecture and released under a commercially permissive Apache 2.0 license, EmbeddingGemma 2 has 740 million parameters, ma…
EmbeddingGemma 2 is the most capable model for on-device multimodal embeddings, natively mapping combinations of text, images, audio, and video into a unified embedding space.