跳到正文
Google DeepMind releases EmbeddingGemma 2, a lightweight multimodal embedding model for on-device useAI 插画
新出现技术报道时间 2026-10-07 03:57

Google DeepMind 发布了 EmbeddingGemma 2,这是一个轻量级的多模态嵌入模型,适用于设备端使用。

Google DeepMind 已推出 EmbeddingGemma 2,这是一个开源的多模态嵌入模型,拥有 7.4 亿个参数,能够在共享嵌入空间中统一文本、图像、音频和视频。该模型基于 Gemma 4 架构,并采用 Apache 2.0 许可证发布,支持设备端推理,仅需 191MB RAM 即可处理纯文本任务,并具备 8K token 的上下文窗口。

深度阅读深度解读

为什么重要. 此次发布通过在设备端启用隐私保护、低延迟的多模态搜索和检索功能,直接在消费者硬件上提升AI能力,有助于开发者构建离线RAG流水线和边缘应用。

01

影响到谁

  1. 1EmbeddingGemma 2
  2. 使用 →事实
    2Gemma 4技术
  3. 支撑 →事实
  4. 使用 →事实
  5. 使用 →事实
    5LiteRT技术
  6. 使用 →事实
    6llama.cpp技术
02

证据

  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    Create real-time decision engines leveraging multimodal context for classification, routing, and predictive capabilities via the MediaPipe Decision Task API .
    查看来源
  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    delivering a significant 9.92-point improvement on code performance (in MTEB Code, from 68.76 to 78.68)
    查看来源
  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    Use your favorite development tools : Serve the model efficiently using transformers, sentence-transformers, MLX , vLLM, llama.cpp , SGLang, Ollama , and LMStudio
    查看来源
  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    When paired with generative models such as Gemma 4, EmbeddingGemma 2 enables on-device RAG pipelines that understand complex multimodal data.
    查看来源
  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    EmbeddingGemma 2 has 740 million parameters, making it optimal for on-device inference.
    查看来源
  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    Because EmbeddingGemma 2 is built on Gemma 4 and shares its text tokenizer and audio encoder, developers can run both models together in a unified pipeline with a lower combined total memory footprint.
    查看来源
  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    Develop cross-platform apps with Google AI Edge MediaPipe for turnkey embedding, retrieval & decision tasks or LiteRT for custom model integration
    查看来源
  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    Store your embedding vectors with Qdrant
    查看来源
  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    Develop cross-platform apps with Google AI Edge MediaPipe for turnkey embedding, retrieval & decision tasks
    查看来源
  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    Build for the browser with transformers.js or WebGPU
    查看来源
  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    in MTEB Code, from 68.76 to 78.68
    查看来源
  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    With more than 20 million downloads, builders have used it to power smarter on-device search tools and privacy-first retrieval augmented generation (RAG) pipelines.
    查看来源
  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    To learn how to build on-device search and RAG systems with LiteRT, read the Google AI Edge blog post .
    查看来源
  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    Today, we’re launching EmbeddingGemma 2 , expanding beyond text to unify code, images, video, and audio in a shared embedding space. Built on the Gemma 4 architecture and released under a commercially permissive Apache 2.0 license, EmbeddingGemma 2 has 740 million parameters, ma…
    查看来源
  • GGoogle DeepMind Blog公司披露2026-10-07 03:57
    EmbeddingGemma 2 is the most capable model for on-device multimodal embeddings, natively mapping combinations of text, images, audio, and video into a unified embedding space.
    查看来源