Back
Google DeepMind Launches EmbeddingGemma 2, a Multimodal On-Device Embedding Model
SiTech AI Team3 min read

Google DeepMind Launches EmbeddingGemma 2, a Multimodal On-Device Embedding Model

Google DeepMind has released EmbeddingGemma 2, an open 740 million parameter model that maps text, images, audio, and video into a single embedding space for on-device search and retrieval.

A Unified Multimodal Embedding Model

Google DeepMind has launched EmbeddingGemma 2, an open, lightweight embedding model that natively maps text, images, audio, and video into a single unified embedding space. Built on the Gemma 4 architecture and released under the Apache 2.0 license, the 740 million parameter model is designed for on-device inference, enabling tasks such as finding a specific video clip from a voice memo or searching hours of audio recordings with a text query, all processed locally by one model.

The release follows the original EmbeddingGemma, which surpassed 20 million downloads and was used to power on-device search tools and privacy-first retrieval augmented generation pipelines. EmbeddingGemma 2 expands beyond text to unify code, images, video, and audio in a shared embedding space.

Modular, Storage-Efficient Design

EmbeddingGemma 2 is modular by design. It requires as little as 270 million parameters for text-only workloads, with optional vision (170 million) and audio (300 million) encoders for full multimodal support. Using Matryoshka Representation Learning, developers can dynamically truncate output vectors from 768 dimensions down to 512, 256, or 128, providing up to 6x storage reduction for local vector databases and memory usage.

With quantization on a Google Pixel 11 Pro, the model requires about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model. Its 8K token context window, four times larger than EmbeddingGemma 1, supports up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations thereof directly on local hardware.

Benchmark Performance

Google DeepMind reports that EmbeddingGemma 2 achieves leading scores among sub-1B multimodal embedders on benchmarks including MTEB Code and MAEB, while matching or outperforming many larger models across text, vision, and audio tasks. Code performance improved by 9.92 points on MTEB Code, from 68.76 to 78.68, making the model well suited for local codebase indexing, semantic code search, and coding agent retrieval. Across image, video, documents, and audio, the company says it sets a new standard in quality-per-parameter for sub-1B models and even outperforms some specialist models more than twice its size.

Availability and On-Device Use Cases

Model weights are available on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform Model Garden availability coming soon. The model can be served with transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, and LMStudio, and embedding vectors can be stored with Qdrant. Fine-tuning guidance is provided by Unsloth.

Developers can try EmbeddingGemma 2 in Google AI Edge Gallery's Instant Media Search and Video Moments Finder, pair it with Gemma 4 for on-device retrieval augmented generation in the Google AI Edge Foresight app, or build real-time decision engines with the MediaPipe Decision Task API. On-device deployment is supported through Google AI Edge MediaPipe, LiteRT, transformers.js, and WebGPU.

Sources: blog.google

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.