EmbeddingGemma 2: An Open, Lightweight Multimodal Embedding Model
Google DeepMind released EmbeddingGemma 2, an open model with 740 million parameters, built on the Gemma 4 architecture and licensed under Apache 2.0. It maps text, code, images, audio and video into one shared embedding space and runs on device. The text-only part needs as little as 270M parameters. With quantization on a Pixel 11 Pro, it uses about 191MB of active RAM for text, or about 567MB for the full multimodal model. Matryoshka Representation Learning cuts vectors from 768 to as few as 128 dimensions, for up to 6x less storage. The context window is 8K tokens. MTEB Code rises from 68.76 to 78.68. These figures come from Google.
On October 6, 2026, Google DeepMind released EmbeddingGemma 2, an open, lightweight, natively multimodal embedding model. It maps text, code, images, audio and video into one shared embedding space. It builds on the Gemma 4 architecture, ships under the commercially permissive Apache 2.0 license, and has 740 million parameters. Its target is on-device inference, on consumer hardware. What was announced The first EmbeddingGemma came out last year as a lightweight option for high-quality text embeddings. Google says it passed 20 million downloads. Builders used it for on-device search tools and privacy-first retrieval augmented generation (RAG) pipelines. EmbeddingGemma 2 widens the input range to code, images, video and audio. One model handles all of them natively. The team does not stitch together several specialist models. Google gives two example uses: find a specific video clip from a voice memo, and search hours of audio recordings with a text query. The model comes from the same technology as the Gemini Embedding models.
Modular architecture The design is modular. A text-only workload needs as little as 270 million parameters. A vision encoder of 170 million parameters is optional. An audio encoder of 300 million parameters is also optional. The three parts add up to the full 740 million parameter multimodal model. A developer loads only what the application needs. A search tool that handles only text does not pay memory for vision or audio weights it never uses. Matryoshka Representation Learning and storage cost The model uses Matryoshka Representation Learning (MRL). The output vector has 768 dimensions by default. A developer can truncate it at runtime to 512, 256 or 128 dimensions. Google states this gives up to a 6x reduction in storage for local vector databases and in memory use. For an application that keeps many vectors on a phone or laptop, this is a direct cost lever. Truncation trades accuracy for size. The announcement does not quantify that loss per dimension, so teams should measure it on their own data.
On-device footprint and context window With quantization, on a Google Pixel 11 Pro, the model needs as little as about 191MB of active RAM for text-only weights. The full multimodal model needs about 567MB. The context window is 8K tokens, four times larger than EmbeddingGemma 1. Google says a single input can hold up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations of these, all processed on local hardware. Benchmarks Google states that EmbeddingGemma 2 achieves leading scores among sub-1B multimodal embedders for its size. The named benchmarks include MTEB Code, from the Massive Text Embedding Benchmark, and MAEB, the Massive Audio Embedding Benchmark. It also matches or outperforms many larger models on text, vision and audio tasks. Multilingual text performance matches the first model. The clearest gain is in code. The MTEB Code score rises from 68.76 to 78.68, an improvement of 9.92 points. Google also says that across image, video, document and audio tasks, it outperforms some specialist models more than twice its size. The full metrics are in the model card. These figures come from the vendor. Independent reproduction by the community has not yet been reported in the announcement.
What it means for developers and enterprises Privacy comes first. Embeddings generated locally do not require data to leave the device. That fits sensitive material such as personal notes, legal files and health records. Latency and cost come next. Local inference needs no network round trip and no per-call embedding API fee. The workflow also gets simpler. Cross-modal retrieval used to need separate models for text, images and audio, and a way to align their vector spaces. One model now covers it. Code work benefits too. The jump on MTEB Code makes local codebase indexing, semantic code search and retrieval for coding agents more practical. Finally, Apache 2.0 allows commercial use, which lowers the legal barrier for enterprise integration.
Outlook and challenges On-device multimodal embeddings point to a common capability: a local semantic index over a person's photos, recordings, videos and code, searchable in one place. Several challenges remain. An 8K window is still short for long videos and long recordings, so teams will need chunking and aggregation strategies. Different quantization schemes may change retrieval quality, and that needs testing on real data. Vector databases, retrieval frameworks and mobile runtimes must add support for multimodal embeddings. A benchmark score is not a business result, so each team should evaluate on its own corpus. Taken together, EmbeddingGemma 2 combines three properties in one release: small, multimodal and commercially usable. AI tool developers should watch it closely.