Skip to content

Google's EmbeddingGemma 2 unifies text, image, video, and audio in one on-device model

· by Pondero Newsdesk

The short version

Google released EmbeddingGemma 2 on October 6, a 740-million-parameter open embedding model that maps four data types into a single vector space and runs locally on a phone.

Google's EmbeddingGemma 2 unifies text, image, video, and audio in one on-device model

Finding a video clip from a half-remembered voice memo, or searching hours of audio with a typed phrase, has normally meant routing data through a cloud model built for one modality at a time. Google's new EmbeddingGemma 2, released October 6, 2026, does that cross-modal matching locally, in a single 740-million-parameter model small enough to run on a phone, per Google's announcement.

What shipped

EmbeddingGemma 2 is built on the Gemma 4 architecture and released under the Apache 2.0 license, so the weights are free to use commercially, per Google. It replaces last year's text-only EmbeddingGemma, which Google says topped 20 million downloads, by natively mapping text, code, images, video, and audio into one shared embedding space rather than stitching together separate encoders. The model is modular: a text-only deployment needs as little as 270 million parameters, with optional 170-million-parameter vision and 300-million-parameter audio encoders layered on for full multimodal support. Context runs to 8,000 tokens, four times the prior version, enough to process roughly 5.5 minutes of audio, 29 images, or 58 video frames, or a mix of those, in a single pass, per Google's model card details in the announcement.

Google also published benchmark gains specific to developers. On MTEB Code, a coding-retrieval benchmark, the model's score rose from 68.76 to 78.68, a 9.92-point jump the company attributes to the new architecture rather than a bigger parameter count. Output vectors can be truncated from 768 dimensions down to 512, 256, or 128 using Matryoshka Representation Learning, which Google says cuts local vector-database storage by up to 6x. Quantized and running on a Pixel 11 Pro, the text-only configuration needs about 191MB of active RAM; the full multimodal model needs about 567MB, per the same post.

What changes for builders

The practical shift is where the compute happens. A developer building retrieval-augmented generation, semantic search, or media-tagging features no longer has to choose between sending audio and video to a cloud API or giving up multimodal matching entirely on a phone or laptop. Because EmbeddingGemma 2 shares its tokenizer and audio encoder with the generative Gemma 4 model, Google says the two can run together in one pipeline with a lower combined memory footprint than running separate systems, which matters for anyone budgeting RAM on constrained hardware. The model is live now on Hugging Face and Kaggle, with support already lined up across MediaPipe, LiteRT, transformers, sentence-transformers, vLLM, llama.cpp, Ollama, LMStudio, and Qdrant, and a Gemini Enterprise Agent Platform Model Garden listing described as coming soon, per Google.

That broad day-one tooling support, paired with a permissive license, lowers the bar for teams that wanted offline, privacy-preserving search but lacked the engineering budget to build a multimodal pipeline from scratch. It does not settle how the benchmark claims hold up against independent testing, since the comparisons published so far are Google's own.

What to watch next

Whether third-party benchmarks confirm EmbeddingGemma 2's reported MTEB Code and audio-retrieval scores once outside developers run their own evaluations, and whether Google brings equivalent multimodal embedding support to Vertex AI for server-side workloads rather than leaving it strictly on-device.

Sources