Text, audio, and video share a search space with EmbeddingGemma 2

Local search spans text, images, audio, and video with EmbeddingGemma 2, Google’s 740-million-parameter model released under Apache 2.0.

Text, code, images, audio, and video can now be searched within a shared representation space using EmbeddingGemma 2. The Google DeepMind model turns content into numerical representations called embeddings, whose proximity reflects similarity in meaning. A written query can retrieve an audio passage, while a voice memo can help locate a video clip. Processing can run directly on the device, offline. The weights are available on Hugging Face under the Apache 2.0 license, which permits commercial use.

The architecture separates components according to the task. Of the full model’s 740 million parameters, 270 million are sufficient for text and code; vision adds 170 million, and audio adds 300 million. Unused modules can remain unloaded. With quantization, published measurements on a Google Pixel 11 Pro show active memory usage as low as approximately 191 MB for text-only weights and 567 MB for the full multimodal model. These figures describe an optimized configuration on specific hardware, rather than the total memory required by a search application and its document database.

Representation storage offers another adjustment: the 768-dimensional vectors can be shortened to 128 dimensions, reducing their size sixfold. That trade-off affects retrieval quality. The model card recommends the smallest setting primarily for text workloads and reports a more substantial decline in multimodal quality. With default settings, the 8,192-token context window accommodates approximately 5.5 minutes of audio, 29 images, or 58 video frames. These inputs share the same budget, so combining formats reduces the amount of each that fits.

Google’s published evaluations show a notable improvement on code: the MTEB Code score rises from 68.76 to 78.68, a gain of 9.92 points over the first version. Those results use the full-precision model. For practical exploration, the Google AI Edge demonstrations offer local media search and tools for finding specific moments in videos. Paired with a generative model such as Gemma 4, the system can also retrieve local files and supply relevant passages for a response grounded in those documents.