Beating AI News Flash: Google DeepMind has open-sourced EmbeddingGemma 2. The model has only 740 million parameters and can directly process text, code, images, video, and audio, placing them into the same vector space for search and RAG. The previous generation EmbeddingGemma could only handle text.
The model can also load only the parts needed. When processing only text and code, it has 270 million parameters; adding vision brings it to 440 million; loading everything gives 740 million. Google tested it on the Pixel 11 Pro, where after quantization the text-only version uses a minimum of about 191MB of active memory, while the full version uses about 567MB. The context has increased from 2K in the previous generation to 8K, allowing it to process up to about 5.5 minutes of audio, 29 images, or 58 video frames at once.
Code retrieval shows the most obvious improvement. The MTEB Code score rose from 68.76 in the previous generation to 78.68, while multilingual text remained basically flat, rising from 61.15 to 61.36. The model also supports compressing the default 768-dimensional vectors to 512, 256, or 128 dimensions, reducing vector storage space to as little as one-sixth.
The license has also been relaxed. The previous generation EmbeddingGemma used Google's own Gemma terms, while EmbeddingGemma 2 has switched to Apache 2.0. Gemma 4, also released by Google this year, likewise uses Apache 2.0, and several recent open models have become noticeably more friendly to commercial use and secondary development.

