— reading now
Flash

Story·Models & Products·2026-10-07 01:55

Google Launches EmbeddingGemma 2: 740M-Param Multimodal Embedding Model Goes Apache 2.0, Runs On-Device

Google announced EmbeddingGemma 2 on October 6 (ET): a 740M-parameter multimodal embedding model, open-sourced under Apache 2.0, with weights live on Hugging Face and Kaggle. It squeezes text, code, images, audio and video into a single shared vector space: the text-only variant needs roughly 191MB of RAM, the full multimodal build about 567MB, the 8K-token context window is 4x larger than the previous generation, and it ships with LiteRT, MediaPipe, LangChain and llama_index support.

What Happened

Google's official blog announced EmbeddingGemma 2 today, echoed by the @GoogleDeepMind X account at around 12:15 PM ET. It's a multimodal embedding model built for on-device use — phones and laptops — on the Gemma 4 architecture. The design is modular: vision or audio components you don't need can be dropped to save memory on the device.

It became one of r/LocalLLaMA's hottest posts of the day (42 upvotes, 8 comments), and within an hour someone had posted it running locally in-browser on WebGPU.

Key Facts

  1. 740M parameters, Apache 2.0: The license is Apache 2.0, not the more restrictive community license the Gemma family has historically carried — commercial use, modification and redistribution are allowed. The first-generation EmbeddingGemma passed 20 million downloads in a year; the multimodal successor is also released under Apache 2.0.
  2. Key numbers: A 9.92-point gain on MTEB Code (68.76 to 78.68); an 8K-token context window, 4x the previous generation; Matryoshka Representation Learning compressing vector storage to one-sixth; a single pass handling 5.5 minutes of audio, 29 images or 58 video frames; roughly 191MB of RAM for text-only, 567MB for the full multimodal build.
  3. The scope of "best-in-class for its size": The official line is best-in-class for its size: the strongest at 740M, able to beat some models twice its size. The comparison does not include Google's own cloud-based Gemini Embedding 2. Google's stated division of labor: EmbeddingGemma for on-device and offline use, the Gemini Embedding API for large-scale server workloads.
  4. One model instead of several pipelines: Cross-modal search used to mean one embedding model per modality, separate indexes per modality, and code to merge the results. Now it is one model, one call, one vector space.

Context

The first EmbeddingGemma launched in September 2025 as an on-device text embedding model and reached 20 million downloads in a year, used for local RAG. Embedding models don't generate or converse; they turn content into vectors for understanding and retrieval, and in agent and RAG pipelines embedding quality affects retrieval results. With this release Google brings multimodal embedding to on-device hardware.

Why it matters

The first generation passed 20 million downloads in a year; the new one puts text, code, images, audio and video in one vector space under Apache 2.0, with the text-only build at about 191MB of RAM.
Useful Tap if this story helped you

SourcesGoogle official blog (2026-10-06, via Future Tools), @GoogleDeepMind on X (2026-10-06), AI Weekly (2026-10-06), didcodexreset, oossa, r/LocalLLaMA community discussion. Compiled from public reporting; not investment advice.

Comments

  1. Loading comments…
Ask the cat