Google's EmbeddingGemma 2 runs search and RAG on device in about 191MB of RAM. Private client data can stay local
For a lot of client projects, the hardest question about AI search is not quality. It is "where does our data go?"
What happened
Google released EmbeddingGemma 2, an open embedding model under the Apache 2.0 license. It turns text, images, audio, video and code into vectors in one shared space, so a text query can find a matching image or audio clip.
From Google's announcement:
- 740 million parameters in total, built on the Gemma 4 architecture. Text only workloads need as little as 270M parameters.
- With quantization, it needs about 191MB of active RAM for text only weights and about 567MB for the full multimodal model, measured on a Pixel 11 Pro.
- Output vectors can be shortened from 768 down to 512, 256 or 128 dimensions, for up to 6x less storage in a local vector database.
- An 8K token context window, four times the first version.
- On MTEB Code, its score rose from 68.76 to 78.68.
The Decoder reports it runs locally without an API key, at roughly 20 to 70 milliseconds per query via WebGPU in the browser, and can power offline RAG apps when paired with a small model like Gemma 4. Google says the first EmbeddingGemma passed 20 million downloads.
My take
Embeddings are the part of a RAG system most people never think about, and they matter more than the chat model for whether the right document gets found.
Three reasons this release matters for builders working with smaller clients:
- Privacy. Clinics, law offices and finance teams often hesitate to send documents to a hosted API. A local embedding model keeps the indexing on their own machine or server.
- Lock in. Simon Willison made the point well: if a hosted embedding model is retired, you pay to re-embed everything you stored. Open weights let you keep running the same model.
- Cost. Shorter vectors and a small model mean a knowledge base search can run on modest hardware.
Before switching, test it on the client's real documents and real questions. Benchmarks are a starting point, not proof it finds the right policy page.
More posts
- Customers' AI agents are getting blocked by human checks and bot defenses. Your client's site may be turning buyers awayOct 7, 2026
- Siena raised $17M to give support, shopping and social agents one shared memory of each customer. Your CRM should do the sameOct 7, 2026
- Atlassian's MCP server now exposes 220+ tools and sorts them into read, write and destructive tiers. Copy that splitOct 7, 2026
