Use EmbeddingGemma 2 modularly, not as a default full model
Choose the smallest encoder and vector size that matches your retrieval surface.
Short answerUse text/code first unless your product already needs cross-modal retrieval across media.
By JasonPublished Oct 10, 2026Last verified Oct 10, 20268 min read

Short answer: small teams should start with EmbeddingGemma 2’s text/code configuration unless they already need image, video or audio retrieval in the first release. The release matters because Google put five content types into one 768-dimensional embedding space, but the full setup is 740 million parameters; that is useful optionality, not a reason to load everything by default. Treat the model as a menu: pick the encoder set and truncation level that match your product’s retrieval job, storage budget and latency target.
How we checked
Researched from the sources listed; Benchdict ran no tests for this article.
- Read the lead source in full: Google Developers Blog: EmbeddingGemma 2: The Developer Guide, fetched 2026-10-07 (text hash 9fe052bc).
- Cross-checked against an independent source: huggingface.co: unsloth/embeddinggemma-2-GGUF · Hugging Face, fetched 2026-10-07 (text hash ec0632dc).
- Followed a link from it: Google AI for Developers: EmbeddingGemma-Modell – Übersicht | Google AI for Developers, fetched 2026-10-07 (text hash 8fb5ec62).
- Cross-checked against an independent source: Unite.AI: DeepMind Debuts EmbeddingGemma 2, Mapping Five Modalities Into One Space, fetched 2026-10-07 (text hash 3d5327d3).
- Cross-checked against an independent source: CellCog: EmbeddingGemma 2: Benchmarks, Specs and How to Run It, fetched 2026-10-07 (text hash 0db1dde6).
Benchdict did not run any tests for this article. Statements about behavior come from the sources above.
Independent reading: 3 site owners besides the lead source's own were read successfully.
---

What changed
EmbeddingGemma 2 is an open Apache 2.0 model based on Gemma 4 that maps text, code, images, video and audio into a shared 768-dimensional embedding space (Google Developers Blog). The important product decision is the modular loadout. You can run text/code only at 270 million parameters, text plus vision at 440 million, text plus audio at 570 million, or the full multimodal setup at 740 million.
That makes the model more like an embedding platform than a single fixed encoder. The full configuration is 174.074% larger than the text/code configuration, so a small team should have a real media-retrieval use case before paying that footprint. Google also says EmbeddingGemma 2 scores 14% higher than EmbeddingGemma 1 on MTEB Code and adds image, video and audio retrieval while retaining multilingual text accuracy.
| Choice | Parameters | Use it when | Avoid it when |
|---|---|---|---|
| Text/code | 270 million | Docs, support search, code search, agentic codebase retrieval | You need image, video or audio retrieval on day one |
| Text + vision | 440 million | Product search, screenshot search, visual knowledge bases | Audio is central to the product |
| Text + audio | 570 million | Voice notes, call archives, audio search | Images or video are central |
| Full multimodal | 740 million | Users search across text, code, image, video and audio together | The first release is text-heavy or storage-constrained |
The vector-size decision
The base embedding space is 768 dimensions, with Matryoshka Representation Learning allowing truncation down to 128 dimensions. Google’s guide says 256 dimensions keeps most text and code quality and about 95% of image, video and speech retrieval quality. At 128 dimensions, text and code retain around 90% quality, while image, video and speech drop to around 75%.
| Vector size | Stated storage implication | Stated quality implication | Practical read |
|---|---|---|---|
| 768 dimensions | One million bfloat16 vectors take roughly 1.5 GB | Full dimension baseline | Use for quality-sensitive retrieval or evaluation baselines |
| 256 dimensions | 3x storage reduction | About 95% multimodal retrieval quality retained | Likely default candidate for mixed-media products |
| 128 dimensions | One million bfloat16 vectors take about 250 MB; 6x reduction | Around 90% text/code quality and around 75% image/video/speech quality | Good for text-heavy, storage-constrained indexes; risky for media-heavy search |
Decision framework
- If your first release is documentation, support or code search, then start with the 270 million-parameter text/code configuration because it covers the core job without loading unused media encoders.
- If users must retrieve screenshots, product images or visual references, then choose text plus vision because the 440 million-parameter load is smaller than the full multimodal setup.
- If the product centers on calls, clips or voice notes, then choose text plus audio because the guide specifies audio inputs and a 570 million-parameter configuration.
- If users will search across text, code, images, video and audio in one workflow, then choose the 740 million-parameter full model because the shared embedding space is the point.
- If storage is the binding constraint, then evaluate 256 dimensions before 128 dimensions because the guide claims 95% multimodal quality at 256 but only around 75% for image, video and speech at 128.
Implementation details that change the estimate
The guide requires sentence-transformers version 6.1.0 or later for the multimodal usage path. Text retrieval examples use different prompt names for queries and documents, so a team should not treat query and document embeddings as interchangeable boilerplate. Media inputs are passed as dictionaries keyed by modality, and interleaved inputs use image, video or audio placeholder tokens. Video is sampled at 1 frame per second by default, and audio should be 16 kHz mono.
The operational upside is migration. Google says a text-only index can later add image or audio embeddings without recomputing previously computed embeddings. For a small team, that reduces the penalty for starting narrow. You can ship text/code search, preserve the index, and add media once the product proves users need it.
Before and after
| Area | Before | Now |
|---|---|---|
| Supported inputs | EmbeddingGemma documentation page describes a multilingual text embedding model | EmbeddingGemma 2 maps text, code, images, video and audio into one embedding space |
| Code benchmark | EmbeddingGemma 1 baseline on MTEB Code | Google says EmbeddingGemma 2 scores 14% higher on MTEB Code |
| Storage | One million 768-dimensional bfloat16 vectors take roughly 1.5 GB | Truncating to 128 dimensions requires about 250 MB |
| Loading footprint | Full multimodal model loads 740 million parameters | Developers can load smaller text/code, text+vision or text+audio configurations |
What we could not confirm
The public material used here does not include exact first-party benchmark tables for every modality. Full tables from Google’s model card would settle that. It also does not state hosting, API, support or enterprise pricing; Apache 2.0 licensing and free weight access do not answer total operating cost. First-party pages do not give a rollout date for Gemini Enterprise Agent Platform Model Garden. They also do not provide measured phone or laptop latency, so “edge” should be read as a design direction rather than a device-specific performance guarantee. The German EmbeddingGemma documentation page still describes a Gemma 3-based multilingual text model with 308 million parameters, which appears to be an older-page mismatch rather than a direct contradiction about EmbeddingGemma 2. The guide also does not name competitors behind claims about leading sub-1B multimodal embedding performance.
Launch checklist
- [ ] Pick the smallest encoder set that covers the first user-facing retrieval workflow.
- [ ] Benchmark 768, 256 and 128 dimensions on your own queries before locking the index size.
- [ ] Use separate query and document prompt names for text retrieval.
- [ ] Normalize video and audio assumptions around 1 frame per second video sampling and 16 kHz mono audio.
- [ ] Keep media expansion in the roadmap only if the product has a real cross-modal search path.
Our position: do not make the full multimodal model your default architecture just because it is available. For most indie products, the first hard question is not “Can the model embed everything?” but “Which retrieval mistake would actually hurt the user?” If the product is documentation search, codebase search or support RAG, the 270 million-parameter text/code setup is the sensible starting point. It keeps the index compatible with later media expansion while avoiding the 740 million-parameter load of the full configuration. The strongest part of EmbeddingGemma 2 is that the upgrade path is cleaner than a separate-model stack: Google says a text-only index can later add image or audio embeddings without recomputing existing embeddings. The overstated part is edge readiness. The guide gives storage figures and low-latency framing, but not exact phone or laptop latency. We would change our mind if first-party benchmark tables and device latency numbers showed the full multimodal path was cheap enough to run everywhere by default.
Use text/code first unless your product already needs cross-modal retrieval across media.
EmbeddingGemma 2 is most useful when treated as a modular retrieval stack: choose encoders and vector dimensions by workload, not by headline capability.
Teams with stable text-only search that cannot evaluate migration quality yet, or teams that need published device latency before committing to local deployment.
Sources
- EmbeddingGemma 2: The Developer GuideGoogle Developers Blog, Oct 6, 2026. Used for: Main facts on model configurations, dimensions, truncation, storage, input handling and claimed improvements.
- unsloth/embeddinggemma-2-GGUF · Hugging Facehuggingface.co, Oct 7, 2026. Used for: Secondary support for Apache 2.0 licensing, shared embedding space and modular loading. (undated page, retrieved 2026-10-07)
- EmbeddingGemma-Modell – Übersicht | Google AI for DevelopersGoogle AI for Developers, Oct 7, 2026. Used for: Context for the documentation-version mismatch around older EmbeddingGemma material. (undated page, retrieved 2026-10-07)
- DeepMind Debuts EmbeddingGemma 2, Mapping Five Modalities Into One SpaceUnite.AI, Oct 6, 2026. Used for: Context that independent coverage described the launch on the same date.
- EmbeddingGemma 2: Benchmarks, Specs and How to Run ItCellCog, Oct 6, 2026. Used for: Context that independent coverage also described the release and Gemma 4 basis.
Jason, Founder & editor, Benchdict. About Benchdict