2026-10-11 16:33 UTC

Google DeepMind's open EmbeddingGemma 2 maps text (incl. code), image, video, and audio into a single 768-dimensional space at 740M parameters for consumer hardware, and becomes the default open on-device multimodal embedding model for local search, RAG, and agent-memory workflows if browser/edge deployments and tooling integrations sustain beyond launch week; a quiet fade closes it.

state: acceleratingheat: mediumuncertainty: mediumconvergesscott: highembedding-models multimodal local-rag on-device-ai agent-memoryGoogle DeepMind

What is this?

Google DeepMind released EmbeddingGemma 2 as an Apache 2.0 open model (740M total parameters across text, vision, and audio towers) that projects text, code, images, video, and audio into a single 768-dimensional space, targeting consumer-hardware and on-device deployment. Four independent implementations have appeared within the launch window: native llama.cpp GGUF support, a WebGPU in-browser demo, a third-party ruNNtime WebGPU port with in-browser image-text retrieval, and a first-party Mac app (DigUp) for multimodal file search. The vendor claims 'best-in-class' performance but no independent benchmarks exist — the model is absent from MTEB, and comparisons against qwen3-embedding (0.6B/4B/8B), jina Omni, and siglip2 are entirely untested; per-modality and mixed-modality retrieval quality remains unknown. The case watches whether tooling integrations (Ollama, sentence-transformers, transformers.js — following v1's adoption playbook) and real local-RAG/agent-memory adoptions sustain beyond launch week.

Why it matters to Scott

Google DeepMind has independently shipped an Apache 2.0 multimodal embedding model for on-device use — the exact pattern Scott's stack already practices piecemeal (open-weights, local inference, embeddings-as-sensor under a wiki-graph). EmbeddingGemma 2 is a direct candidate to replace BGE-M3 in Scott's dev-wiki recall and code/session indexing (search), and the case's sustaining condition (Ollama, sentence-transformers, transformers.js integrations) maps to his actual tooling chain. The quality vacuum (no MTEB, no independent benchmarks vs qwen3-embedding/jina Omni) directly engages his evaluation-driven development and capability-audit frameworks.
dev:technology.bge-m3dev:project.dev-wikidev:project.searchip:framework.sovereign-software-assuranceip:framework.rag-wiki-substrate-ruleip:concept.evaluation-driven-developmentdev:technology.ollamadev:technology.sentence-transformersip:concept.model-perishabilitydev:concept.endpoint-independent-vector-identityradar:awareness-local-agent-memoryradar:cuemap-deterministic-agent-memoryradar:fraise-temporal-agent-memoryradar:fastrecall-cross-model-memory-apiradar:futureos-context-compaction-recallradar:agent-memory-add-search-evaluationradar:aa-agentperf-local-benchmarkradar:apogee-local-first-browser-agentradar:adaptive-kv-cache-streamingradar:fab-finance-agents-benchmarkradar:gamow-labs-labbench-wetlab-decisionsradar:abliterated-weights-agent-backdoorradar:0pirate-ast-anonymizer-mcp-proxyradar:49ide-spatial-agent-workspaceradar:514-coding-agent-simulation-infraradar:anthropic-claude-eval-pluginradar:baseten-open-models-in-codex
queries asked of Scott's wikis
  • open-weights embedding model strategy and model sovereignty
  • local on-device multimodal embedding for agent memory and RAG
  • embedding model evaluation benchmarks MTEB and independent quality testing
  • browser WebGPU deployment patterns for embedding models
  • tooling integration playbook: Ollama sentence-transformers transformers.js adoption
  • BGE-M3 replacement criteria for code-session indexing and dev-wiki recall

Measured heat

now 20 pts/hpeak 358 pts/hcomments 2/hpeers p97momentum: steady3 platformsage 120h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-06 16:55 (minted)⭐ origin echo-reconstructedOpen multimodal embedding model mapping 'text (incl. code), images, video, and audio inputs—and combinations thereof' into one unified 768-d
Google DeepMind on github (echo) · attributed from reddit.post.1wz5va3, reddit.post.1wz6uxn · published time unknown
—
10-06 15:41first on r/LocalLLaMA · published · lag ?google/embeddinggemma-2 · Hugging Face
jacek2023
—
10-06 16:03first on hacker news · published · lag ?EmbeddingGemma 2
ilreb
—
10-06 15:41amplified on r/LocalLLaMAreddit.post.1wz5va3
jacek2023
peak 508 · 116 comments · 19% of case engagement
10-06 16:03amplified on hacker newshn.story.49980487
ilreb
peak 436 · 46 comments · 26% of case engagement
10-06 16:20amplified on r/LocalLLaMAreddit.post.1wz6uxn
xenovatech
peak 88 · 14 comments · 3% of case engagement
10-06 16:42amplified on r/LocalLLaMAreddit.post.1wz7faa
Recoil42
peak 176 · 34 comments · 6% of case engagement
10-06 18:03amplified on r/LocalLLaMAreddit.post.1wz9jvx
sn2006gy
peak 88 · 20 comments · 3% of case engagement
10-06 21:20amplified on hacker newshn.story.49984322
hmokiguess
peak 11 · 1 comments · 1% of case engagement
6 more amplifiers in ainews.case_chain
10-06 16:23our radar first saw it · lag ?discovery anchor: reddit.post.1wz5va3—
pace: p95 vs 1247 stories at the 96h mark (now 120h old) — ahead of anthropic-rnd-automation-index (1.0x), behind deepseek-4-1-flash-underreaction (1.0x)

Evidence (13) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditgoogle/embeddinggemma-2 · Hugging Face
LocalLLaMA
jacek2023514116
🟠 redditEmbeddingGemma 2 running locally in-browser on WebGPU
LocalLLaMA
xenovatech8815
🟧 echo.github ⭐Open multimodal embedding model mapping 'text (incl. code), images, video, and audio inputs—and combinations thereof' into one unified 768-dGoogle DeepMind——
🟠 redditEmbeddingGemma 2 is a best-in-class open model for natively multimodal embeddings
LocalLLaMA
sn2006gy8920
🟠 redditIntroducing EmbeddingGemma 2: A best-in-class open model for natively multimodal embeddings | Google
LocalLLaMA
Recoil4216834
🟧 hnEmbeddingGemma 2ilreb43646
🟧 hnGoogle EmbeddingGemma 2hmokiguess111
🟠 redditImage-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU
LocalLLaMA
FinancialAd1961757
🟧 hnEmbedding Gemma2 Use Caseskarimtn20
🟠 redditOpen-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in them
LocalLLaMA
A-Rahim1137174
🟧 hnDigUp: Search files on Mac by what's in them, using EmbeddingGemma 2AbuAssar10
🟠 redditlocal semantic file search for Linux (Rust, llama.cpp, EmbeddingGemma 2)
LocalLLaMA
somthing_tn70
🟧 hnShow HN: Running EmbeddingGemma2 on the Browserglpcc10

Interpretation history

Decision trace