Google DeepMind released EmbeddingGemma 2 as an Apache 2.0 open model (740M total parameters across text, vision, and audio towers) that projects text, code, images, video, and audio into a single 768-dimensional space, targeting consumer-hardware and on-device deployment. Four independent implementations have appeared within the launch window: native llama.cpp GGUF support, a WebGPU in-browser demo, a third-party ruNNtime WebGPU port with in-browser image-text retrieval, and a first-party Mac app (DigUp) for multimodal file search. The vendor claims 'best-in-class' performance but no independent benchmarks exist — the model is absent from MTEB, and comparisons against qwen3-embedding (0.6B/4B/8B), jina Omni, and siglip2 are entirely untested; per-modality and mixed-modality retrieval quality remains unknown. The case watches whether tooling integrations (Ollama, sentence-transformers, transformers.js — following v1's adoption playbook) and real local-RAG/agent-memory adoptions sustain beyond launch week.
Google DeepMind has independently shipped an Apache 2.0 multimodal embedding model for on-device use — the exact pattern Scott's stack already practices piecemeal (open-weights, local inference, embeddings-as-sensor under a wiki-graph). EmbeddingGemma 2 is a direct candidate to replace BGE-M3 in Scott's dev-wiki recall and code/session indexing (search), and the case's sustaining condition (Ollama, sentence-transformers, transformers.js integrations) maps to his actual tooling chain. The quality vacuum (no MTEB, no independent benchmarks vs qwen3-embedding/jina Omni) directly engages his evaluation-driven development and capability-audit frameworks.
dev:technology.bge-m3dev:project.dev-wikidev:project.searchip:framework.sovereign-software-assuranceip:framework.rag-wiki-substrate-ruleip:concept.evaluation-driven-developmentdev:technology.ollamadev:technology.sentence-transformersip:concept.model-perishabilitydev:concept.endpoint-independent-vector-identityradar:awareness-local-agent-memoryradar:cuemap-deterministic-agent-memoryradar:fraise-temporal-agent-memoryradar:fastrecall-cross-model-memory-apiradar:futureos-context-compaction-recallradar:agent-memory-add-search-evaluationradar:aa-agentperf-local-benchmarkradar:apogee-local-first-browser-agentradar:adaptive-kv-cache-streamingradar:fab-finance-agents-benchmarkradar:gamow-labs-labbench-wetlab-decisionsradar:abliterated-weights-agent-backdoorradar:0pirate-ast-anonymizer-mcp-proxyradar:49ide-spatial-agent-workspaceradar:514-coding-agent-simulation-infraradar:anthropic-claude-eval-pluginradar:baseten-open-models-in-codex
queries asked of Scott's wikis
- open-weights embedding model strategy and model sovereignty
- local on-device multimodal embedding for agent memory and RAG
- embedding model evaluation benchmarks MTEB and independent quality testing
- browser WebGPU deployment patterns for embedding models
- tooling integration playbook: Ollama sentence-transformers transformers.js adoption
- BGE-M3 replacement criteria for code-session indexing and dev-wiki recall
2026-10-11T16:06:30Z
Launch week (~120h) ending with strong browser/edge deployment velocity (5 independent implementations across native llama.cpp, WebGPU×2, Mac, Linux; 2 shipping end-user apps) but sustaining condition unmet: zero tooling integrations (Ollama, sentence-transformers, transformers.js). Quality vacuum persists (no MTEB, no independent benchmarks vs qwen3-embedding/jina Omni/siglip2). Measured heat cooled to 20 pts/h absolute but remains 96.9th percentile steady. Direct candidate for Scott's BGE-M3 replacement in dev-wiki recall and code/session indexing.
2026-10-11T15:42:37Z
evidence attached: hn.story.50043847 — Browser deployment of EmbeddingGemma 2 demonstrates on-device multimodal embedding adoption, directly relevant to the case's hypothesis about becoming the default open on-device model.
2026-10-11T02:35:43Z
Fifth independent deployment appears: a Linux semantic search tool using EmbeddingGemma 2 via llama.cpp (indexes PDFs, Office, scans, screenshots, audio, video with hybrid vector+FTS5 ranking). This is the second end-user application after DigUp Mac app, and the first demonstrating the llama.cpp path producing a shipping tool — a step toward the Ollama/sentence-transformers/transformers.js integrations that define the sustaining condition. Implementation periphery now spans 5 deployments across native llama.cpp, WebGPU (webml + ruNNtime), Mac, and Linux. Launch week nearly complete; tooling integrations still absent; quality vacuum persists (no MTEB, no independent benchmarks vs qwen3-embedding/jina Omni/siglip2, per-modality retrieval unknown).
2026-10-11T02:33:19Z
evidence attached: reddit.post.1x2vz3y — Independent builder releases local semantic search tool using EmbeddingGemma 2 via llama.cpp, demonstrating practical adoption for on-device multimodal retrieval.
2026-10-10T20:17:56Z
DigUp Mac app launch generated a velocity spike (220 pts on reddit.post.1x2eeds) and cross-platform top-decile spread (magnitude_valve_eligible), confirming the fourth independent deployment and first end-user application. Implementation periphery now spans native llama.cpp, WebGPU (webml + ruNNtime), and a shipping Mac app. Sustaining condition unchanged: Ollama/sentence-transformers/transformers.js integrations (v1's playbook) still absent. Quality vacuum persists: no MTEB, no independent benchmarks vs qwen3-embedding/jina Omni/siglip2, per-modality retrieval unknown. Agent-memory topic band hot (71 episodes) directly adjacent to Scott's stack.
2026-10-10T18:47:50Z
evidence attached: hn.story.50035576 — Independent tool (DigUp) adopting EmbeddingGemma 2 for local semantic file search — direct corroboration of on-device multimodal embedding adoption.
2026-10-10T14:36:47Z
grounded: converges/high — Google DeepMind has independently shipped an Apache 2.0 multimodal embedding model for on-device use — the exact pattern Scott's stack already practices pieceme
2026-10-10T14:20:12Z
First-party Mac app (DigUp) shipping on-device EmbeddingGemma 2 for multimodal file search adds a real adoption artifact — the fourth independent deployment (native llama.cpp, webml WebGPU, ruNNtime WebGPU, now DigUp) and the first end-user application. The implementation periphery is expanding while engagement velocity cools (9.5 pts/h, 79th pct peer). Sustaining condition still unmet: Ollama/sentence-transformers/transformers.js integrations (v1's playbook) absent. Quality vacuum persists: no MTEB, no independent benchmarks vs qwen3-embedding/jina Omni/siglip2, per-modality retrieval unknown. Agent-memory topic band hot (71 episodes) directly adjacent to Scott's stack.
2026-10-10T13:39:44Z
evidence attached: reddit.post.1x2eeds — First-party artifact (DigUp Mac app) demonstrating on-device EmbeddingGemma 2 integration for multimodal file search — independent adoption evidence for the accelerating case.
2026-10-08T11:48:25Z
Scott up-voted (explicit attention); the ruNNtime WebGPU port saw a 7.7× velocity spike confirming the third-party browser implementation is gaining independent traction; a GitHub use-cases repo appeared (hn.story.50003695) but remains a single low-engagement post, not the tooling integrations (Ollama, sentence-transformers, transformers.js) the case's sustaining condition requires. Multi-platform top-decile spread confirmed (magnitude_valve_eligible). Quality vacuum persists: no MTEB entry, no independent benchmarks vs qwen3-embedding/jina Omni/siglip2, per-modality retrieval quality unknown. Agent-memory topic band is hot (64 open episodes), directly adjacent to Scott's stack.
2026-10-08T10:42:27Z
evidence attached: hn.story.50003695 — GitHub repository demonstrating use cases for EmbeddingGemma 2; directly extends the accelerating open-case about the same model.
2026-10-07T10:37:48Z
A second, fully third-party implementation landed in-window: the ruNNtime WebGPU port of text+vision towers with in-browser image-text retrieval — the case's adoption condition now shows two independent browser/edge deployments plus native llama.cpp, which moves it corroborated→accelerating on substance (a new implementation post-corroboration in a new tooling ecosystem), not engagement. The case's meaning shifts from 'is it real/deployable' to 'will it sustain and is it any good': velocity cooled to ~35/h (94th pct) off its ~267/h peak while the quality questions (MTEB absence, qwen3/jina/siglip comparisons) remain entirely unanswered.
2026-10-07T10:25:58Z
evidence attached: reddit.post.1wzs1t5 — Third-party WebGPU browser port with vision tower is exactly the browser/edge deployment and tooling-integration evidence the case's adoption condition hinges on.
2026-10-07T04:53:51Z
Launch spread continued (HN main 162→240, Reddit main 325→374) but nothing substantive arrived: the newly attached HN post is a duplicate whose only comment links back to the main thread, and thread content is repetitive quality questions (MTEB absence, unanswered qwen3-embedding and jina Omni comparisons). Momentum is cooling off its peak yet still 99th-percentile across three platforms, so heat holds at medium; state holds at corroborated — the implementations that earned it predate this look and engagement alone does not promote.
2026-10-06T23:36:37Z
evidence attached: hn.story.49984322 — Independent HN traction on the EmbeddingGemma 2 release adds spread to the open launch-sustainment case.
2026-10-06T22:48:46Z
The case crosses from 'watching' to 'corroborated' on substance, not engagement: two independent implementations landed in-window — llama.cpp support merged (PR #30054 + official GGUFs) and the WebGPU in-browser demo running the model locally — alongside cross-platform spread (Reddit, LocalLLaMA, HN). The open question shifts from 'is it real' to 'is it good': vendor 'best-in-class' claims remain unverified with no independent benchmarks vs the qwen3-embedding line.
2026-10-06T20:42:17Z
evidence attached: hn.story.49980487 — Independent HN traction on the model's launch feeds the open case's launch-week spread evidence (though it is spread, not adoption).
2026-10-06T20:42:17Z
evidence attached: reddit.post.1wz7faa — LocalLLaMA launch echo with strong engagement (71+ score) — independent community spread evidence for the adoption question the case watches.
2026-10-06T20:42:17Z
evidence attached: reddit.post.1wz9jvx — Well-received LocalLLaMA echo of the release shows community traction the watching case needs, though it's vendor-blog-derived, not independent corroboration of the quality claim.
2026-10-06T17:04:57Z
grounded: converges/high — Google DeepMind has independently shipped what Scott's stack already practices piecemeal: an open, on-device embedding model for exactly the local-recall worklo
2026-10-06T16:55:20Z
case created — First-party open release with real community pickup (74 pts) and an immediate WebGPU in-browser demo already running the model locally.