Benzi (github.com/oooscoos/Benzi) is a released MCP server by oooscoos that compiles a codebase with tree-sitter into an index of symbols, call edges, inheritance, and data flow, so coding agents like Claude Code can query structure directly instead of pulling chunks from embedding-based retrieval; the ~2x faster/cheaper claim is first-party only with no independent benchmark, and its Reddit launch was received hostilely (ad accusations, cloud-privacy concerns). The category around it is demonstrably active: a 'Codebase-Memory' arXiv paper documents the same tree-sitter-knowledge-graph approach with real adoption metrics (900+ stars and ~100 forks within four weeks of its Feb 2026 release, auto-detected by ten coding agents, scaling to the Linux kernel at 10x lower token cost and 2.1x fewer tool calls), and independent builders β CodeRLM on HN, an MIT fully-local code knowledge graph, a tree-sitter analysis skill in the MCP Market β keep shipping their own versions. The recurring unsolved problem is agent-side, not index-side: CodeRLM's author reports Claude won't reach for indexed lookups unprompted and ships hooks/plugins to force it, while most production pipelines remain hybrid (tree-sitter for structural chunking, embeddings for recall), so 'replace vector RAG entirely' is still the contrarian position rather than a proven default.
A second independent builder (MIT, fully-local call graph, motivated because 'grep finds mentioners, not actual callers') plus the Codebase-Memory arXiv adoption metrics corroborate the world converging on the position Scott's canon already holds β precompiled structural navigation over blind embedding queries ('the index is the data'; Walk a Wiki) β in exactly the domain where his own tooling sits on the other side: dev:project.search is embedding-only ChromaDB, so this bears directly on whether it gains a structural index layer beside his existing tree-sitter skeletonizer. CodeRLM's finding that agents won't reach for indexed lookups unprompted (hooks/plugins to force it) and the hybrid production norm extend his map-injection/harness doctrine and hold the substrate rule's stratify-don't-migrate boundary β Benzi's replace-RAG-entirely claim remains unbenchmarked β so this is publishable dated receipts plus a concrete dev decision, but cold engagement and zero independent measurement keep it at watch-not-act.
ip:framework.rag-wiki-substrate-ruleip:framework.the-index-is-the-data-self-cleaning-wiki-graphip:source.walk-a-wiki-cant-drive-a-rag-ebookip:concept.rag-as-sensorip:concept.map-injectiondev:project.searchdev:concept.deterministic-code-skeletonradar:benzi-repository-map-harnessradar:concept.code-intelligenceradar:concept.repository-intelligenceradar:concept.rag-knowledge-systemsradar:cognee-codebase-memory-efficiencyradar:novgraph-persistent-codebase-memory
queries asked of Scott's wikis
- substrate rule structure as navigation layer embeddings as recall
- Walk a Wiki named links symbol-addressed navigation
- coding agent harness forcing tool usage hooks
- tree-sitter AST code index call graph notes
- embedding retrieval token cost economics
- search project index layer codebase context
now 0 pts/hpeak 29 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 348h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
2026-10-07T17:30:05Z
The only change is Scott's own up-vote β an endorsement of the case, not new world evidence; reobservations are null and no evidence moved. The benchmark-led briefing evidently landed, and Scott is affirming this as a live decision input: relevance moves mediumβhigh on his explicit nudge, since the 1,500-run forced-vs-unforced finding is independent empirical support for his map-injection/stratify doctrine and bears directly on dev:project.search. The world itself is unchanged, so heat stays low and this is not material.
2026-10-07T13:46:04Z
The sensor firing is incremental engagement on Repowise β the second released full-index MCP for Claude Code attached last look β a +3/+3 bump with one positive comment; it confirms the periphery keeps expanding but adds no adoption or measurement and doesn't touch the live forced-usage question. Meaning unchanged since the benchmark repricing: corroborated builder pattern, displacement thesis measured-negative as shipped, cold watch with unchanged escalation triggers.
2026-10-07T13:27:20Z
evidence attached: reddit.post.1wzw6v6 β A second released structure-resolved codebase index (call graph via MCP, local, no embeddings) shows the pattern spreading beyond Benzi.
2026-10-07T01:13:48Z
The independent measurement this case was waiting for arrived and cut against the thesis: a 1,500-run benchmark of 8 Claude Code add-ons found code-graph tools added cost with no accuracy gain and were invoked ~1% of the time β independently replicating CodeRLM's agent-usage finding while memory add-ons helped. The case's meaning shifts from 'untested replacement claim' to 'measured-negative as shipped'; the live question is now harness-forced usage rather than whether structure beats embeddings unforced, which is precisely Scott's stratify/forcing doctrine. Caveats keep this from closing the case: 40 questions only, blind-judge calibration unverified, a 10K hook-output limit reduced tool outputs to ~2KB previews (integration confound), and Benzi itself wasn't among the tested tools.
2026-10-06T23:36:37Z
evidence attached: reddit.post.1wzdfam β 1,500-run controlled benchmark finds code-graph tools add cost with no accuracy gain and agents call them ~1% of the time β direct counter-evidence to structure-resolved code intelligence displacing grep/Read.
2026-10-04T15:31:42Z
The comment periphery widened but only sideways: sibling MIT tools surfaced (tirth8205/code-review-graph, colbymchenry/codegraph with one user's daily-use testimony for blast-radius queries, DeusData/codebase-memory-mcp) plus the standard 'isn't this just LSP servers' objection β all reinforcing that this is a crowded, contested category without touching the case's actual unknowns. The velocity spike was a transient +5-point bump on the second post with zero new comments, already decayed to 0.0 pts/h; no benchmark, no Benzi adoption, so the meaning is unchanged: cold corroborated pattern watch, escalate only on independent measurement or notable adoption.
2026-10-03T17:34:39Z
grounded: converges/medium β A second independent builder (MIT, fully-local call graph, motivated because 'grep finds mentioners, not actual callers') plus the Codebase-Memory arXiv adoptio
2026-10-03T17:25:27Z
A second, unrelated builder shipping an MIT, fully-local call-graph tool for Claude Code lifts the case from a lone, poorly-received vendor claim to a corroborated emerging pattern: independent implementers are converging on structure-resolved indexes over grep/embedding retrieval. What stays uncorroborated is the claim itself β no independent benchmark, no adoption of Benzi, and its launch reception (27% upvote ratio, ad/security accusations) was hostile, while the new tool differentiates on exactly Benzi's weaknesses (local, permissive license).
2026-10-03T17:23:53Z
evidence attached: reddit.post.1wws6o6 β Independent second implementation (MIT, fully local call-graph for Claude Code instead of grep) corroborates structure-resolved code intelligence as an emerging pattern.
2026-09-27T04:32:09Z
grounded: converges/medium β Benzi independently arrives at the architecture Scott's canon already argues β symbol-addressed navigation over blind embedding queries (Walk a Wiki's named-lin
2026-09-27T04:24:30Z
case created β Released first-party artifact with a distinctive, resolvable claim that resolved code structure can replace embedding-RAG for coding-agent context; seeded low given a single low-engagement post, and the scout's SWE-bench numbers (78.2% vs 73.7%) appear nowhere in the evidence so the hypothesis was rewritten to only what the post establishes.