2026-10-11 16:38 UTC

Anchor's creator claims its released local-first pipeline converts organizational sources into traceable, reviewable model.yaml semantic models served through MCP, giving agents shared business definitions without sending raw source rows to ontology inference.

state: seedheat: lowuncertainty: highconvergesscott: mediumagent-ontologies knowledge-systems agent-memorylucatropeatrybacked

What is this?

The supplied case describes Anchor as a creator-announced open-source ontology layer that turns organizational sources into reviewable, traceable model.yaml semantic models for agents to access through MCP. It names lucatropea and trybacked, but the supplied search snippets do not establish their roles or independently confirm Anchor's release, local-first implementation, confidence review, or claim that raw source rows are excluded from ontology inference. The search results document adjacent semantic-layer offerings from Credible/Malloy and CorralData, not Anchor itself; Anchor's specific capabilities therefore remain creator claims in this material.

Why it matters to Scott

Anchor’s claimed source-to-reviewed-semantic-model pipeline converges with Scott’s BI for Soft Data architecture and offers a concrete comparison for FDE BI’s reviewed mappings and MCP IP Wiki’s shared knowledge endpoint: whether a model.yaml contract can supply reusable business definitions without exposing raw rows during inference. This warrants implementation review rather than a validation headline: the supplied material establishes neither independent verification nor recovery of Scott’s unstructured causal knowledge, and his Metadata-First, Not Metadata-Safe distinction remains essential when assessing statistics-based disclosure.
ip:framework.bi-for-soft-dataip:concept.metadata-first-not-metadata-safedev:project.fde-bidev:project.mcp-ip-wikiradar:ontology-bootstrap-agent-contextradar:dbctx-postgres-context-compilerradar:nerra-shared-company-context
queries asked of Scott's wikis
  • shared business definitions semantic contracts for agents
  • agent-maintained knowledge provenance human review confidence thresholds
  • local-first knowledge pipelines privacy metadata-only inference
  • MCP semantic models knowledge-system integration
  • ontology induction structured data versus document RAG

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 553h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-18 15:40 (minted)⭐ origin echo-reconstructedAnchor publishes an open semantic-model protocol, local CLI, statistics-based ontology proposals, configurable confidence review, provenance
trybacked on github (echo) · attributed from hn.story.49755135 · published time unknown
—
09-18 14:43first on hacker news · published · lag ?I built Anchor. The open-source ontology layer for AI agents
lucatropea
—
09-18 16:03first on r/ClaudeAI · published · lag ?Power Remote BI MCP x Claude Projects
DevelopmentFormal399
—
09-18 14:43amplified on hacker news 👑hn.story.49755135
lucatropea
peak 1 · 0 comments · 50% of case engagement
09-18 16:03amplified on r/ClaudeAIreddit.post.1wjuaa4
DevelopmentFormal399
peak 1 · 1 comments · 50% of case engagement
09-18 15:20our radar first saw it · lag ?discovery anchor: hn.story.49755135—
pace: p32 vs 1032 stories at the 336h mark (now 553h old) — ahead of addom-local-coding-harness (1.5x), behind agentsec-static-config-auditing (0.8x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnI built Anchor. The open-source ontology layer for AI agents
Retrieved article excerpt

Open article · Retrieved 2026-09-18T15:23:03.779047+00:00

Anchor

**The institutional memory of every organization**  
Open protocol for organizational semantic models

[License: Apache 2.0](https://camo.githubusercontent.com/1eba057adc6218aed457b0a4f66c89c7c6ad66a0be6455df0a554ab82ebf78d2/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6963656e73652d4170616368655f322e302d626c75653f7374796c653d666f722d7468652d6261646765)
[TypeScript](https://camo.githubusercontent.com/2ab4127dedd4ccf7d3dbb0d440cacca4f720cbb4f06c74350a37f6925b828523/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f547970655363726970742d3331373843363f7374796c653d666f722d7468652d6261646765266c6f676f3d74797065736372697074266c6f676f436f6c6f723d7768697465)
[model.yaml v1](https://camo.githubusercontent.com/6e3f09f4dbe91bb99887508eb34189a21c1ee23aa20a5205099afa182b013488/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6d6f64656c2e79616d6c2d76312d4342333833373f7374796c653d666f722d7468652d6261646765)
[MCP](https://camo.githubusercontent.com/3750945cdeeb0160fbd475d5762fb766e346420c8b3e84435549245f5c29de76/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4d43502d737464696f2d3030303030303f7374796c653d666f722d7468652d6261646765)

Every organization runs on data spread across systems that were never built to share a vocabulary. ERP, exports, spreadsheets, and documents each tell a partial story; without a shared layer of meaning, humans argue over definitions and agents invent new ones every session.

Anchor does not move data or replace systems. It builds the **ontology layer** above them — the same primitive enterprise platforms treat as foundational: map sources to **entities**, wire **relations**, capture **business definitions**, and govern what is true with provenance and confidence. That layer *is* institutional memory when it is written down, versioned, and shared.

The output is `model.yaml`: a committable semantic model. Humans confirm what the machine is unsure about through risk-ranked review; agents query what has been confirmed through MCP.

---

## Why

Organizations have data everywhere and meaning nowhere. Three systems disagree on customer count because *customer* was never defined — not in the database, but in the ontology that should sit above it. Anchor brings that layer within reach for ordinary organizations: local-first, evidence-backed, and small enough to stay true.

---

## How it works

Anchor keeps three questions separate: **what the data shows**, **what it means**, and **what the organization has agreed is true**. Evidence is computed locally and reproducibly. Meaning is inferred, but only from compressed statistics. Truth is decided by people, and recorded with the reasoning behind it.

### Architecture

```
flowchart TB
    subgraph boundary["Your infrastructure"]
        direction TB
        sources["Sources"]
        evidence["Evidence"]
        proposal["Proposal"]
        review["Review"]
        model["Semantic model"]
        mcp["MCP interface"]

        sources --> evidence --> proposal --> review --> model --> mcp
    end

    inference["Inference endpoint"]
    agents["Agents"]

    proposal -.->|compressed statistics| inference
    inference -.->|structured proposal| proposal
    mcp --> agents

    classDef external stroke-dasharray: 5 4
    class inference,agents external
    style boundary fill:none
```

 Loading

Everything on the solid path runs where your data already lives. The dotted path is the only network call in the pipeline, and it carries column names, types, distributions, and patterns — never rows, cell values, or document text.

### Stages

| Stage | Function | Inference |
| --- | --- | --- |
| **Ingest** | Normalize sources into a queryable local snapshot | None |
| **Classify** | Assign document types from workspace naming rules | Ambiguous files only |
| **Extract** | Derive structured mentions and facts from documents | None |
| **Index** | Embed document chunks for semantic search | Embeddings only |
| **Profile** | Compute statistical evidence per table and column | None |
| **Propose** | Derive entities, properties, relations, and business rules | Structured tables |
| **Review** | Arbitrate uncertain inferences | None |
| **Serve** | Answer ontology queries | None |

Six of the eight stages involve no inference at all. Ingest resolves encodings, delimiters, regional number formats, and nested archives without a model call. Profiling derives null rates, distinct counts, value patterns, candidate keys, and the cross-table value overlap that surfaces foreign-key candidates.

### Design invariants

These hold on every run and are enforced in code, not by convention.

| Invariant | Guarantee |
| --- | --- |
| **Read-only sources** | Anchor never writes to, moves, or mutates the files you point it at |
| **Data locality** | Inference sees compressed statistics only — never rows, cell values, or document text |
| **Full traceability** | Every entity, property, relation, and rule records the source table, column, and justifying evidence |
| **Validated state** | Pipeline state is read and written through versioned schemas; malformed state fails closed |
| **Inert caching** | Responses are cached under a content hash of model, prompt, and schema — caching changes cost, not results |
| **No silent uncertainty** | Anything below threshold becomes a review question or a recorded doubt, never an unannounced fact |

### Governance

Confidence decides whether a person is asked. The answer decides what is written.

| Confidence | Behaviour | Outcome |
| --- | --- | --- |
| At or above `0.95` | Accepted without a question | Confirmed |
| Between `0.7` and `0.95` | Raised for review | Confirmed, renamed, or removed |
| Below `0.7` | Recorded as a doubt or question | Proposed, or removed on rejection |

Reviewers answer Yes, No, or Rename. A rejection removes the element outright — it never reaches the model. Both thresholds are configurable per workspace.

### Change management

Every source file is fingerprinted by content hash. Re-running the pipeline reprocesses only what changed; unchanged files cost nothing. Consecutive runs can be compared to show what moved in the ontology — entities added, relations dropped, confidence shifted — so model drift stays reviewable instead of invisible.

### Deployment

**Local-first.** One workspace per organization or project. The model, the data snapshot, and every run artifact stay on the machine. Serving the model over MCP requires no account and no network.

**Hosted and ephemeral.** The same pipeline runs as a multi-tenant service where uploaded bytes are garbage-collected after each run, only derived artifacts persist, and every deletion is written to an append-only audit log. See [Hosted deployment](https://github.com/trybacked/anchor#hosted-deployment).

---

Commands, configuration, and artifacts: [Operational workflow](https://github.com/trybacked/anchor#operational-workflow) · [The model](https://github.com/trybacked/anchor#the-model).

---

## Install

**Requirements:** Node.js ≥ 22 · pnpm · [Vercel AI Gateway](https://vercel.com/ai-gateway) API key

```
git clone https://github.com/trybacked/anchor.git
cd anchor && pnpm install && pnpm build
cd apps/cli && pnpm link --global
```

Create `.env` in your **workspace root** (the folder containing `.backed/`, or any parent of your cwd — Anchor walks up to find it):

```
AI_GATEWAY_API_KEY=...                          # required
REVIEW_CONFIDENCE_THRESHOLD=0.95                # optional
# SEMANTIC_MODEL=zai/glm-5.3-flash
# SEMANTIC_EMBEDDING_MODEL=openai/text-embedding-3-small
```

---

## Operational workflow

The local path: one workspace folder per organization or project, driven by the `backed` CLI. Configuration lives in that folder, and inference calls go only to the endpoint you set in `.env`. To run the same pipeline as a service instead, see [Hosted deployment](https://github.com/trybacked/anchor#hosted-deployment).

### 1. Initialize (`backed init`)

Run once per folder (or again to change document rules):

```
mkdir -p sources
backed init
```

**Prompts:** optional document filename rules (explained in the wizard) → sources folder.

**Output:** `.backed/config.yaml`. Edit before `backed model` to fine-tune rules.

Each PDF filename becomes a **slug** (lowercase, punctuation → underscores). Rules check whether the slug **contains** your keyword.

```
sourcesDir: ./sources
documentTypeHints:
  # keyword in filename slug → type id + display name (no LLM when confidence ≥ 0.85)
  - match: invoice # matches invoice_acme_2026.pdf, acme_invoice_q1.pdf, …
    documentType: invoice # id in model.yaml / MCP
    documentTypeLabel: Invoice # label in review
    confidence: 0.95
  - match: inv # second keyword, same type — add one rule per keyword
    documentType: invoice
    documentTypeLabel: Invoice
    confidence: 0.95
  - match: notice
    documentType: notice
    documentTypeLabel: Notice
    confidence: 0.9
```

Example: `public_notice_board.pdf` → slug `public_notice_board` → matches `notice`. Confidence ≥ 0.85 → deterministic (no LLM). No match → one LLM call per file. **No built-in rules at runtime** - only `config.yaml`. Empty list = LLM for every document.

### 2. Build (`backed model`)

```
backed model              # sources from config
backed model ./exports    # override sources (updates config)
backed model --full       # re-infer everything
```

| Stage | When | LLM? | Output |
| --- | --- | --- | --- |
| Ingest | Always | No | `.backed/data.duckdb` |
| Documents | PDF/TXT/DOCX | Ambiguous files only | `documents.json` |
| Mentions + facts | Documents | No | `document_mentions`, `document_facts`, `entity_profiles` in DuckDB |
| Chunk + embed | Documents | Embeddings only | vectors in DuckDB |
| Profile | Always | No | `profile.json` |
| Proposal | Always | Structured tables | `proposal.json` |

LLM responses are cached on disk in `.backed/cache/llm/`, keyed by model + prompt + schema. The cache only reduces cost and latency — it never changes validated outputs. Delete the folder or run `backed model --full` to re-infer from scratch.

Fact extraction is deterministic and runs during `backed model`. Upgrading `@backed/semantic` does not mutate an existing snapshot — **re-run** `backed model` on workspaces that already have document corpora when fact parsing improves.

**PDF-only folders:** `doc_`\* tables get deterministic ontology (no column-classification LLM). **Mixed folders:** CSV gets LLM ontology; documents stay deterministic.

Requires `AI_GATEWAY_API_KEY` — see [Install](https://github.com/trybacked/anchor#install).

### 3. Review (`backed review`)

Confirms or rejects proposals → writes `model.yaml`.

### 4. Serve (`backed serve`)

Authenticated MCP stdio — five deterministic operations on `model.yaml` (see [MCP surface](https://github.com/trybacked/anchor#mcp-surface)).

### 5. Data changes

```
backed model && backed diff
```

### Cheat sheet

```
backed init && backed model && backed review && backed serve
```

Agent pattern: `list_entities` → `get_entity` → `search_model` / `get_definition`.

---

## Hosted deployment

The service path: the same pipeline exposed as a multi-tenant HTTP API, for organizations that submit a corpus rather than run a CLI. Partners upload files, the service infers the model, and the uploaded bytes are destroyed when the run completes.

### How it differs from the local path

| Aspect | Local workflow | Hosted deployment |
| --- | --- | --- |
| Interface | `backed` CLI | HTTP API and TypeScript SDK |
| Tenancy | One workspace per folder | Many isolated tenants per instance |
| Raw files | Stay on your machine | Deleted after every run |
| Review | Interactive prompts | API-driven |
| Audit | Run artifacts on disk | Append-only deletion log and content ledger |

### Ephemeral by construction

Each tenant workspace separates processing from persistence:

```
tenants/<tenantId>/
├── work/      ← uploaded bytes and pipeline scratch, garba
lucatropea10
🟧 echo.github ⭐Anchor publishes an open semantic-model protocol, local CLI, statistics-based ontology proposals, configurable confidence review, provenancetrybacked——
🟠 redditPower Remote BI MCP x Claude Projects
ClaudeAI
DevelopmentFormal39911

Interpretation history

Decision trace