Retrieved article excerpt
Open article · Retrieved 2026-09-18T15:23:03.779047+00:00
Anchor
**The institutional memory of every organization**
Open protocol for organizational semantic models
[License: Apache 2.0](https://camo.githubusercontent.com/1eba057adc6218aed457b0a4f66c89c7c6ad66a0be6455df0a554ab82ebf78d2/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6963656e73652d4170616368655f322e302d626c75653f7374796c653d666f722d7468652d6261646765)
[TypeScript](https://camo.githubusercontent.com/2ab4127dedd4ccf7d3dbb0d440cacca4f720cbb4f06c74350a37f6925b828523/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f547970655363726970742d3331373843363f7374796c653d666f722d7468652d6261646765266c6f676f3d74797065736372697074266c6f676f436f6c6f723d7768697465)
[model.yaml v1](https://camo.githubusercontent.com/6e3f09f4dbe91bb99887508eb34189a21c1ee23aa20a5205099afa182b013488/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6d6f64656c2e79616d6c2d76312d4342333833373f7374796c653d666f722d7468652d6261646765)
[MCP](https://camo.githubusercontent.com/3750945cdeeb0160fbd475d5762fb766e346420c8b3e84435549245f5c29de76/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4d43502d737464696f2d3030303030303f7374796c653d666f722d7468652d6261646765)
Every organization runs on data spread across systems that were never built to share a vocabulary. ERP, exports, spreadsheets, and documents each tell a partial story; without a shared layer of meaning, humans argue over definitions and agents invent new ones every session.
Anchor does not move data or replace systems. It builds the **ontology layer** above them — the same primitive enterprise platforms treat as foundational: map sources to **entities**, wire **relations**, capture **business definitions**, and govern what is true with provenance and confidence. That layer *is* institutional memory when it is written down, versioned, and shared.
The output is `model.yaml`: a committable semantic model. Humans confirm what the machine is unsure about through risk-ranked review; agents query what has been confirmed through MCP.
---
## Why
Organizations have data everywhere and meaning nowhere. Three systems disagree on customer count because *customer* was never defined — not in the database, but in the ontology that should sit above it. Anchor brings that layer within reach for ordinary organizations: local-first, evidence-backed, and small enough to stay true.
---
## How it works
Anchor keeps three questions separate: **what the data shows**, **what it means**, and **what the organization has agreed is true**. Evidence is computed locally and reproducibly. Meaning is inferred, but only from compressed statistics. Truth is decided by people, and recorded with the reasoning behind it.
### Architecture
```
flowchart TB
subgraph boundary["Your infrastructure"]
direction TB
sources["Sources"]
evidence["Evidence"]
proposal["Proposal"]
review["Review"]
model["Semantic model"]
mcp["MCP interface"]
sources --> evidence --> proposal --> review --> model --> mcp
end
inference["Inference endpoint"]
agents["Agents"]
proposal -.->|compressed statistics| inference
inference -.->|structured proposal| proposal
mcp --> agents
classDef external stroke-dasharray: 5 4
class inference,agents external
style boundary fill:none
```
Loading
Everything on the solid path runs where your data already lives. The dotted path is the only network call in the pipeline, and it carries column names, types, distributions, and patterns — never rows, cell values, or document text.
### Stages
| Stage | Function | Inference |
| --- | --- | --- |
| **Ingest** | Normalize sources into a queryable local snapshot | None |
| **Classify** | Assign document types from workspace naming rules | Ambiguous files only |
| **Extract** | Derive structured mentions and facts from documents | None |
| **Index** | Embed document chunks for semantic search | Embeddings only |
| **Profile** | Compute statistical evidence per table and column | None |
| **Propose** | Derive entities, properties, relations, and business rules | Structured tables |
| **Review** | Arbitrate uncertain inferences | None |
| **Serve** | Answer ontology queries | None |
Six of the eight stages involve no inference at all. Ingest resolves encodings, delimiters, regional number formats, and nested archives without a model call. Profiling derives null rates, distinct counts, value patterns, candidate keys, and the cross-table value overlap that surfaces foreign-key candidates.
### Design invariants
These hold on every run and are enforced in code, not by convention.
| Invariant | Guarantee |
| --- | --- |
| **Read-only sources** | Anchor never writes to, moves, or mutates the files you point it at |
| **Data locality** | Inference sees compressed statistics only — never rows, cell values, or document text |
| **Full traceability** | Every entity, property, relation, and rule records the source table, column, and justifying evidence |
| **Validated state** | Pipeline state is read and written through versioned schemas; malformed state fails closed |
| **Inert caching** | Responses are cached under a content hash of model, prompt, and schema — caching changes cost, not results |
| **No silent uncertainty** | Anything below threshold becomes a review question or a recorded doubt, never an unannounced fact |
### Governance
Confidence decides whether a person is asked. The answer decides what is written.
| Confidence | Behaviour | Outcome |
| --- | --- | --- |
| At or above `0.95` | Accepted without a question | Confirmed |
| Between `0.7` and `0.95` | Raised for review | Confirmed, renamed, or removed |
| Below `0.7` | Recorded as a doubt or question | Proposed, or removed on rejection |
Reviewers answer Yes, No, or Rename. A rejection removes the element outright — it never reaches the model. Both thresholds are configurable per workspace.
### Change management
Every source file is fingerprinted by content hash. Re-running the pipeline reprocesses only what changed; unchanged files cost nothing. Consecutive runs can be compared to show what moved in the ontology — entities added, relations dropped, confidence shifted — so model drift stays reviewable instead of invisible.
### Deployment
**Local-first.** One workspace per organization or project. The model, the data snapshot, and every run artifact stay on the machine. Serving the model over MCP requires no account and no network.
**Hosted and ephemeral.** The same pipeline runs as a multi-tenant service where uploaded bytes are garbage-collected after each run, only derived artifacts persist, and every deletion is written to an append-only audit log. See [Hosted deployment](https://github.com/trybacked/anchor#hosted-deployment).
---
Commands, configuration, and artifacts: [Operational workflow](https://github.com/trybacked/anchor#operational-workflow) · [The model](https://github.com/trybacked/anchor#the-model).
---
## Install
**Requirements:** Node.js ≥ 22 · pnpm · [Vercel AI Gateway](https://vercel.com/ai-gateway) API key
```
git clone https://github.com/trybacked/anchor.git
cd anchor && pnpm install && pnpm build
cd apps/cli && pnpm link --global
```
Create `.env` in your **workspace root** (the folder containing `.backed/`, or any parent of your cwd — Anchor walks up to find it):
```
AI_GATEWAY_API_KEY=... # required
REVIEW_CONFIDENCE_THRESHOLD=0.95 # optional
# SEMANTIC_MODEL=zai/glm-5.3-flash
# SEMANTIC_EMBEDDING_MODEL=openai/text-embedding-3-small
```
---
## Operational workflow
The local path: one workspace folder per organization or project, driven by the `backed` CLI. Configuration lives in that folder, and inference calls go only to the endpoint you set in `.env`. To run the same pipeline as a service instead, see [Hosted deployment](https://github.com/trybacked/anchor#hosted-deployment).
### 1. Initialize (`backed init`)
Run once per folder (or again to change document rules):
```
mkdir -p sources
backed init
```
**Prompts:** optional document filename rules (explained in the wizard) → sources folder.
**Output:** `.backed/config.yaml`. Edit before `backed model` to fine-tune rules.
Each PDF filename becomes a **slug** (lowercase, punctuation → underscores). Rules check whether the slug **contains** your keyword.
```
sourcesDir: ./sources
documentTypeHints:
# keyword in filename slug → type id + display name (no LLM when confidence ≥ 0.85)
- match: invoice # matches invoice_acme_2026.pdf, acme_invoice_q1.pdf, …
documentType: invoice # id in model.yaml / MCP
documentTypeLabel: Invoice # label in review
confidence: 0.95
- match: inv # second keyword, same type — add one rule per keyword
documentType: invoice
documentTypeLabel: Invoice
confidence: 0.95
- match: notice
documentType: notice
documentTypeLabel: Notice
confidence: 0.9
```
Example: `public_notice_board.pdf` → slug `public_notice_board` → matches `notice`. Confidence ≥ 0.85 → deterministic (no LLM). No match → one LLM call per file. **No built-in rules at runtime** - only `config.yaml`. Empty list = LLM for every document.
### 2. Build (`backed model`)
```
backed model # sources from config
backed model ./exports # override sources (updates config)
backed model --full # re-infer everything
```
| Stage | When | LLM? | Output |
| --- | --- | --- | --- |
| Ingest | Always | No | `.backed/data.duckdb` |
| Documents | PDF/TXT/DOCX | Ambiguous files only | `documents.json` |
| Mentions + facts | Documents | No | `document_mentions`, `document_facts`, `entity_profiles` in DuckDB |
| Chunk + embed | Documents | Embeddings only | vectors in DuckDB |
| Profile | Always | No | `profile.json` |
| Proposal | Always | Structured tables | `proposal.json` |
LLM responses are cached on disk in `.backed/cache/llm/`, keyed by model + prompt + schema. The cache only reduces cost and latency — it never changes validated outputs. Delete the folder or run `backed model --full` to re-infer from scratch.
Fact extraction is deterministic and runs during `backed model`. Upgrading `@backed/semantic` does not mutate an existing snapshot — **re-run** `backed model` on workspaces that already have document corpora when fact parsing improves.
**PDF-only folders:** `doc_`\* tables get deterministic ontology (no column-classification LLM). **Mixed folders:** CSV gets LLM ontology; documents stay deterministic.
Requires `AI_GATEWAY_API_KEY` — see [Install](https://github.com/trybacked/anchor#install).
### 3. Review (`backed review`)
Confirms or rejects proposals → writes `model.yaml`.
### 4. Serve (`backed serve`)
Authenticated MCP stdio — five deterministic operations on `model.yaml` (see [MCP surface](https://github.com/trybacked/anchor#mcp-surface)).
### 5. Data changes
```
backed model && backed diff
```
### Cheat sheet
```
backed init && backed model && backed review && backed serve
```
Agent pattern: `list_entities` → `get_entity` → `search_model` / `get_definition`.
---
## Hosted deployment
The service path: the same pipeline exposed as a multi-tenant HTTP API, for organizations that submit a corpus rather than run a CLI. Partners upload files, the service infers the model, and the uploaded bytes are destroyed when the run completes.
### How it differs from the local path
| Aspect | Local workflow | Hosted deployment |
| --- | --- | --- |
| Interface | `backed` CLI | HTTP API and TypeScript SDK |
| Tenancy | One workspace per folder | Many isolated tenants per instance |
| Raw files | Stay on your machine | Deleted after every run |
| Review | Interactive prompts | API-driven |
| Audit | Run artifacts on disk | Append-only deletion log and content ledger |
### Ephemeral by construction
Each tenant workspace separates processing from persistence:
```
tenants/<tenantId>/
├── work/ ← uploaded bytes and pipeline scratch, garba