2026-10-11 17:14 UTC

Zep claims its production Konig data plane maintains sub-100ms p95 retrieval across thousands to tens of millions of independently governed memory graphs while tiering idle graphs into object storage, potentially making agent-memory costs track activity rather than provisioned capacity.

state: seedheat: mediumuncertainty: mediumconvergesscott: lowagent-memory knowledge-graphs ai-infrastructureZepDaniel Chalef

What is this?

The case describes Konig as Zep’s production graph-database data plane for agent memory, with claimed sub-100ms p95 retrieval across millions of separately governed graphs and object-storage tiering for idle graphs. The sole supplied web snippet is a discussion of Zep’s temporal knowledge-graph approach, not a Konig announcement or benchmark; its speaker explicitly says they have not deployed Zep at huge scale. Consequently, the snippets do not establish Konig’s production status, performance, isolation or tiering behavior, Daniel Chalef’s role, or the proposed activity-based cost advantage.

Why it matters to Scott

At the architectural level, Zep’s claimed graph-memory service converges with Scott’s durable, queryable memory tier in “Wiki Is the Kernel,” but does not establish his distinctive agent-maintained wiki approach or an actionable improvement to OpenClaw’s graph memory. The potentially consequential delta is governed graph storage whose costs track activity; the supplied evidence does not establish that economics, isolation or latency claim, and the radar hits track adjacent developments rather than Konig itself.
ip:framework.wiki-is-the-kerneldev:project.openclawradar:concept.agent-memoryradar:concept.knowledge-graphsradar:polign-stateless-agent-memoryradar:verity-permission-aware-agent-memory
queries asked of Scott's wikis
  • agent memory graph retrieval latency at scale
  • persistent memory infrastructure activity-based costs
  • multi-tenant memory isolation and per-user governance
  • cold memory object storage tiering retrieval tradeoffs
  • agent-maintained wikis versus temporal knowledge graphs

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 1106h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

08-26 14:00⭐ origin echo-reconstructedZep describes Konig as its production graph-memory data plane, reporting p95 retrieval below 100ms at tens of millions of graphs, per-graph
Daniel Chalef on blog (echo) · attributed from hn.story.49712610
—
09-15 13:57first on hacker news · published · +480.0hWe built a graph database service for agent memory
roseway4
—
09-15 13:57amplified on hacker news 👑hn.story.49712610
roseway4
peak 2 · 0 comments · 98% of case engagement
09-15 14:21our radar first saw it · +480.4hdiscovery anchor: hn.story.49712610—

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnWe built a graph database service for agent memory
Retrieved article excerpt

Open article · Retrieved 2026-09-15T14:22:34.108966+00:00

Featured

# Why we built a graph database service for agent memory

Konig, Zep's graph database service, is the data plane beneath our agent memory platform. This post covers why we built it, how it works, and what we learned along the way.

- [Daniel Chalef](https://blog.getzep.com/author/daniel/)

#### [Daniel Chalef](https://blog.getzep.com/author/daniel/)

27 Aug 2026
• 15 min read

Why we built a graph database service for agent memory

## Key takeaways

- Governed agent context at scale is a workload general-purpose graph databases handle poorly. Konig is Zep's purpose-built graph database for agent memory: millions of knowledge graphs, one per user, team, or project, most in cold storage and all temporal and governed.
- Retrieval latency holds near-constant as the number of graphs grows: p95 retrieval stays under 100ms from a thousand graphs to tens of millions, and Zep's full end-to-end retrieval stays under 200ms at that scale. Cost tracks activity, not capacity: hot graphs serve from RAM and idle graphs are evicted to object storage.
- A single query fuses vector, full-text, graph, and pattern signals into one ranked answer. Graph analytics that normally run offline as batch jobs, like PageRank, run inline in milliseconds.
- Every graph has its own full-text and vector indexes inside its snapshot: exact kNN with SIMD kernels by default, an IVF index with SPFresh/LIRE incremental maintenance as graphs grow. There's no shared search cluster to operate.
- Governance and security are built into the data plane: attribute-based access control (ABAC) on every node and edge, per-graph isolation, customer-managed keys (CMEK/BYOK), and bi-temporal facts with provenance to source.

Agents act on context: what they know about the users they serve, the business they operate in, and the work they have already done. That context is scattered across sources. The churning account, the unpaid invoice, and last week's difficult call are the same customer, but each fact is in a different system, and no one system records the connection. Unifying context from across these sources offers agents insights and efficiencies that tool calling does not. The unified data is entities and the relationships between them.

A graph is the right structure for this data. Relationships are first-class, so traversal replaces the joins a relational or document store would require. Provenance is built in: each fact is an edge that references the source data it was extracted from, so tracing why an agent holds a belief is a traversal. Time fits the model as well, since an edge can store when a relationship held in the world and when the system recorded it.

Unifying business data across sources models well on a graph.

Zep builds a specialized version of these graphs, *one per subject*: the user, customer, team, or project an agent's task concerns. A single customer can have millions of them.

Graphs are difficult to build and serve at scale. That problem is why we built a database service.

The first version of Konig was a holiday project. After several challenging months of fighting outages and poor performance with our production graph database, I prototyped a graph engine over December and January. It was built on in-memory adjacency lists and showed it could meet our scale needs. Over the first half of this year more of the team joined, and their graph algorithm, search, and distributed systems experience turned the prototype into the production system this post describes.

## The requirements

Zep's customers are enterprises and fast-growing AI-native startups. A single customer has many subjects for its agents to work on, many sources of data about each, and many agents acting on that data. A single deployment maintains separate context for thousands or millions of subjects at once. Many deployments run in regulated industries, where an incorrect answer can become a compliance incident.

As a result, our requirements for Konig included:

1. **Scale:** millions of graphs, each with its own evolving memory, and retrieval latency that holds as graph size and graph count grow.
2. **Governance:** access control, retention, provenance, and audit, applied via policy down to nodes and edges.
3. **Security:** isolation between graphs and customer-controlled encryption, in a deployment model that matches the customer's compliance requirements.
4. **Cost:** the system should not keep millions of idle graphs in memory or on cluster storage to cover a peak that rarely occurs.

## Why we built a graph database service

We did not set out to build a graph database. Zep ran on an existing commercial database, and for a long time that worked. As we scaled, it stopped working.

The database failed under load, and adding hardware stopped helping because the failures were structural. It was slow under our access patterns in ways we could not tune away. It could not express what our customers required: many graphs, per-graph isolation, customer-held encryption keys, or a cost model that tracked activity instead of provisioned cluster capacity.

The database was good technology built for a different problem. General-purpose graph databases often assume one large graph: mostly resident in memory, traversed by complex queries, against a schema fixed in advance. Agent context is the inverse workload: millions of smaller graphs, most cold at any moment, each temporal and governed independently. Graphs have a reputation for being slow and hard to scale. In our case the cause was the database's fit to this workload, not the graph model itself.

Fighting a database built for a different workload made less sense than building for our workload, so we built Konig. It is the data plane on which [Graphiti](https://github.com/getzep/graphiti?ref=blog.getzep.com) and other Zep service components run.

Zep architecture: Konig is the data plane for the Zep service.

## Why a narrow scope was easier to build

Building a general-purpose database is a multi-year undertaking. A purpose-built engine for a single workload is far smaller, both to build and to maintain, because it never has to serve every use case.

The query interface is the clearest example. Konig exposes a gRPC API: mutate, search, traverse, and a small set of lookups. It has no general-purpose query language such as Cypher. That removes a large part of what makes a database hard to build. There is no grammar to parse and no query optimizer to keep fast across arbitrary queries. The parsing and planning layer that consumes significant effort in a general-purpose database does not exist in Konig.

We made the same trade in the infrastructure. Konig runs on our cloud provider's managed services instead of storage and failover primitives of our own: object storage for snapshots, and a managed, multi-AZ datastore for the write-ahead log and metadata. Building those ourselves would have been a major undertaking. Using managed infrastructure shortened the path to production.

## Millions of graphs, not one big one

Konig does not partition one graph across tenants. Each subject, such as a user, project, or team, gets its own graph. This graph is the unit of storage, access, encryption, and retention.

With one graph per subject, isolation is structural. One tenant's context cannot appear in another's results. Governance, encryption, and retention attach to a boundary that already exists, with none of the per-tenant query filters common in SaaS.

Bounding each query to one graph enables otherwise offline algorithms to run in real time. PageRank queries return in low milliseconds. Splitting memory into one graph per subject bounds the working set each query and algorithm touches. These algorithms run inline, per request, and latency stays flat as the graph count grows.

The design gives up global ordering and traversals that span every graph at once. On the real-time retrieval path we do not need them: a single retrieval targets one graph, and an agent reasons across many graphs through separate calls. Some customers do need cross-graph analysis, and we are building it as a separate, non-real-time workload off the low-latency retrieval path.

The Konig architecture, with multiple, tiered layers of graph storage

## Hot, warm, and cold storage

Most graphs are idle most of the time. Konig treats that as the basis of its cost model and keeps each graph in one of three tiers.

Hot graphs live in RAM and serve most queries at microsecond latency. Snapshots are periodically written to both local ephemeral NVMe and object storage. A graph evicted from RAM due to inactivity stays warm while its snapshot remains on local NVMe, where it reloads in low milliseconds. After longer idleness the local copy is pruned and the graph goes cold, leaving only the snapshot in object storage.

A cold graph costs almost nothing to retain and reloads from object storage on its next request. After optimizing how we use our cloud provider's object storage, we've gotten cold load down to low hundreds of milliseconds. Promotion to hot is lazy; no background process warms graphs that are not in use.

Cost tracks active graphs, not total graphs. A deployment with a million graphs and one percent hot pays for one percent of the memory; the rest is in object storage at object-storage prices. This is the data-lake pattern applied to agent context: the working set stays in fast storage and the long tail costs little.

## Scaling by adding shards

Graphs are distributed across shards by rendezvous hashing. Each graph is scored against every shard by hashing its key with the shard's ID, and the highest-scoring shard owns it. Adding or removing a shard therefore relocates only about 1/N of graphs; the rest stay in place. There is no resharding step, no rebalancing window, and no coordinator assigning ownership.

Hash scores determine which shard owns a graph.

A new shard becomes ready almost immediately. On startup it does not bulk-load its assigned graphs or replay their logs. It registers a heartbeat, marks itself ready, and starts with an empty in-memory map. Each graph loads on demand on the first request for it: Konig loads the snapshot, then replays the write-ahead log written since. Adding capacity means adding a node and letting it load its graphs as requests arrive.

This fixes the original failure, where adding hardware stopped helping. Even clustered, the old database kept the full dataset on every node and could not shard the workload across them, so scaling meant larger machines. Konig scales horizontally. There is no cluster topology to design in advance and no peak load to capacity-plan against.

A query touches one graph, so its latency does not depend on the total number of graphs. In production, p95 search latency holds near-constant as the graph count grows. From a thousand graphs to tens of millions, p50 retrieval is unchanged and p95 stays under 100 milliseconds, rising only marginally as the count keeps growing. End-to-end Zep latency is under 200 milliseconds, and the system sustains thousands of mutations and queries per second.

## How a query runs

Reads enter through the typed API and run against the in-memory graph. Entities and edges are held in compact, integer-indexed arrays with adjacency lists, so following a relationship is a pointer dereference instead of an index lookup. Over that representation, Konig combines every relevance signal into a single ranked result.

Konig offers several approaches to context retrieval from the graph. These include lexical relevance via BM25 and vector similarity for semantic meaning. Both come from the per-graph search indexes described in the next section.

An example of a search strategy executed by Konig.

These lexical and semantic search results may function as seeds to graph-structural queries such as BFS (Breadth-First Search) and Personalized PageRank. Doing so optimizes these operations by narrowing them to subgraphs.

For many search op
roseway420
🟧 echo.blog ⭐Zep describes Konig as its production graph-memory data plane, reporting p95 retrieval below 100ms at tens of millions of graphs, per-graph Daniel Chalef——

Interpretation history

Decision trace