2026-10-11 17:14 UTC

kapa.ai claims its released Company Knowledge Bench — 1,000 eval cases from real production queries showing frontier-model agent+grep matching a tuned retrieval pipeline at 0.61 (five times slower) and its optimized agentic retriever leading at 0.65 — makes agentic retrieval the winning pattern for messy enterprise knowledge and the benchmark the reference for measuring it; external citation, adoption, or replication of the finding resolves it, silence confirms it as one biased vendor post.

state: seedheat: lowuncertainty: mediumconvergesscott: highagent-evaluation rag-benchmarks agent-retrieval enterprise-knowledgekapa.aiFinn Bauer

What is this?

kapa.ai — an enterprise AI knowledge-assistant vendor, with Finn Bauer listed as a key figure (the supplied snippets don't establish his role) — has released 'Company Knowledge Bench': 1,000 evaluation cases built from real production queries, with agent-generated labels validated by ~170 hours of human review. Its headline result: a frontier-model agent running raw grep/lexical search matches a tuned retrieval pipeline at 0.61 (at roughly five times the latency), while kapa's own optimized agentic retriever tops the leaderboard at 0.65 — which is exactly why the hypothesis sets external citation or replication as the resolution test. Importantly, none of the supplied web results directly cover this release; they map the surrounding 2026 debate, which is demonstrably live and contested: tuned BM25 plus a strong agent loop matching dense retrieval on BrowseComp-Plus (83.1%/94.7%) while collapsing to 3.86% with a weaker agent, an 'Is Grep All You Need?'-inspired factorial benchmark isolating retriever × delivery × harness, and agentic-over-fixed-pipeline architectures showing large gains on hard multi-hop queries at 2–10× cost. So the field-level question kapa is answering is real, but no third-party pickup of the benchmark itself is visible in the supplied material.

Why it matters to Scott

kapa.ai — an enterprise-RAG vendor with every incentive to defend tuned retrieval pipelines — publishes first-party evidence that an agent running raw grep matches its tuned pipeline (0.61 at ~5× latency), independently arriving at the embeddings-demotion position Scott's wikis already hold (search-what-churns, advisory embedding recall, the no-vector-DB dev-wiki): a dated receipt for the 'RAG was built for chatbots' thesis in the enterprise domain, strengthening the convergence Cursor's retreat and the FRAMES results started. The agent-generated labels graded against agent-driven retrievers hand him a characteristic correlated-checkers critique, and a leaderboard where the vendor's own optimized retriever wins is textbook evidence-class-ladder material — so this is both receipt and critique opportunity, not mere illustration.
ip:concept.search-what-churnsdev:concept.advisory-embedding-recalldev:concept.llm-navigated-wikiip:source.rag-was-built-for-chatbots-agents-need-a-wiki-ebookip:concept.independent-vendor-convergenceip:concept.correlated-checkers-pitfallip:concept.evidence-class-ladderradar:agentic-retrieval-frames-benchmarkradar:cursor-sqlite-doc-reconstructionradar:pageindex-vectorless-ragradar:concept.ragradar:concept.benchmark-integrity
queries asked of Scott's wikis
  • grep ripgrep agent retrieval vs vector embeddings
  • Cursor embedding retreat retrieval without vectors
  • LLM-generated eval labels judge validation methodology
  • agentic retrieval vs tuned RAG pipeline benchmark
  • retrieval strategy agent-maintained wiki memory
  • vendor-published benchmark conflict of interest

Measured heat

now 0 pts/hpeak 13 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 242h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-01 14:00⭐ origin echo-reconstructedIntroducing Company Knowledge Bench: 1,000 eval cases built from real production queries with agent-generated labels validated against 170 h
Finn Bauer (kapa.ai) on blog (echo) · attributed from hn.story.49933381
—
10-02 13:37first on hacker news · published · +23.6hBenchmarking retrieval for agents on messy real-world company knowledge
emil_sorensen
—
10-02 13:37amplified on hacker news 👑hn.story.49933381
emil_sorensen
peak 27 · 3 comments · 100% of case engagement
10-02 14:20our radar first saw it · +24.4hdiscovery anchor: hn.story.49933381—
pace: p56 vs 1188 stories at the 168h mark (now 242h old) — ahead of claude-auto-mode-classifier-outage (1.0x), behind anthropic-opus55-cache-read-repricing (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnBenchmarking retrieval for agents on messy real-world company knowledge
Retrieved article excerpt

Open article · Retrieved 2026-10-02T14:26:21.768740+00:00

# Benchmarking retrieval for agents on messy real-world company knowledge

Introducing Company Knowledge Bench

Oct 2, 2026

Finn Bauer

How teams do retrieval is changing fast, in every domain. Cursor recently stopped [searching your code via embeddings](https://forum.cursor.com/t/what-do-you-think-about-cursor-removing-the-codebase-indexing-settings/165899) in favour of relying only on grep and indexed search. Others are swapping their traditional rerankers for new models like [Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev).

One of the most important kinds of knowledge that agents rely on is company knowledge: documentation, tickets, chat messages, internal wikis, and code. Yet most teams cannot tell which retrieval works best on it, because they have no way to measure how their retrieval performs on real data.

Kapa is a platform for indexing your company knowledge and letting your agents search them for context. You connect your sources, Kapa turns them into one searchable knowledge base, and any agent can query it for the information it needs. Retrieval is the core of what we build, and we change how we do it constantly.

The **Company Knowledge Bench** is what we built for ourselves to improve our own system: 1,000 eval cases annotated from real production data. In this post we explain how it works, and score a few common retrieval implementations on it.

Company Knowledge Bench results: score against time per query

Fixed retrieval, agentic grep and Kapa: 7 retrievers on 1,000 eval cases. Higher and further left is better.

Fixed retrieval (traditional RAG)Agentic retriever with grepKapaBest score for the time

Retrieval score0.40.50.60.70 s5 s10 s15 sTime per query, secondsHybrid search 0.41Hybrid search + rerank 0.50Query decomposition 0.56Kapa Default 0.61Kapa Deep 0.65Agent + grep (Luna) 0.54Agent + grep (Sol) 0.61

The takeaway: a frontier model with nothing but grep matches a tuned modern retrieval pipeline at 0.61, but takes five times as long. An optimized agentic retriever (Kapa Deep) does better still: 0.65 in about five seconds.

A note on bias: retrieval is core to our product, and all seven retrievers were built by us and use documents ingested with our pipeline.

## Why public benchmarks do not work for us

Public benchmarks do not work for us because none of them cover all the use cases and types of queries that we see.

**Use cases**

Teams index their sources in Kapa and connect an agent to the retrieval, and that agent can serve quite different use cases, each with its own documents and its own people asking questions:

- **Developers** ask about a product over its docs, API specs, code and GitHub issues, often through Claude Code.
- **Sales and other employees** ask about products, processes and customers over Slack, Confluence, Notion and Google Drive.
- **Support teams** draft replies to new tickets from old tickets, help center articles and internal handbooks.

**Queries**

As these use cases show, many different kinds of agents send queries to Kapa’s retrieval, and the queries look different depending on who wrote them:

- **People** put everything into one message: several questions, a pasted error, a reference to something said earlier.
- **Agents** like Claude Code write their own, breaking a complex question into short, precise searches, like `webhook retry backoff config`

Public retrieval benchmarks are mostly too narrow, built around one domain like law or medicine, or too artificial, built from synthetic documents and questions. None of them cover this range, so we built our own.

## How we score retrieval for agents

Before we can measure good retrieval, we have to define what it is.

A few terms first. A **query** is what is sent to the retriever. The **corpus** is everything a team has indexed in Kapa. A **chunk** is a short piece of it, such as a section of a page. A **retriever** takes a query and returns the most relevant chunks from the corpus for it.

Our definition:

> The retriever’s goal is to collect a minimal set of chunks that completely answers the query, using the highest-quality sources available.

The Company Knowledge Bench rests on three properties:

- **Completeness.** The agent you hand the chunks to needs nothing else to answer the query. If the set is incomplete, the agent gives an incomplete or incorrect answer.
- **Minimality.** Removing any chunk would make the set incomplete. Extra chunks the query does not need cost money and make it harder for the model to reason.
- **Source preference.** A corpus often holds several sets of chunks that could answer the same query, and they are almost never of equal quality. Our benchmark only accepts the preferred ones. Deciding which source is preferred is hard and has a lot of grey zones, but in general it comes down to two factors:

  - **Source authority.** A dedicated reference page outranks a tutorial, an issue thread, or a blog post that repeats the same fact.
  - **Currency.** A current source outranks an outdated one. A Slack thread from last week beats a Confluence page last edited three years ago.

Beyond these, the benchmark enforces rules for specific situations. For example:

- **The query is ambiguous.** The retriever has to return chunks for every reasonable reading of it. Picking one reading is the agent’s job.
- **A problem has several valid solutions.** The retriever has to return chunks for each documented solution, so the agent can explain the options and their trade-offs.
- **The corpus holds nothing that answers the query.** The retriever has to return the strongest evidence there is: a statement that the feature is not supported, a complete list the feature is missing from, or a documented workaround. If none of those exists, the right result is nothing at all.

All of it is written down precisely in our benchmark labelling handbook.

## What an eval case looks like

Our benchmark is a set of eval cases. Each one consists of a real production **query**, a snapshot of the **corpus** as it was when the query was asked, and a **retrieval criterion**: a boolean expression that specifies which chunks from the corpus are valid to retrieve for the query.

Here is a simplified eval case, scored against the chunks one retriever returned:

**Query**

```
"Which plans include SSO, and how do I turn it on?"
```

```
"Which plans include SSO, and how do I turn it on?"
```

```
"Which plans include SSO, and how do I turn it on?"
```

```
"Which plans include SSO, and how do I turn it on?"
```

**Retrieval criterion**

```
{
  "operator": "AND",
  "operands": ["chunk_1", "chunk_2"]
}
```

```
{
  "operator": "AND",
  "operands": ["chunk_1", "chunk_2"]
}
```

```
{
  "operator": "AND",
  "operands": ["chunk_1", "chunk_2"]
}
```

```
{
  "operator": "AND",
  "operands": ["chunk_1", "chunk_2"]
}
```

**Chunks the retriever returned**

```
chunk_1  Pricing page
         "SSO is included in the Team and Enterprise plans."

chunk_2  SSO setup guide
         "To enable SSO, go to Settings > Security, choose your identity
         provider, and paste in its metadata URL."

chunk_3  Community forum post
         "You get SSO on the Team and Enterprise plans."

chunk_4  Changelog
         "Dark mode is now available in the dashboard."
```

```
chunk_1  Pricing page
         "SSO is included in the Team and Enterprise plans."

chunk_2  SSO setup guide
         "To enable SSO, go to Settings > Security, choose your identity
         provider, and paste in its metadata URL."

chunk_3  Community forum post
         "You get SSO on the Team and Enterprise plans."

chunk_4  Changelog
         "Dark mode is now available in the dashboard."
```

```
chunk_1  Pricing page
         "SSO is included in the Team and Enterprise plans."

chunk_2  SSO setup guide
         "To enable SSO, go to Settings > Security, choose your identity
         provider, and paste in its metadata URL."

chunk_3  Community forum post
         "You get SSO on the Team and Enterprise plans."

chunk_4  Changelog
         "Dark mode is now available in the dashboard."
```

```
chunk_1  Pricing page
         "SSO is included in the Team and Enterprise plans."

chunk_2  SSO setup guide
         "To enable SSO, go to Settings > Security, choose your identity
         provider, and paste in its metadata URL."

chunk_3  Community forum post
         "You get SSO on the Team and Enterprise plans."

chunk_4  Changelog
         "Dark mode is now available in the dashboard."
```

**Score**

- **Pass.** The criterion requires chunk\_1 and chunk\_2, and both were returned.
- **Precision: 2 of 4, or 50%.** chunk\_3 and chunk\_4 are not in the criterion.

The forum post is not in the criterion because of source preference. It answers the same part of the query as the pricing page, which plans include SSO, and the pricing page has more source authority. The changelog is simply not relevant. Both count against precision, but for different reasons: one is irrelevant, the other is relevant but not preferred.

In total, we created **1,000 eval cases** for the Company Knowledge Bench, and they are what the results later in this post are based on.

## How we created the benchmark

Labelling a thousand eval cases by hand is not feasible. Sampling them properly makes it even harder: the queries come from every type of source we ingest, internal as well as public corpora, and the industries we serve. So to achieve this scale, the labelling has to be done by agents.

Of course, letting agents create the benchmark only works if they do it correctly. A wrong criterion means a wrong score for every retriever we test against it. So before the agents can build our benchmark, we need a second benchmark that measures how well they build it: a benchmark for benchmark creation.

That second benchmark is a set of 170 eval cases that humans labelled completely by hand, following the same benchmark labelling handbook. Each has a real query, a corpus, and a retrieval criterion a person wrote after working through the corpus themselves, so together they encode the handbook’s rules.

So the process was: build agents that can do the same task the humans did for this set, run them on its queries, and compare the criteria they write with the human ones. We kept improving the agents until they reached a high enough agreement rate with the human labellers. Only then did we let them generate the 1,000 eval cases of the benchmark from production queries.

The labelling is split between two kinds of agents:

- **Candidate agents** scan the full corpus for chunks that might belong in the retrieval criterion. They are built for recall and return a large set of possibly relevant chunks.
- **Criteria agents** take those candidates and turn them into the actual retrieval criterion, following the rules in the benchmark labelling handbook.

The agents do not have to reproduce the human criteria exactly for the benchmark to give a good signal, but they do have to come close. Getting there took many iterations, on the agents and on the handbook itself. We made sure the agents reached sufficient agreement with the human labellers, and validated them extensively, before moving on to generate the benchmark.

The result is labels that are close to what a human would write, and unlike hand labels they scale well beyond the 1,000 eval cases we generated for this benchmark.

## How the retrievers compare

We ran seven retrievers on the Company Knowledge Bench. All of them search the same corpora, ingested and chunked the same way, so the only thing that differs is the retrieval strategy.

### Fixed pipelines

These are the traditional RAG approaches of the last few years. They run the same steps for every query, with no model deciding what to do next.

- **Hybrid search.** The baseline most RAG pipelines start from. The query is matched against a keyword index and an embedding index at once, using `gemini-embedding-001` for the embeddings, and th
emil_sorensen273
🟧 echo.blog ⭐Introducing Company Knowledge Bench: 1,000 eval cases built from real production queries with agent-generated labels validated against 170 hFinn Bauer (kapa.ai)——

Interpretation history

Decision trace