2026-10-11 16:37 UTC

Manticore claims its released engine-native chunking embeds and searches long documents beyond model input limits while returning document-level results, eliminating separate chunking pipelines and improving retrieval at increased memory and ingestion cost.

state: watchingheat: lowuncertainty: mediumconvergesscott: mediumrag vector-search document-chunkingManticore SearchDmitrii Kuzmenkov

What is this?

The case describes Manticore Search announcing automatic, table-defined chunking inside its search engine: long documents would be embedded as smaller chunks but returned as document-level results ranked by their nearest chunk. The supplied web snippets explain why chunking is needed—embedding input limits and loss of retrieval specificity—but none covers Manticore's announcement, so they do not verify release availability, ranking behavior, pipeline simplification, or the claimed recall and resource-cost tradeoffs. The case names Dmitrii Kuzmenkov, but the supplied material does not establish his role; the benchmark figure in the evidence title is truncated.

Why it matters to Scott

Manticore’s claimed separation of embedding chunks from returned documents partially converges with Scott’s hierarchical auto-merge retrieval, although nearest-chunk ranking differs from promoting a parent when several matching leaves cluster beneath it. It offers a concrete comparison for his search project’s custom source-native chunking pipeline—whether engine-managed splitting preserves meaning boundaries and provenance while reducing plumbing—but the supplied snippets do not verify availability or the recall/cost claims.
dev:concept.hierarchical-auto-merge-retrievaldev:concept.source-native-semantic-chunkingdev:project.searchradar:concept.retrievalradar:concept.ragradar:concept.vector-search
queries asked of Scott's wikis
  • RAG ingestion engine-native chunking versus separate pipelines
  • document-level retrieval nearest-chunk ranking
  • long-document embeddings truncation retrieval recall evaluation
  • vector indexing memory ingestion cost tradeoffs
  • agent wiki search chunk granularity source context

Measured heat

now 0 pts/hpeak 3 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 650h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-14 14:00⭐ origin echo-reconstructedManticore announces table-defined automatic chunking with document-level nearest-chunk ranking, reporting recall@5 improving from 55.1% to 8
Dmitrii Kuzmenkov on blog (echo) · attributed from hn.story.49738766
—
09-17 10:30first on hacker news · published · +68.5hBetter Vector Search for Long Documents: Chunking Inside Manticore Search
GloriaVinogrado
—
09-17 10:30amplified on hacker news 👑hn.story.49738766
GloriaVinogrado
peak 80 · 14 comments · 97% of case engagement
09-29 08:35amplified on hacker newshn.story.49890003
GloVin
peak 3 · 0 comments · 3% of case engagement
09-17 11:20our radar first saw it · +69.3hdiscovery anchor: hn.story.49738766—
pace: p66 vs 1032 stories at the 336h mark (now 650h old) — ahead of compute-cheap-h100-h200-pricing (1.0x), behind openai-german-wiki-incident (1.0x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnBetter Vector Search for Long Documents: Chunking Inside Manticore Search
Retrieved article excerpt

Open article · Retrieved 2026-09-17T11:21:47.622879+00:00

blog-post

## Better Vector Search for Long Documents: Chunking Inside Manticore Search

[author image](https://manticoresearch.com/author/dmitrii-kuzmenkov/ "Dmitrii Kuzmenkov")

Author: [Dmitrii Kuzmenkov](https://manticoresearch.com/author/dmitrii-kuzmenkov/)  
Published: Sep 15, 2026 - 33 Min read

[View as markdown](https://manticoresearch.com/blog/auto-chunking.md)

Say you are building search over your team's internal documentation — guides, runbooks, postmortems. You have a table with [auto embeddings](https://manticoresearch.com/blog/auto-embeddings/)
: you insert text, Manticore runs the model and fills the vector column for you. (If that is new to you, start with [vector search in Manticore](https://manticoresearch.com/blog/vector-search/)
.) You load a 4,000-word document. The insert succeeds. The search works. Everything looks fine.

Except the model you picked has a 512-token input window, and that document is about 5,000 tokens long. The model read the first 380 words and threw away the other 3,600. Nothing in the document past that point can ever be retrieved, and nothing anywhere told you. The embedding may not represent the document as a whole either.

Until now, you would usually split the document into several pieces yourself, create embeddings for each one, and then work out how to combine the results if you wanted document search rather than chunk search. Manticore now handles this in the table definition: add `chunk_strategy` to the vector column in `CREATE TABLE`, and Manticore splits each document into chunks, embeds every chunk, and searches all of them:

```
DROP TABLE IF EXISTS docs;

CREATE TABLE docs (
  title text,
  content text,
  chunks float_vector_array knn_type='hnsw' hnsw_similarity='cosine'
    model_name='Xenova/all-MiniLM-L6-v2' from='title,content'
    chunk_strategy='sentence' max_tokens='256' overlap_tokens='32'
);
```

That is the whole feature. No ingest pipeline, no splitter library, no second table for chunks, no `GROUP BY` to fold chunk hits back into documents.

## TL;DR

- **Five strategies**: `truncate` (the old default), `mean`, `fixed`, `recursive`, `sentence`. Set with `chunk_strategy` on a model-backed vector column.
- **`truncate` and `mean`** produce one vector per document and work on a `float_vector` column. **`fixed`, `recursive` and `sentence`** produce many, so they need a [`float_vector_array`](https://manual.manticoresearch.com/Creating_a_table/Data_types#Float-vector-array)
  column.
- **A document is still one search result.** Chunks compete individually, and Manticore returns the document once, with `knn_dist()` reporting the distance to its closest chunk. `k` counts documents, not chunks.
- **Tuning knobs**: `max_tokens` (chunk size), `overlap_tokens` (shared tokens between neighbors), `max_chunks` (ceiling per document).
- **Measured on the Manticore manual** (189 pages, ~298k words): for content buried past the model's window, recall@5 went from **55.1% → 83.3%** and MRR from **0.44 → 0.70**, at ~2.5× the RAM and ~4× the ingest time.
- **Queries are never chunked.** A query is short enough to embed as a whole; only stored documents are split.

## The problem, shown with a small example

Suppose you have four documents:

1. **Backup and restore runbook** — about 700 words, roughly 900 tokens. Backup schedules, retention, restore drills, credentials, capacity planning. The *last* section explains how to rotate the TLS certificate used by the replication port.
2. **Monitoring and alerting guide** — unrelated.
3. **Getting started with the CLI** — unrelated.
4. **TLS and certificates for the HTTP API** — a short page that is *entirely* about certificates, and never mentions rotation or replication.

You can create the table and add the documents using the commands below.

So, what we have is: one table, three vector columns with the same source text — one column per strategy. A single `INSERT` fills all three, so the comparison conditions are identical:

```
DROP TABLE IF EXISTS docs;
CREATE TABLE docs (
  title text,
  body text,
  v_truncate float_vector       knn_type='hnsw' hnsw_similarity='cosine'
    model_name='Xenova/all-MiniLM-L6-v2' from='title,body',
  v_mean     float_vector       knn_type='hnsw' hnsw_similarity='cosine'
    model_name='Xenova/all-MiniLM-L6-v2' from='title,body' chunk_strategy='mean',
  v_sentence float_vector_array knn_type='hnsw' hnsw_similarity='cosine'
    model_name='Xenova/all-MiniLM-L6-v2' from='title,body'
    chunk_strategy='sentence' max_tokens='128' overlap_tokens='32'
);
```

Insert the four documents

```
INSERT INTO docs (id, title, body) VALUES
  (1, 'Backup and restore runbook',
   'Nightly backups run at 02:00 UTC from the standby node. The job snapshots every table directory, writes a manifest, and uploads the result to object storage. Retention is thirty daily copies, twelve monthly copies, and one yearly copy. A restore drill runs on the first Monday of each month against a scratch cluster. The drill counts as passed only when a full-text search over the restored data returns the same document count as production. Anything less is treated as a failed drill and investigated the same week. Before a restore, freeze the target cluster so that no writes land while files are being replaced. Copy the manifest first and verify its checksum. If the checksum does not match, stop: a partial restore is worse than no restore, because the cluster will start and silently serve half the corpus. After the files are in place, unfreeze and let replication catch up. Watch the queue depth. If it does not drain within ten minutes, the node is probably still reading from cold storage and needs a warm-up pass before it can serve traffic. Backup failures page the on-call engineer. The three most common causes are an expired object storage credential, a disk that filled up while the snapshot was being written, and a table left frozen by a previous failed run. All three are recoverable without data loss. Check the job log first, then the disk, then the freeze state of every table. Capacity planning for backups is boring but it matters. A daily copy of the search cluster is roughly the size of the data directory plus fifteen percent for the manifest and metadata. Multiply by the retention count, add the transfer cost, and you have the monthly bill. Most teams discover too late that the yearly copies dominate the storage line. Object storage lifecycle rules do most of the retention work. Daily copies move to infrequent access after seven days and expire after thirty. Monthly copies move to archive after sixty days. Yearly copies never expire automatically; deleting one is a manual action that requires a second approver. Credentials for the backup job live in the secret manager and are issued to a role, not to a person. The role can write new objects and list the bucket. It cannot delete, and it cannot read objects older than the current day. That last restriction is the cheapest defence against a compromised backup runner turning into a data exfiltration path. Documentation for each table lives next to its schema: what the table is for, who owns it, how large it is expected to get, and whether it can be rebuilt from an upstream source. A table that can be rebuilt does not need thirty daily copies. Roughly half of most clusters turns out to be derived data that nobody had marked as derived. Verification is not the same as the job exiting zero. The job can succeed while producing an unusable copy: an empty table, a truncated upload, a manifest that references a file that was never written. The verification step reads the manifest back, checks every referenced object exists and matches its recorded size, and compares row counts on three sampled tables against production. Rotating the replication TLS certificate is a separate procedure and the step people most often get wrong. The certificate that secures the replication port is not the same as the one the HTTP API uses, and replacing one does not replace the other. Generate the new key and signing request on the node that will be rotated first, sign them with the cluster certificate authority, and place the files next to the existing ones rather than on top of them. Then update the node configuration to point at the new paths and reload. Do one node at a time and confirm that the cluster reports every peer as synced before moving on. A half-rotated cluster where two nodes trust different authorities will keep accepting writes on both sides and diverge quietly. When every node has been rotated, remove the old key material and revoke the retired certificate at the authority.'),
  (2, 'Monitoring and alerting guide',
   'Every node exports metrics over an HTTP endpoint that a scraper collects once per fifteen seconds. The dashboards are grouped into four rows: traffic, latency, saturation, and errors. Traffic is queries per second broken down by table. Latency is the ninety-fifth and ninety-ninth percentile of query time, measured server side. Alerting is deliberately thin. Paging alerts fire on sustained error rate above one percent for five minutes, on ninety-ninth percentile latency above two seconds for ten minutes, and on a node dropping out of the cluster. Everything else is a ticket, not a page. Teams that page on every anomaly stop reading pages within a month. Log retention is fourteen days hot and ninety days cold. The query log records the query text, the table, the match count, and the elapsed time. Turning it on costs a few percent of throughput and is almost always worth it, because most performance investigations start with a slow query nobody knew was being issued.'),
  (3, 'Getting started with the CLI',
   'The command line client connects over the MySQL wire protocol, so any MySQL client works and you do not need to install anything special. Point it at port 9306 and you get an interactive shell. The shell understands the usual conveniences: history, tab completion of table names, and vertical output when a row is too wide for the terminal. Start by listing tables, then look at one with SHOW CREATE TABLE. The output is the exact statement that would recreate the table, including every option that was applied implicitly, which makes it the fastest way to find out what a table actually does rather than what someone documented two years ago. Bulk loading from the shell is possible but rarely what you want. For anything above a few thousand rows, use the HTTP bulk endpoint or one of the log shipper integrations, both of which batch and retry for you.'),
  (4, 'TLS and certificates for the HTTP API',
   'The HTTP API can be served over TLS. You supply a certificate, a private key, and optionally a chain file, and the listener starts speaking HTTPS instead of HTTP. Clients that present a certificate of their own can be authenticated by it, which is the usual way to lock an internal API down without putting a password in every config file. Certificates for the HTTP API come from wherever your organisation gets certificates: a public authority, an internal authority, or an automated issuer. The file format is PEM. Both the certificate and the key must be readable by the user the server runs as, and the key must not be world readable or the listener refuses to start. Debugging TLS problems is mostly about reading the handshake. A client that reports an unknown authority is missing the chain. A client that reports a hostname mismatch is connecting by an address that is not in the certificate. A client that hangs is usually talking TLS to a plaintext port.');
```

Now ask a question whose answer lives in the runbook's last section, once per strategy:

```
SELECT title, knn_dist() FROM docs
WHERE knn(v_truncate, 4, 'how do I rotate the TLS certificate used for replication');

SELECT title, knn_dist() FROM docs
WHERE knn(v_mean, 4, 'how do I rotate the TLS certificate used for replication');

SELECT title, knn_dist() FROM docs
WHERE k
GloriaVinogrado8014
🟧 echo.blog ⭐Manticore announces table-defined automatic chunking with document-level nearest-chunk ranking, reporting recall@5 improving from 55.1% to 8Dmitrii Kuzmenkov——
🟧 hnFull-Text Search Still Works. It Just Doesn't Get You to an AnswerGloVin30

Interpretation history

Decision trace