2026-10-11 18:01 UTC

Scry claims its released ClickHouse-backed index exposes internet data as queryable relations, initially documenting Reddit submissions, enabling agents to filter and aggregate source records rather than rely solely on ranked search results.

state: seedheat: lowuncertainty: mediumknownscott: lowagent-search retrieval clickhouseScryXyra

What is this?

The supplied case describes Scry as releasing a ClickHouse-backed internet index, advertised in a Show HN title as 500 TB with congestion pricing, with an initially documented reddit.posts table claimed to cover nearly all Reddit submissions over a specified period. The web results contain no Scry-specific corroboration: they do not establish the release, dataset coverage, pricing, or Xyra’s role. ClickHouse’s own snippets do establish that its text indexes support filtering source rows and combining those filters with SQL aggregation and joins, making the proposed retrieval approach technically plausible without verifying Scry’s implementation or agent capabilities.

Why it matters to Scott

The architectural position is already held in Scott’s RAG/Wiki Substrate Rule and implemented in Search Conversations: retain source records for recall-shaped workloads rather than depend solely on synthesized or similarity-ranked results; the supplied radar hits do not track Scry’s release itself. Scry’s claimed relational internet index is another example of that position, but the uncorroborated coverage and implementation claims establish neither a consequential endorsement nor a concrete change to Scott’s projects.
ip:framework.rag-wiki-substrate-ruledev:project.search-conversationsradar:concept.retrievalradar:concept.agent-search
queries asked of Scott's wikis
  • agent retrieval SQL tools versus ranked web search
  • RAG structured filtering aggregation source records
  • corpus completeness provenance retrieval blind spots
  • ClickHouse analytics agent knowledge systems
  • agent tool query budgets congestion pricing

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 569h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-17 23:21 (minted)⭐ origin echo-reconstructedScry documents one live table, reddit.posts, as a near-census of Reddit submissions from 2005-06-23 to 2026-09-17, with schema, capture limi
Scry on blog (echo) · attributed from hn.story.49748041 · published time unknown
—
09-17 23:15first on hacker news · published · lag ?Show HN: 500 TB internet index in ClickHouse, with congestion pricing
Xyra
—
09-17 23:15amplified on hacker news 👑hn.story.49748041
Xyra
peak 61 · 26 comments · 100% of case engagement
09-17 23:20our radar first saw it · lag ?discovery anchor: hn.story.49748041—
pace: p66 vs 1032 stories at the 336h mark (now 569h old) — ahead of compute-cheap-h100-h200-pricing (1.0x), behind openai-german-wiki-incident (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: 500 TB internet index in ClickHouse, with congestion pricing
Retrieved article excerpt

Open article · Retrieved 2026-09-17T23:21:20.594323+00:00

### One table live

`reddit.posts` is a near-census of Reddit submissions from 2005-06-23 to 2026-09-17, one row per observation of a post (the bulk month and the live tail each land their own row, so count posts with `uniqExact(id)`); posts removed before capture are absent. Completeness is measured — post ids are one global counter, so held-versus-allocated is exact per month.

|  |  |  |
| --- | --- | --- |
| `id` | String | one global base36 counter |
| `subreddit` | String |  |
| `author` | String |  |
| `created_utc` | DateTime |  |
| `title` | String |  |
| `selftext` | String | body of a text post; empty for link posts |
| `score` | Int32 | net upvotes |
| `num_comments` | Int32 |  |
| `upvote_ratio` | Float32 |  |
| `domain` | String | where a link post points |
| `url` | String |  |
| `search_text_lc` | String | lower(title + selftext), token-indexed |

The comment tree is `reddit.comments`, joined on `link_id = concat('t3_', id)`. Every relation is documented like this — columns, indexes, extent, known holes — at [`GET /v1/scry/schema`](https://scry.io/docs/schema-and-provenance). [Source catalog →](https://scry.io/sources)
Xyra6126
🟧 echo.blog ⭐Scry documents one live table, reddit.posts, as a near-census of Reddit submissions from 2005-06-23 to 2026-09-17, with schema, capture limiScry——

Interpretation history

Decision trace