2026-10-11 16:37 UTC

Antfly CTO AJ Roetker claims its v0.2 Zig rewrite integrates search and inference with holistic resource management and simulation testing, potentially making the database more portable and reliable across embedded and distributed deployments.

state: watchingheat: lowuncertainty: highnovelscott: lowknowledge-systems retrieval inference-databases systems-testingAntflyAJ Roetker

What is this?

Antfly is a retrieval database for AI agents, built by co-founders James McDermott and AJ Roetker, who previously worked on data systems at Lytics. Its website describes keyword, vector, and graph retrieval in one query plan, with chunking, embedding, and reranking inference inside the engine; its GitHub snippet describes a zero-dependency Zig implementation. Older Go descriptions alongside the Zig repository are consistent with a rewrite, but the supplied snippets do not establish the v0.2 release details, holistic resource management, Raft trace validation, deterministic simulation testing, or demonstrated portability and reliability improvements.

Why it matters to Scott

Scott’s search and dev-wiki projects make retrieval infrastructure adjacent to his work, but the supplied hits establish neither Antfly adoption nor a position on integrated search/inference engines that this rewrite challenges or independently validates. The claimed resource-management and simulation-testing benefits remain unestablished in the grounding, and database simulation alone does not converge with his historical design-replay frameworks; no supplied radar page tracks this Antfly development.
queries asked of Scott's wikis
  • RAG stack consolidation search inference unified engine
  • agent memory retrieval graphs debugging
  • local inference private data embedded deployment
  • retrieval inference shared resource management
  • distributed systems deterministic simulation Raft validation

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 722h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-11 14:00⭐ origin echo-reconstructedAntfly announces its v0.2 engine rewrite from Go to Zig and describes portability, resource control, Raft trace validation, and deterministi
AJ Roetker on blog (echo) · attributed from hn.story.49714157
—
09-15 15:35first on hacker news · published · +97.6hA search-and-inference database from scratch in pure Zig
kingcauchy
—
09-15 15:35amplified on hacker news 👑hn.story.49714157
kingcauchy
peak 66 · 19 comments · 100% of case engagement
09-18 20:20our radar first saw it · +174.3hdiscovery anchor: hn.story.49714157—
pace: p66 vs 519 stories at the 720h mark (now 722h old) — ahead of anthropic-claude-sandbox-breakouts (1.0x), behind llama-cpp-rdna4-flash-attention (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnA search-and-inference database from scratch in pure Zig
Retrieved article excerpt

Open article · Retrieved 2026-09-18T20:22:45.502356+00:00

[Back to Research](https://antfly.io/research)September 12, 2026

# A Search-and-Inference Database from Scratch in Pure Zig

Why we rewrote Antfly's Go engine in Zig while the startup was still early: first principles, caring about the model instead of the embeddings, TigerBeetle-style simulation testing, and what the rewrite made possible.

AJ Roetker

[AJ Roetker](https://antfly.io/about/aj-roetker)

CTO

[Zig](https://antfly.io/research/tags/Zig)[distributed-systems](https://antfly.io/research/tags/distributed-systems)[search](https://antfly.io/research/tags/search)[raft](https://antfly.io/research/tags/raft)[vector-db](https://antfly.io/research/tags/vector-db)[ML](https://antfly.io/research/tags/ML)[performance](https://antfly.io/research/tags/performance)

My colleague [Rowan](https://antfly.io/about/rowan-copley) summarized our ambitious goal for Antfly over a year ago: perfect search! This goal is silly, ambitious, unobtainable, and perfect for us. Hearing grandiose technologists talk about perfect search still strikes me as something out of an episode of *Silicon Valley*, but I love an impossible goal, and my inner tech hipster appreciates the irony. Aim for perfect search and you miss, but you miss somewhere interesting. So I'll walk you through why we did the thing you are never supposed to do with a new startup: we rewrote the product. [Antfly v0.1](https://antfly.io/research/distributed-search-engine-go) launched a document storage, full-text, vector, and graph indexing engine in Go. Antfly v0.2 launches the same engine in Zig, with zero dependencies (well, none that get to run the show, but more on that below).

## First principles[#](https://antfly.io/research/antfly-zig#first-principles)

I don't hear people talking about first principles as often as I used to, but they're still important to me and to how I make decisions. When building the first version of Antfly, the idea was to fill a gap in the database market: a schema-ish, friendly query engine for indexing like Elasticsearch, closer in scale to Postgres than Iceberg, as easy to use as Mongo, and as easy to operate as Google's [Spanner](https://research.google/pubs/spanner-googles-globally-distributed-database-2/) or Bigtable. (And I thought perfect search was too lofty... HA!)

When building the first version I didn't have the same sort of software tools (cough cough Codex, Claude, Aider, Pi) available to assist in development, and the hard problems I wanted to tackle were the ones [CockroachDB had similarly chosen Go for](https://www.cockroachlabs.com/blog/why-go-was-the-right-choice-for-cockroachdb/): distributed systems and concurrency. Rust can guarantee memory safety through lifetimes, but memory safety wasn't the hard part, and Go is infinitely more readable to me than Rust ever was. Plus, Go had the most battle-tested [Raft](https://raft.github.io/raft.pdf) implementation out there ([etcd's](https://github.com/etcd-io/raft)), and I was only one person, working on weekends and on my Fridays, trying to build an Elasticsearch DSL on top of [Bleve](https://blevesearch.com/) as well.

## Caring about the model, not the embeddings[#](https://antfly.io/research/antfly-zig#caring-about-the-model-not-the-embeddings)

Fast forward a little bit and embeddings started to become a thing. [word2vec](https://arxiv.org/abs/1301.3781) had shown that a vector could actually carry meaning, and the first practical embedding models meant an average dev could build Google-lite semantic search for their app. At work, our scale of vector storage was so small that a top-k could be an exhaustive search, but I read about all the fun algorithms behind Pinecone, Vertex AI Vector Search, Elastic, Mongo, CockroachDB, and pgvector after an old coworker rolled his eyes at the ridiculous Pinecone seed round. Pinecone might have been overvalued, but I think one of its most interesting ideas was overlooked: users could care about the model instead of having to care about the embeddings. It made me think about Postgres, and how for the most part a user can avoid knowing about B-trees and other indexing algorithms, or how in Bleve and Lucene you can avoid knowing about S2 indexes, finite state transducers, and so on. Vector databases and indexes, on the other hand, required you to know and care about [HNSW](https://arxiv.org/abs/1603.09320), [SPFresh](https://arxiv.org/abs/2410.14452), [RaBitQ](https://arxiv.org/abs/2405.12497), and the rest.

I decided to try my hand at implementing these algorithms from their papers and blog posts in the initial version of Antfly, and it was wildly successful. I was able to make a semantically searchable Wikipedia using my laptop, Antfly, and [Ollama](https://ollama.com/)! Embedding generation was so slow that the database being a little bit slower was not a big concern for the initial implementation.

## Enter Zig[#](https://antfly.io/research/antfly-zig#enter-zig)

Concurrently with all this, I had tried my hand at implementing [VSR](https://dspace.mit.edu/handle/1721.1/71763), Protobuf, and an LSM in [Zig](https://ziglang.org/) a few years prior, after stumbling upon [TigerBeetle](https://tigerbeetle.com/). I really liked the readability of Zig, the concurrency primitives, the people implementing the language, and the interoperability with C. But Zig was too green for my weekend database project and lacked a lot of the heavy lifting: Raft, full-text indexing, a portable, battle-tested LSM, yada yada yada. I took a lot of the spirit of the TigerBeetle folks with me, though, and put myself to work incorporating [VOPR](https://docs.tigerbeetle.com/concepts/safety/) testing with [TLA+](https://lamport.azurewebsites.net/tla/tla.html) trace validation (Rowan's [post on formal verification with coding agents](https://antfly.io/research/agent-formal-verification) covers that), built on the new [Go mock time](https://go.dev/blog/synctest) and on prior art from [etcd's Raft TLA+ spec and trace validation](https://github.com/etcd-io/raft/tree/main/tla).

Then [`std.Io`](https://kristoff.it/blog/zig-new-async-io/) started making a big splash across the technoverse when Zig decided to make some serious overhauls to the language. I had been reading the GitHub design threads on the subject for a while and thought it was all pretty cool from an engineering perspective. Antfly had just released v0.1 and I had a breath of air to start thinking about what came next. So I started to play around with Zig and the [new `std.Io` work](https://andrewkelley.me/post/zig-new-async-io-text-version.html), trying to implement our LSM using the async I/O and `std.Io.Evented` machinery Zig had started to expose. At the same time, I wanted to see how far I could take the software factorization of our code, and it felt like Raft, TLA+ specs, clear traces, and language-agnostic tests were the perfect hill to climb. So I went to work directing traffic and building Raft, full-text indexing, an LSM, and... HTTP/2 (and our Raft transport, which was over [HTTP/3](https://www.rfc-editor.org/rfc/rfc9114) and [QUIC](https://www.rfc-editor.org/rfc/rfc9000)). I had also rewritten most of our Go-based end-to-end tests in Python, both to make sure the coding agents couldn't "cheat" by reaching into the Go code and to make sure our Python SDK was solid. This turned out to be the perfect language-agnostic framework for transpiling to Zig too.

## Why rewrite[#](https://antfly.io/research/antfly-zig#why-rewrite)

A conversation with [James](https://antfly.io/about/james-mcdermott) and another with [Drew](https://antfly.io/about/drew-lanenga) really solidified the idea for me when we talked about what we were optimizing for: asymmetric outcomes and reliability. Most startups land on their face, so the expensive bet, writing every high-performance dependency ourselves, is the one worth making. A user MUST know and trust that their database works, and a user wants that database to fly, not sprint (it needs to be astonishingly fast). People were already building [all sorts of interesting projects](https://antfly.io/showcase) on top of Antfly. To run it from any other programming language or in the browser, though, we'd need something like Rust or Zig to give the code a C-compatible, WASM- and WebGPU-compatible interface. I wanted Antfly to be the grand unified theory of databases: a machete for old-school use cases and traditional apps, and the perfect Swiss army knife for the AI and semantic use cases nobody has seen yet (perfect search didn't seem lofty enough anymore). Something a developer could embed in a sandboxed environment (laptop, unit tests, Lambda), run at average application workloads (PostgreSQL, Cockroach, Mongo), or run at analytic scale (data warehouse, serverless). And I wanted Antfly to be even more reliable in all of those environments while maintaining or exceeding the performance goals we had set for ourselves. A rewrite of everything from the ground up gives you a chance to build all your dependencies in a purpose-aware shape, baking in resource management, priority scheduling, and testing ideology from day one.

So why Zig and not Rust? Four reasons, all back to first principles: portability, C interop, speed, and testability.

Portability and C interop: Zig is a C compiler with a libc for every target. Cross-compiling Antfly with CUDA, ONNX, and Wasmtime linked in is one flag, and calling them is `@cImport`. Rust cross-compiles pure Rust fine and breaks on the first C dependency, which is why `cargo-zigbuild` is Zig.

Speed: same LLVM, so codegen is a wash. The difference is the fast version is the default. Every allocation takes an allocator, nothing allocates or branches behind your back, and SIMD and comptime specialization are built in. In Rust the hot paths end up in `unsafe`. Next to Go it's not close: no GC, no scheduler, no CGO.

Testability: `std.Io` makes the world a parameter. Hand a package a simulated `Io` and it gets VOPR: disk, network, and clock faults, all of it. In Rust you rip out tokio for madsim or turmoil and hope the dependency tree cooperates.

People usually think formal verification and memory safety protect you from a whole class of bugs, and they do, but much like the validity of your lifetime in Rust, it depends on the context... Rust has a borrow checker and Zig doesn't, but the bugs that bug me aren't use-after-frees. They're a replica that fell behind, a message that arrived twice, an fsync that lied. The borrow checker doesn't cover that, and tools like [Kani](https://github.com/model-checking/kani) and [Verus](https://github.com/verus-lang/verus) prove things about Rust functions, not about whether your Raft is linearizable. Rust or Zig, you still need something like [Antithesis](https://antithesis.com/) throwing faults at the whole system. In either language you want every bounds and overflow check on while you test, because one missed check is someone's data. Rust turns overflow checks off in release like everyone else; Zig makes safety one build mode you can flip. So we verify the protocol, TLA+ specs, trace validation against the running Raft, and VOPR on every package, none of which cares what language you wrote it in. Memory safety we buy in testing: every suite runs in Debug and ReleaseSafe with the testing allocator and the same simulator, and we ship ReleaseFast.

There was one more big reason to rewrite that I had wanted from the Go version anyway. We had already started to own forks of all our major dependencies. If we owned every dependency outright, we could control memory, CPU, and GPU resources holistically across the whole process, the same argument TigerBeetle makes in [Tiger Style](https://github.com/tigerbeetle/tigerbeetle/blob/main/docs/TIGER_STYLE.md) with its zero-dependencies policy. With Zig, zero dependencies is something you can actually keep, not a purity goalpost that moves every time you need something to go faster. Linking a C library is a non-event, so C
kingcauchy6619
🟧 echo.blog ⭐Antfly announces its v0.2 engine rewrite from Go to Zig and describes portability, resource control, Raft trace validation, and deterministiAJ Roetker——

Interpretation history

Decision trace