2026-10-11 16:37 UTC

Elastic claims its atune harness combines profiling, statistically gated microbenchmarks, real-workload validation, and human review to discover useful Elasticsearch optimizations, potentially reducing the engineering attention required to improve mature infrastructure.

state: seedheat: mediumuncertainty: mediumconvergesscott: mediumcoding-agents agent-harnesses software-optimizationElasticThomas VeaseyChris Hegarty

What is this?

The case describes atune as an Elastic AI-agent harness for optimizing Elasticsearch, combining profiling, statistically gated microbenchmarks, real-workload validation, and human review. The supplied web snippets establish Elastic's Elasticsearch platform and mention earlier performance measurement using the Rally macrobenchmarking framework, but none directly documents atune or the cited article. Consequently, the harness design, reported optimization discoveries, roles of Thomas Veasey and Chris Hegarty, and possible reduction in engineering attention remain unverified by these search results.

Why it matters to Scott

Elastic's claimed combination of statistical benchmark gates, real-workload validation and human review converges with Scott's Evaluation-Driven Development and Governance Barbell, offering a potential publishing comparison about verification-led agents optimizing mature infrastructure rather than merely generating code. The supplied snippets do not verify atune or reduced engineering attention, so this remains a claim-level opportunity; related radar optimization cases do not establish that this Elastic development is already tracked.
ip:concept.evaluation-driven-developmentip:concept.mechanically-different-verifiersip:framework.governance-barbellradar:codex-autoresearch-gpu-kernel-speedupradar:glm-infra-agent-serving-optimizationradar:graphsignal-agent-readable-gpu-profiler
queries asked of Scott's wikis
  • coding agent harnesses empirical feedback loops
  • statistical benchmark gates real workload validation
  • human review bottlenecks agent engineering attention
  • autonomous code optimization mature infrastructure
  • agent evaluation benchmark gaming correctness performance tradeoffs

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 746h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-10 14:00⭐ origin echo-reconstructedDescribes Elastic's atune optimization harness and reports discoveries involving ES|QL string conversion, NEON dot products, and a gzip-libr
Thomas Veasey and Chris Hegarty on blog (echo) · attributed from hn.story.49740769
—
09-17 13:52first on hacker news · published · +167.9hTrust, but benchmark: How we let an AI agent optimize Elasticsearch
eatonphil
—
09-17 13:52amplified on hacker news 👑hn.story.49740769
eatonphil
peak 1 · 0 comments · 106% of case engagement
09-17 14:20our radar first saw it · +168.3hdiscovery anchor: hn.story.49740769—
pace: p11 vs 519 stories at the 720h mark (now 746h old) — behind addom-local-coding-harness (0.5x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnTrust, but benchmark: How we let an AI agent optimize Elasticsearch
Retrieved article excerpt

Open article · Retrieved 2026-09-17T14:22:19.794246+00:00

[Blog](https://www.elastic.co/search-labs/blog)

# Trust, but benchmark: How we let an AI agent optimize Elasticsearch

We share how we built a harness that automatically identifies and implements optimizations in the Elasticsearch codebase.

September 11, 2026

[Thomas Veasey](https://www.elastic.co/search-labs/author/thomas-veasey)[Chris Hegarty](https://www.elastic.co/search-labs/author/chris-hegarty)

[Inside Elastic](https://www.elastic.co/search-labs/blog/category/inside-elastic)[Agentic AI](https://www.elastic.co/search-labs/blog/category/agentic-ai)

Jump to

Share

Test Elastic's leading-edge, out-of-the-box capabilities. Dive into our sample notebooks in the [Elasticsearch Labs repo](https://github.com/elastic/elasticsearch-labs?tab=readme-ov-file), start a [free cloud trial](https://cloud.elastic.co/registration), or try Elastic on your [local machine now](https://github.com/elastic/start-local).

[Elasticsearch](https://elastic.co/elasticsearch) executes a diverse set of workloads, including sustained heavy index building and real-time search and analytics. Delivering excellent performance across the board requires going broad in coverage while simultaneously diving deep enough into the codebase to understand optimization opportunities for each workload. Traditionally, human attention has been the bottleneck in this process; there simply aren't enough engineering hours to scrutinize every hot code path looking for inefficiency across a large and evolving surface area.

However, with the rapid progression of coding agents, performance optimization has become a task we can tackle semiautomatically. Unlike many software engineering challenges, optimizing code offers a cheap and objective verifier. If you ask an AI model to make code faster, there’s a hard number at the end telling you exactly what happened, backed by profiling tools that explain why. This makes performance a perfect candidate for automation, provided you can actually trust the numbers.

If you simply point a coding agent at a benchmark, you typically get low signal-to-noise: wins that fall inside the variance of the environment, or variations caused by thermal throttling rather than better code. To capture optimizations that actually benefit Elasticsearch users, we had to bridge the gap between "checkable in principle" and "checked in practice." We built a highly trustworthy measurement loop: a [harness](https://en.wikipedia.org/wiki/Agent_harness) that assumes the agent will be wrong a good fraction of the time but reliably catches and proves it when it’s right and then helps guide it where to look next.

Once the machinery is in place, the results speak for themselves. By letting this harness loose on the codebase, we've already begun uncovering meaningful wins across the stack. In part 2 of this post, we’ll dive into some examples it has found so far, including string conversion inefficiencies in Elasticsearch Query Language ([ES|QL](https://www.elastic.co/docs/explore-analyze/query-filter/languages/esql)), an improvement to our NEON vector dot product implementation, and an upgrade opportunity for the gzip library we were using. In this part, we’ll take a look at the design choices we made and how they relate to the broader topic of effective harness development.

## The AI code optimization pipeline architecture

The first step in any software engineering problem is to identify the correct high-level components. We made an architectural choice that turned out to be very helpful for this problem: separate understanding where opportunities exist from the loop making code changes. The agent starts with a real workload but only uses it to mine information about where to seek performance improvements. At this stage, it’s instructed to go broad and consider a range of performance-related signals. Once it has found and classified the hot spots, the agent reads the context of the code around them to understand the optimization opportunities. We use a separate task to condense the ranked list of hot spots into artifacts that a loop can iterate against in minutes: a microbenchmark that we prove exercises the hot path in its real operating regime. Finally, we use a proposer-verifier loop to actually make changes to the codebase to improve performance on the benchmark. This hands off to validation to assess the impact on real workloads at the end. Our CLI (`atune`) supplies the tools this process needs, and the rest is largely automated by a set of task-specific instructions.

For context, our high-level architecture looks like the following. Pink boxes are the humans, and teal boxes are the agent. There are three task types, one skeleton loop, and one referee.

AI code optimization pipeline: exploration to benchmark, human approval, exploitation, validation and PR review

*An exploration task profiles a real workload and produces ranked opportunities; a human promotes one into an exploitation task, which iterates against an approved microbenchmark and commits each accepted experiment; a validation run on the real workload guards the result before a human reviews and opens the PR. Where no benchmark covers the hot path, a benchmark task authors one and a human approves it into a registry. A performance atlas informs every task and accumulates what each one learns.*

## Why performance optimization suits autonomous agents

Three properties make a task ideally suited for autonomous work, and it's worth being explicit about them because they provide a checklist you can use to evaluate automation candidates. You want:

1. An objective verdict so that the agent can be held to something other than its own opinion.
2. A dense guiding signal so that it knows where to look next instead of guessing.
3. A bounded blast radius so that being wrong is affordable.

Performance gives you all three. Benchmarks provide the verdict and profilers provide the gradient, while a rejected patch costs you wall-clock time rather than correctness. The change gets reverted, and the reason gets recorded. Life goes on. Regarding the first two, we've come to think the gradient matters more than the verdict. ″This got faster″ is binary, whereas a profile hints at what to try next. An agent can generate its next hypothesis conditioned on a rich guiding signal, the richer the better, rather than grinding through a list.

## Signals are what the agent gets to see

A useful mental model is that the CLI is the agent's sensory apparatus. That changes how you design each command. Rather than exposing a capability, you design it to return a clear and concise answer to a question about the task at hand, and, when relevant, an explanation the model can reason over, instead of raw data it has to parse and interpret. These are the signals that the agent acts on, and we ended up with the following for our harness:

|  |  |
| --- | --- |
| **Signal** | **Question it answers** |
| Facet-decomposed macro profile | Where in the code does real workload time go, per query type? |
| Allocation and lock sampling, in the same capture | Is the cost cycles, garbage, or contention? |
| Cost-composition classification | Is this in scope compute, other product code, GC, JIT tax, or parked threads? |
| Input-shape instrumentation | What does the workload actually feed this code? |
| Statistical verdict | Did this change help, at this measured noise floor? |
| Allocation-rate comparison | Did the new code end up allocating more? |
| Interpreted disassembly | Why did that result happen? |
| End-to-end A/B guard, with differential profile attribution | Did anything appear to break, and was it us? |
| Environment check | Is this machine even fit to measure right now? |
| Upstream duplicate search | Has somebody already reported or fixed this? |

Four of these signals are worth dwelling on, because in each case the tool encodes a judgment that the agent would otherwise have had to keep making by hand.

Facet-decomposition is the clearest example. Blended CPU shares hide breadth: if you profile a mixed query workload, the grouping hash map insert and the percentiles sketch update can both show up as single-digit percentages of the total run time and look comparable. They aren't comparable at all, because the hash map insert is paid for by nearly every aggregation query, while the sketch update is only paid for when somebody asks for percentiles. So the profiler runs each named query facet as its own race, and every opportunity the agent records carries a breadth field (universal, broad, or narrow) and gets ranked by headroom × tractability × breadth. A universal 3% [beats](https://www.elastic.co/beats) a narrow 10%. Putting the ranking function in the tool prevents it from having to be rediscovered on every run.

The cost-composition classification works as a router. Rather than handing the agent a flat top-N list of frame names, it buckets every sampled stack into ″in scope compute,″ ″other product compute,″ ″GC,″ ″JIT and safepoint overhead,″ ″off CPU waiting,″ and ″parked threads.″ Each of those buckets implies a different kind of investigation. If GC is above about 15%, the real target is allocation rate, and the CPU top frames will actively mislead you, because they show where objects were collected rather than where they were created. If JIT and safepoint overhead is above about 25%, you're looking at a ceiling rather than an opportunity, since no in-scope code change will move it. If threads are parked and core utilization is low, this is a concurrency problem and CPU flame graphs are the wrong instrument entirely. We wrote that mapping into the playbook as a table, so the model (even a cheap one) reads a profile the way that an experienced engineer would, rather than reaching straight for the top frame. The classification is a prior for forming a hypothesis, though, not a substitute for evidence, so the agent still has to cite specific frames when it proposes an experiment.

Interpreted disassembly is a tool we hadn't originally provided, but it most definitely earns its place. It helps answer *why*, the question that unblocks the next hypothesis. Flame graphs tell you where the time goes; they rarely tell you why a change made things worse. So `atune asm` runs the benchmark briefly with the JIT told to print the assembly for one hot method, captures both sides of the working-tree diff, reduces each to the final C2 compilation, normalizes the addresses, and diffs them. The diff alone would likely still be 4,000 lines of aarch64, so on top of it sits an interpretation layer: a per-mnemonic delta, a net instruction count, the compilation tier that was actually captured, and a vectorization signal that counts vector register references on each side and raises a warning when they're eliminated, halved, or narrowed from [Advanced Vector Extensions](https://en.wikipedia.org/wiki/Advanced_Vector_Extensions) (AVX) to [Streaming SIMD Extensions](https://en.wikipedia.org/wiki/Streaming_SIMD_Extensions) (SSE) width. The playbook then maps mnemonic patterns to causes:

|  |  |
| --- | --- |
| **Pattern in the diff** | **Likely cause** |
| `b.eq``/``b.ne` up, `csel` down | New unpredictable branches |
| Clusters of `str``/``ldr` against the stack pointer | The compiler ran out of registers |
| NEON loads replaced by scalar compares | The vector path degraded |

In one experiment, the agent fused two [SIMD](https://en.wikipedia.org/wiki/Single_instruction,_multiple_data) mask extractions into one, and the benchmark regressed by 26%. The vectorization warning explained it in about 10 seconds. Without that tool, the agent has a dead end and no working model of the machine; with it, it has a corrected model and several new ideas.

The fourth signal is less a single tool than a habit; the instruments check themselves. Core utilization is derived two independent ways, from sample density and from process sampling, so the two can be compared. The classification is rejected if t
eatonphil10
🟧 echo.blog ⭐Describes Elastic's atune optimization harness and reports discoveries involving ES|QL string conversion, NEON dot products, and a gzip-librThomas Veasey and Chris Hegarty——

Interpretation history

Decision trace