2026-10-11 17:14 UTC

Kairo maintainer peter941221 claims its released research workbench measures 1.30–2.62× CUDA Graph throughput gains on specified RTX 5090 NVFP4 workloads and selects only exact measured serving profiles, enabling workload-specific optimization without assuming universal speedups.

state: seedheat: mediumuncertainty: mediumconvergesscott: lowlocal-inference inference-economics cuda-graphs inference-routingpeter941221

What is this?

The supplied case describes Kairo as an LLM inference research workbench maintained by peter941221, announced on Show HN with pinned RTX 5090 measurements and correctness gates. The maintainer reportedly claims 1.30–2.62× CUDA Graph throughput gains over eager execution for specified NVFP4 workloads, with routing restricted to exact measured serving profiles rather than assuming general speedups. None of the supplied web results mentions Kairo or its maintainer, so they do not independently establish the release, measurements, or routing behavior; they provide only adjacent consumer-GPU optimization material.

Why it matters to Scott

Kairo’s claimed restriction to exact measured serving profiles converges with Scott’s Hardware-aware local inference principle of making precision and compilation explicit runtime policy. However, the supplied material neither independently verifies the gains nor establishes applicability to his gamepc serving stack, so this remains an implementation example rather than a demonstrated reason to change what he builds; no supplied radar page already tracks Kairo.
dev:concept.hardware-aware-local-inferencedev:project.gamepcradar:concept.inference-optimizationradar:concept.inference-benchmarkingradar:concept.nvfp4
queries asked of Scott's wikis
  • local inference economics workload-specific benchmarks
  • measurement-driven inference routing exact profile matching
  • fail-closed runtime policies correctness gates
  • CUDA Graphs eager execution serving optimization
  • consumer GPU low-precision inference NVFP4

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 662h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-14 03:26 (minted)⭐ origin echo-reconstructedKairo publishes pinned RTX 5090 Graph-versus-eager measurements, correctness gates, and a runtime policy that selects exact measured workloa
peter941221 on github (echo) · attributed from hn.story.49691112 · published time unknown
—
09-14 02:21first on hacker news · published · lag ?Show HN: Kairo – Fail-closed LLM inference routing from RTX 5090 measurements
peter941221
—
09-14 02:21amplified on hacker news 👑hn.story.49691112
peter941221
peak 5 · 0 comments · 100% of case engagement
09-14 03:20our radar first saw it · lag ?discovery anchor: hn.story.49691112—
pace: p23 vs 1032 stories at the 336h mark (now 662h old) — ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Kairo – Fail-closed LLM inference routing from RTX 5090 measurements
Retrieved article excerpt

Open article · Retrieved 2026-09-14T03:21:29.453098+00:00

# Kairo

> Evidence-driven Blackwell inference research: measure a workload, prove a
> result, and route only what was measured.

[Phase 1 write-up](https://peter941221.github.io/Kairo/001-when-cuda-graphs-actually-help.html) ·
[Phase 1 results](https://github.com/peter941221/Kairo/blob/main/docs/PHASE_1_RESULTS.md) ·
[Experiment protocol](https://github.com/peter941221/Kairo/blob/main/docs/EXPERIMENT_PROTOCOL.md) ·
[Reproduce Phase 1](https://github.com/peter941221/Kairo/blob/main/docs/REPRODUCE_PHASE1.md) ·
[Blog](https://peter941221.github.io/Kairo/) ·
[Contributing](https://github.com/peter941221/Kairo/blob/main/CONTRIBUTING.md) ·
[Apache-2.0](https://github.com/peter941221/Kairo/blob/main/LICENSE)

## Why Kairo

On a single RTX 5090, CUDA Graph replay can improve NVFP4 serving throughput by
more than 2x for some workloads—and by far less for another valid workload.
The useful conclusion is not “always enable Graphs.” Kairo captures the
conditions behind a result, passes a correctness gate, and promotes only exact
measured workload buckets into a fail-closed runtime policy.

Kairo is a research workbench and the beginning of a path toward native
Blackwell inference components. It is not yet a general-purpose inference
engine or a replacement for vLLM, CUTLASS, or cuBLAS.

## Phase 1 at a glance

Local measurements on one RTX 5090 32GB, using pinned model and software
revisions. Read the linked protocols before comparing values.

| Measured workload | Graph | Eager | Result |
| --- | --- | --- | --- |
| Qwen3.8-27B-NVFP4, c32, prompt 512, context 1K | 645.08 tok/s | 308.51 tok/s | 2.09x |
| Qwen3.8-27B-NVFP4, c8, prompt 2048, context 4K | 322.78 tok/s | 123.22 tok/s | 2.62x |
| Qwen3.8-27B-NVFP4, c16, prompt 2048, context 4K | 260.07 tok/s | 199.52 tok/s | 1.30x |
| Qwen3-8B-NVFP4, c16, prompt 512, context 1K | 2177.49 tok/s | 1073.63 tok/s | 2.03x |

The 1.30x result is intentional context: CUDA Graphs are not promoted as a
global default. [Phase 1 results](https://github.com/peter941221/Kairo/blob/main/docs/PHASE_1_RESULTS.md) includes the scope,
evidence links, and claims Kairo does not make.

## What Kairo does

- Defines versioned workload and experiment protocols.
- Verifies declarative kernel blueprints before a build or launch.
- Records cache, correctness, artifact, and launch timing separately.
- Selects serving profiles only for exact measured workload buckets.
- Fails closed to `manual` when no measured route covers a request.
- Probes native SM120 data movement and matrix-compute paths, including
  correctness-aware negative results.

## Quick start

Clone the repository:

```
git clone https://github.com/peter941221/Kairo.git
cd Kairo
```

The portable contracts do not require a GPU:

```
$env:PYTHONPATH = "src"
python -m unittest discover -s tests -q
```

Validate the sample kernel blueprint:

```
wsl.exe -d Ubuntu-24.04 -- bash -lc 'cd /mnt/c/Projects/Kairo && bash scripts/wsl/run_lab.sh validate-blueprint --file experiments/blueprints/fp16-tma-wmma.yaml'
```

Ask the runtime policy about a measured route:

```
wsl.exe -d Ubuntu-24.04 -- bash -lc 'cd /mnt/c/Projects/Kairo && bash scripts/wsl/run_lab.sh recommend-runtime --model qwen38 --concurrency 32 --prompt-tokens 512 --context-tokens 1024 --generation-tokens 128'
```

GPU-serving paths require a compatible RTX 5090 environment, local model
weights, and the pinned runtime described in [the environment contract](https://github.com/peter941221/Kairo/blob/main/docs/ENVIRONMENT.md).
Set model and environment paths through the documented `KAIRO_*` variables;
do not commit local paths or credentials.

For the full evidence-to-serving-route procedure, including the expected route
and the exact c32/1K reproduction command, see [Reproduce Phase 1](https://github.com/peter941221/Kairo/blob/main/docs/REPRODUCE_PHASE1.md).

## Evidence and reproducibility

Every public result should answer: under which model, quantization, workload,
hardware, software revision, and correctness tolerance did it win?

- [Phase 1 results](https://github.com/peter941221/Kairo/blob/main/docs/PHASE_1_RESULTS.md) is the concise public result index.
- [Experiment protocol](https://github.com/peter941221/Kairo/blob/main/docs/EXPERIMENT_PROTOCOL.md) defines controls and the
  promotion gate.
- [Baseline status](https://github.com/peter941221/Kairo/blob/main/docs/BASELINE_STATUS.md) is the detailed chronological lab
  record and includes both promoted and rejected experiments.
- [`experiments/protocols/`](https://github.com/peter941221/Kairo/blob/main/experiments/protocols) contains machine-readable
  result and method records.
- [`evidence/`](https://github.com/peter941221/Kairo/blob/main/evidence) contains reviewed, redacted per-repeat public records;
  `python scripts/render_phase1_results.py` regenerates the published table and
  SVG chart in [`docs/generated/`](https://github.com/peter941221/Kairo/blob/main/docs/generated).
- [Public release checklist](https://github.com/peter941221/Kairo/blob/main/docs/PUBLIC_RELEASE.md) lists the artifact,
  review, and repository gates before public visibility changes.

## Repository map

```
blog/                 publication drafts and article sources
docs/                 result index, protocol, environment, and release guides
experiments/          versioned workloads, protocols, and kernel blueprints
scripts/wsl/          local RTX 5090 probes, runners, and analyzers
src/kairo_lab/        dependency-light validation, cache, dispatch, and policy
tests/                portable contract tests
.kairo-local/         ignored local credentials, logs, and machine facts
```

## Roadmap

**Phase 1 — complete research baseline.** Controlled protocols, correctness
gates, measured Graph/eager routes, route selection, and native capability
probes.

**Phase 2 — native components.** Use Phase 1 evidence to choose narrowly
scoped data movement, layout, subgraph, or execution-path work. A candidate is
not promoted without correctness, boundary coverage, repeatability, and a fair
end-to-end comparison.

## Contributing

See [CONTRIBUTING.md](https://github.com/peter941221/Kairo/blob/main/CONTRIBUTING.md). In brief: do not extrapolate measured
wins, do not commit credentials or private artifacts, and retain negative
results that establish a meaningful boundary.

## License

Copyright 2026 Kairo contributors. Licensed under the
[Apache License 2.0](https://github.com/peter941221/Kairo/blob/main/LICENSE). The license grants broad reuse rights, including
an express patent grant for contributed work; it does not grant trademark rights.
peter94122150
🟧 echo.github ⭐Kairo publishes pinned RTX 5090 Graph-versus-eager measurements, correctness gates, and a runtime policy that selects exact measured workloapeter941221——

Interpretation history

Decision trace