2026-10-11 16:38 UTC

Opper AI's Jevman Pac-Man benchmark becomes a recurring reference for comparing latency-sensitive decision models (Jev, KEV, Clef, GPT-6 Luna, Laya) in real-time control tasks.

state: corroboratedheat: mediumuncertainty: lowconvergesscott: highdecision-models agent-evaluation jev-benchmark real-time-control pac-man-benchmarkOpper AIfelix089

What is this?

Opper AI released 'Jevman', an open-source benchmark where six decision-model APIs (jev-1.13, KEV, Clef, Clef-Flash, GPT-6 Luna Decisions, Laya) play Pac-Man in real time against scripted ghosts with a 2-second per-decision deadline. The benchmark is hosted at github.com/opper-ai/jevman-benchmark and was announced via Show HN. Opper serves these models under a unified 'System One' API; the models span proprietary (OpenAI's GPT-6 Luna), cloud-hosted open-weight (Cloudflare's Clef/Clef-Flash on Workers AI), and community fine-tunes (KEV, a Qwen3.5-4B tune by Jared Palmer; Laya from ConvAI Innovations). Third-party coverage (Arize, explainx.ai, AIMultiple) treats it as a recurring reference for comparing latency-sensitive decision models.

Why it matters to Scott

An independent third-party benchmark (Jevman) arrives at Scott's exact evaluation methodology: trace-backed, latency-constrained comparison of decision-model APIs in real-time control. The benchmark includes jev-1.13 (TypeSafe Jev, Scott's ~0.3s 'System One' decisions API) alongside proprietary, cloud-hosted open-weight, and community fine-tune models β€” directly testing the model-barbell / cheap-front-door / task-aware routing architecture Scott builds and the sovereignty posture he argues for. This is a dated-receipts moment: independent evaluation validates his evaluation-driven development principle and fast-slow split architecture for real-time AI.
dev:project.jevdev:technology.typesafe-jevdev:concept.trace-backed-agent-comparisonip:concept.evaluation-driven-developmentip:framework.fast-slow-splitip:concept.real-time-ai-systemsip:concept.latency-accuracy-asymmetryip:framework.the-lane-doctrinedev:concept.cheap-model-front-doordev:concept.task-aware-model-routingdev:concept.cost-tiered-llm-routingdev:concept.plan-forecast-model-routingip:framework.sovereign-software-assuranceip:concept.vendor-lock-inip:concept.model-perishabilityip:framework.decision-authority-infrastructureip:concept.decision-attestation-packagedev:concept.deterministic-agent-control-planedev:concept.deterministic-first-ai-councildev:technology.litellmdev:technology.openrouterradar:agent-memory-leaderboard-validationradar:agent-review-studio-local-evaluationradar:agentgauntlet-failure-benchmarkradar:agentabstain-benchmark-validityradar:ai-benchmark-saturation-distortionradar:1password-scam-agent-benchmarkradar:aa-agentperf-local-benchmarkradar:500-dollar-9b-rl-catalog-reviewradar:afm3-prompt-conditioned-pruningradar:accelerated-understanding-neural-operator
queries asked of Scott's wikis
  • decision-model evaluation benchmarks
  • system-one fast-inference architecture
  • open-weight model sovereignty local inference
  • agent evaluation frameworks real-time control
  • latency-constrained model routing
  • model API unification patterns

Measured heat

now 0 pts/hpeak 5 pts/hcomments 0/hpeers p16momentum: steady2 platformsage 99h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

10-07 13:00⭐ origin echo-reconstructedSix decision models played 100 games each against classic Pac-Man ghosts in real-time with 2-second decision deadlines; benchmark is open so
Opper AI on blog (echo) Β· attributed from hn.story.50007993
β€”
10-08 16:34first on hacker news Β· published Β· +27.6hShow HN: Jevman – AI decision models play Pac-Man
felix089
β€”
10-08 16:34amplified on hacker news πŸ‘‘hn.story.50007993
felix089
peak 65 Β· 16 comments Β· 100% of case engagement
10-08 17:35our radar first saw it Β· +28.6hdiscovery anchor: hn.story.50007993β€”
pace: p66 vs 1247 stories at the 96h mark (now 99h old) β€” ahead of compute-cheap-h100-h200-pricing (1.0x), behind royalcities-controllable-audio-synths (1.0x)

Evidence (2) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Jevman – AI decision models play Pac-Man
Retrieved article excerpt

Open article Β· Retrieved 2026-10-08T23:06:53.347561+00:00

# jevman: AI decision models play Pac-Man

Six of them played 100 games each against the arcade ghosts. See how they rank, watch their games or play against them.

jev 1.13**0**
High score**0**
Ghosts**Classic**

β–ΆPlay vs AI

β–²
β—€
β–Ό
β–Ά

Activity

Watch

## Leaderboard

100 games per model against the classic ghosts.

=1 means tied: those scores are within each other's margin of error.

## Benchmark your own model

jevman is open source, and any model behind an HTTP endpoint can play it: a hosted model, a fine-tune or one running on your laptop.

1. 01

   ### Put your model behind an endpoint

   At every junction the game sends the maze as JSON and asks for a direction, with odds if your model has them. The repo has a 34-line example endpoint to start from.
2. 02

   ### Run the benchmark

   One command plays the games under the leaderboard's rules: real time, 2 seconds per answer, the classic ghosts, three lives or 5 minutes.
3. 03

   ### Send a pull request

   CI replays every game you submit and checks its score. Your model then joins the leaderboard as self-reported.

[opper-ai/jevman-benchmarkAGPL-3.0](https://github.com/opper-ai/jevman-benchmark/blob/main/CONTRIBUTING.md#benchmark-your-own-model)
Is your model on Opper or another public API? [Open an issue](https://github.com/opper-ai/jevman-benchmark/issues/new) and we'll run it ourselves.

## Method

Games
:   Each model played 100 games against the classic scripted ghosts, each until it lost all three lives. Games are capped at 5 minutes so every run ends, but none came close: the longest lasted 2 minutes 24 seconds.

Decisions
:   At every junction the game asks the model one question, with the maze, pellets and ghosts as state. The model returns a probability per direction, and Pac-Man takes its pick.

Deadline
:   An answer that takes longer than 2 seconds is replaced by a simple backup rule, counted under backup moves.

Scores
:   The ranking uses the mean score with a 95% margin of error (Β±2 standard errors), and models within each other's margin are tied. The high score is a model's best single game.

Checks
:   Every game is recorded and replays exactly, so any result can be checked.
felix0896516
🟧 echo.blog ⭐Six decision models played 100 games each against classic Pac-Man ghosts in real-time with 2-second decision deadlines; benchmark is open soOpper AIβ€”β€”

Interpretation history

Decision trace