Opper AI's Jevman Pac-Man benchmark becomes a recurring reference for comparing latency-sensitive decision models (Jev, KEV, Clef, GPT-6 Luna, Laya) in real-time control tasks.
state: corroboratedheat: mediumuncertainty: lowconvergesscott: highdecision-models agent-evaluation jev-benchmark real-time-control pac-man-benchmarkOpper AIfelix089
What is this?
Opper AI released 'Jevman', an open-source benchmark where six decision-model APIs (jev-1.13, KEV, Clef, Clef-Flash, GPT-6 Luna Decisions, Laya) play Pac-Man in real time against scripted ghosts with a 2-second per-decision deadline. The benchmark is hosted at github.com/opper-ai/jevman-benchmark and was announced via Show HN. Opper serves these models under a unified 'System One' API; the models span proprietary (OpenAI's GPT-6 Luna), cloud-hosted open-weight (Cloudflare's Clef/Clef-Flash on Workers AI), and community fine-tunes (KEV, a Qwen3.5-4B tune by Jared Palmer; Laya from ConvAI Innovations). Third-party coverage (Arize, explainx.ai, AIMultiple) treats it as a recurring reference for comparing latency-sensitive decision models.
Why it matters to Scott
An independent third-party benchmark (Jevman) arrives at Scott's exact evaluation methodology: trace-backed, latency-constrained comparison of decision-model APIs in real-time control. The benchmark includes jev-1.13 (TypeSafe Jev, Scott's ~0.3s 'System One' decisions API) alongside proprietary, cloud-hosted open-weight, and community fine-tune models β directly testing the model-barbell / cheap-front-door / task-aware routing architecture Scott builds and the sovereignty posture he argues for. This is a dated-receipts moment: independent evaluation validates his evaluation-driven development principle and fast-slow split architecture for real-time AI.
dev:project.jevdev:technology.typesafe-jevdev:concept.trace-backed-agent-comparisonip:concept.evaluation-driven-developmentip:framework.fast-slow-splitip:concept.real-time-ai-systemsip:concept.latency-accuracy-asymmetryip:framework.the-lane-doctrinedev:concept.cheap-model-front-doordev:concept.task-aware-model-routingdev:concept.cost-tiered-llm-routingdev:concept.plan-forecast-model-routingip:framework.sovereign-software-assuranceip:concept.vendor-lock-inip:concept.model-perishabilityip:framework.decision-authority-infrastructureip:concept.decision-attestation-packagedev:concept.deterministic-agent-control-planedev:concept.deterministic-first-ai-councildev:technology.litellmdev:technology.openrouterradar:agent-memory-leaderboard-validationradar:agent-review-studio-local-evaluationradar:agentgauntlet-failure-benchmarkradar:agentabstain-benchmark-validityradar:ai-benchmark-saturation-distortionradar:1password-scam-agent-benchmarkradar:aa-agentperf-local-benchmarkradar:500-dollar-9b-rl-catalog-reviewradar:afm3-prompt-conditioned-pruningradar:accelerated-understanding-neural-operator
queries asked of Scott's wikis
- decision-model evaluation benchmarks
- system-one fast-inference architecture
- open-weight model sovereignty local inference
- agent evaluation frameworks real-time control
- latency-constrained model routing
- model API unification patterns
Measured heat
now 0 pts/hpeak 5 pts/hcomments 0/hpeers p16momentum: steady2 platformsage 99h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
How the heat travelled
pace: p66 vs 1247 stories at the 96h mark (now 99h old) β ahead of compute-cheap-h100-h200-pricing (1.0x), behind royalcities-controllable-audio-synths (1.0x)
Evidence (2) β β canonical anchor
| source | object | author | score | comments |
| π§ hn | Show HN: Jevman β AI decision models play Pac-ManRetrieved article excerptOpen article Β· Retrieved 2026-10-08T23:06:53.347561+00:00 # jevman: AI decision models play Pac-Man
Six of them played 100 games each against the arcade ghosts. See how they rank, watch their games or play against them.
jev 1.13**0**
High score**0**
Ghosts**Classic**
βΆPlay vs AI
β²
β
βΌ
βΆ
Activity
Watch
## Leaderboard
100 games per model against the classic ghosts.
=1 means tied: those scores are within each other's margin of error.
## Benchmark your own model
jevman is open source, and any model behind an HTTP endpoint can play it: a hosted model, a fine-tune or one running on your laptop.
1. 01
### Put your model behind an endpoint
At every junction the game sends the maze as JSON and asks for a direction, with odds if your model has them. The repo has a 34-line example endpoint to start from.
2. 02
### Run the benchmark
One command plays the games under the leaderboard's rules: real time, 2 seconds per answer, the classic ghosts, three lives or 5 minutes.
3. 03
### Send a pull request
CI replays every game you submit and checks its score. Your model then joins the leaderboard as self-reported.
[opper-ai/jevman-benchmarkAGPL-3.0](https://github.com/opper-ai/jevman-benchmark/blob/main/CONTRIBUTING.md#benchmark-your-own-model)
Is your model on Opper or another public API? [Open an issue](https://github.com/opper-ai/jevman-benchmark/issues/new) and we'll run it ourselves.
## Method
Games
: Each model played 100 games against the classic scripted ghosts, each until it lost all three lives. Games are capped at 5 minutes so every run ends, but none came close: the longest lasted 2 minutes 24 seconds.
Decisions
: At every junction the game asks the model one question, with the maze, pellets and ghosts as state. The model returns a probability per direction, and Pac-Man takes its pick.
Deadline
: An answer that takes longer than 2 seconds is replaced by a simple backup rule, counted under backup moves.
Scores
: The ranking uses the mean score with a 95% margin of error (Β±2 standard errors), and models within each other's margin are tied. The high score is a model's best single game.
Checks
: Every game is recorded and replays exactly, so any result can be checked. | felix089 | 65 | 16 |
| π§ echo.blog β | Six decision models played 100 games each against classic Pac-Man ghosts in real-time with 2-second decision deadlines; benchmark is open so | Opper AI | β | β |
Interpretation history
2026-10-09T12:08:00Z
Scott explicitly up-voted this case; independent Jevman benchmark validates his exact evaluation methodology (trace-backed, latency-constrained, real-time control) and fast-slow/model-barbell architecture. HN engagement steady at 82nd percentile for age. Two independent lines of evidence now: third-party benchmark + Scott's directed attention.
2026-10-08T23:52:48Z
grounded: converges/high β An independent third-party benchmark (Jevman) arrives at Scott's exact evaluation methodology: trace-backed, latency-constrained comparison of decision-model AP
2026-10-08T23:41:23Z
case created β Released benchmark pitting multiple decision-model APIs against each other in real-time Pac-Man; independent evaluation artifact in decision-model space.
Decision trace
- 10-10 10:28attention_communicatedOpper AI's Jevman benchmark has a material update (revision 2, cause: material_reprice) since the 18:08 briefing. The benchmark runs trace-backed, latency-constrained evaluation of six decision-m
- 10-10 10:28attention_routeScott flagged this case with feedback_interrupt during the 18:08 briefing, and a material reprice has landed since (last_material_at 23:08 Sydney, ~1.5 hours ago). The benchmark directly validates his
- 10-10 00:33attention_routeScott flagged this case with feedback_interrupt during the 18:08 briefing, and a material reprice has landed since (last_material_at 23:08 Sydney, ~1.5 hours ago). The benchmark directly validates his
- 10-09 23:11attention_routeScott flagged this case with feedback_interrupt during the 18:08 briefing, and a material reprice has landed since (last_material_at 23:08 Sydney, ~2 min ago). The benchmark directly validates his eva
- 10-09 23:08attention_candidatematerial_reprice
- 10-09 23:08repriceScott explicitly up-voted this case; independent Jevman benchmark validates his exact evaluation methodology (trace-backed, latency-constrained, real-time control) and fast-slow/model-barbell architec
- 10-09 18:42feedback_interruptScott vote via UI
- 10-09 18:08attention_communicatedOpper AI's Jevman runs trace-backed, latency-constrained evaluation of decision-model APIs in real-time Pac-Man (2s deadline). Includes jev-1.13 (TypeSafe Jev, Scott's ~0.3s System One decis
- 10-09 18:08attention_routeFurther reading for 6 PM briefing: independent benchmark validating Scott's decision-model evaluation methodology and fast-slow architecture. Results are live and reproducible. Briefing allows ti
- 10-09 15:39sensor_dirtycomment_update
- 10-09 13:18attention_routeIndependent benchmark validating Scott's decision-model evaluation methodology and fast-slow architecture. Results are live and reproducible. Briefing allows time to review the leaderboard and co
- 10-09 13:11attention_candidatecreate
- 10-09 10:52groundAn independent third-party benchmark (Jevman) arrives at Scott's exact evaluation methodology: trace-backed, latency-constrained comparison of decision-model APIs in real-time control. The benchm
- 10-09 10:41createReleased benchmark pitting multiple decision-model APIs against each other in real-time Pac-Man; independent evaluation artifact in decision-model space.