Redditor GuiltyBookkeeper4849 claims Artificium's publicly inspectable autonomous-agent run is approaching a solution to C(25,15,5) using fewer than the reported best 42 groups, potentially demonstrating a verifiable mathematical improvement from sustained agent search.
state: seedheat: lowuncertainty: mediumknownscott: lowlong-horizon-agents agent-research verifiable-mathArtificiumGuiltyBookkeeper4849
What is this?
The case describes Artificium as an open-source autonomous-agent harness whose creator reportedly exposes a live experiment observatory. Redditor GuiltyBookkeeper4849 claims a run exceeding 50 hours and 100 million tokens is getting closer to improving C(25,15,5), targeting fewer than a reported best of 42 groups—not that an improvement has been achieved. None of the supplied web snippets directly covers Artificium or this experiment, so they do not establish its creator’s identity, the reported benchmark, the run’s progress, or a verifiable mathematical result.
Why it matters to Scott
The evaluative position is already held in Scott’s The Mature Token Law and Agent Observability: inspect the run and audit conversion of compute into validated outcomes, rather than treating 50 hours and 100 million tokens as achievement. No supplied radar hit tracks this specific Artificium development, but the uncorroborated progress claim establishes neither a mathematical improvement nor a harness mechanism that would extend Scott’s architecture or change what he builds; it remains an example to watch, not a demonstrated convergence or challenge.
ip:framework.the-mature-token-lawip:concept.agent-observabilityradar:proofcouncil-llm-agent-open-mathradar:concept.long-horizon-agentsradar:concept.agent-observability
queries asked of Scott's wikis
- long-horizon agent harnesses persistence stopping criteria
- agent research independent verification exact certificates
- agent memory failed approaches process lessons
- autonomous search token budgets cost versus progress
- public agent observability reproducible execution traces
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 554h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p63 vs 1032 stories at the 336h mark (now 554h old) — ahead of cache-tax-idle-session-warming (1.0x), behind hunterbench-live-pentesting-benchmark (1.0x)
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-09-21T22:30:24Z
grounded: known/low — The evaluative position is already held in Scott’s The Mature Token Law and Agent Observability: inspect the run and audit conversion of compute into validated
2026-09-21T22:25:16Z
origin walked (codex/luna, conf 0.95): anchor reddit.post.1wmqkx4 -> echo.other.aca27d2d0d by GR (gr0010 / gr.bio; Reddit handle u/GuiltyBookkeeper4849)
2026-09-21T22:23:46Z
case created — The linked live experiment supplies a bounded, checkable research target, but the excerpt establishes neither a better construction nor measurable progress toward one.
Decision trace
- 09-27 12:47review_screenjev screen: no material development (noul=0.17)
- 09-22 08:30groundThe evaluative position is already held in Scott’s The Mature Token Law and Agent Observability: inspect the run and audit conversion of compute into validated outcomes, rather than treating 50 hours
- 09-22 08:25promote_anchororigin walk conf 0.95
- 09-22 08:23createThe linked live experiment supplies a bounded, checkable research target, but the excerpt establishes neither a better construction nor measurable progress toward one.