2026-10-11 16:38 UTC

Redditor GuiltyBookkeeper4849 claims Artificium's publicly inspectable autonomous-agent run is approaching a solution to C(25,15,5) using fewer than the reported best 42 groups, potentially demonstrating a verifiable mathematical improvement from sustained agent search.

state: seedheat: lowuncertainty: mediumknownscott: lowlong-horizon-agents agent-research verifiable-mathArtificiumGuiltyBookkeeper4849

What is this?

The case describes Artificium as an open-source autonomous-agent harness whose creator reportedly exposes a live experiment observatory. Redditor GuiltyBookkeeper4849 claims a run exceeding 50 hours and 100 million tokens is getting closer to improving C(25,15,5), targeting fewer than a reported best of 42 groups—not that an improvement has been achieved. None of the supplied web snippets directly covers Artificium or this experiment, so they do not establish its creator’s identity, the reported benchmark, the run’s progress, or a verifiable mathematical result.

Why it matters to Scott

The evaluative position is already held in Scott’s The Mature Token Law and Agent Observability: inspect the run and audit conversion of compute into validated outcomes, rather than treating 50 hours and 100 million tokens as achievement. No supplied radar hit tracks this specific Artificium development, but the uncorroborated progress claim establishes neither a mathematical improvement nor a harness mechanism that would extend Scott’s architecture or change what he builds; it remains an example to watch, not a demonstrated convergence or challenge.
ip:framework.the-mature-token-lawip:concept.agent-observabilityradar:proofcouncil-llm-agent-open-mathradar:concept.long-horizon-agentsradar:concept.agent-observability
queries asked of Scott's wikis
  • long-horizon agent harnesses persistence stopping criteria
  • agent research independent verification exact certificates
  • agent memory failed approaches process lessons
  • autonomous search token budgets cost versus progress
  • public agent observability reproducible execution traces

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 554h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-18 14:00⭐ origin echo-reconstructedThe primary artifact is the creator’s live experiment observatory, not a repost. It describes Artificium as “my open-source general agent ha
GR (gr0010 / gr.bio; Reddit handle u/GuiltyBookkeeper4849) on other (echo) · attributed from reddit.post.1wmqkx4
—
09-21 21:49first on r/LocalLLaMA · published · +79.8h50+ Hours and 100M+ Tokens Later, Open Source Autonomous Agent is GETTING CLOSER at Solving an Open Math problem
GuiltyBookkeeper4849
—
09-21 21:49amplified on r/LocalLLaMA 👑reddit.post.1wmqkx4
GuiltyBookkeeper4849
peak 53 · 14 comments · 100% of case engagement
09-21 22:20our radar first saw it · +80.3hdiscovery anchor: reddit.post.1wmqkx4—
pace: p63 vs 1032 stories at the 336h mark (now 554h old) — ahead of cache-tax-idle-session-warming (1.0x), behind hunterbench-live-pentesting-benchmark (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit50+ Hours and 100M+ Tokens Later, Open Source Autonomous Agent is GETTING CLOSER at Solving an Open Math problem
LocalLLaMA
GuiltyBookkeeper48495314
🟧 echo.other ⭐The primary artifact is the creator’s live experiment observatory, not a repost. It describes Artificium as “my open-source general agent haGR (gr0010 / gr.bio; Reddit handle u/GuiltyBookkeeper4849)——

Interpretation history

Decision trace