ProofCouncil is an LLM-based agent system for solving open mathematical problems, introduced in a recent arXiv paper. It uses an author-critic workflow and achieved a partial solution to Erdős Problem 539, an open problem in number theory, verified by human experts and partially formalized in Lean. The system also outperformed proprietary models on the FirstProof benchmark, but the provided snippets do not name the authors or describe independent evaluation beyond the paper's own claims.
2026-08-13T21:33:14Z
The system-specific validation watch has gone dormant: repeated checks produced no independent expert evaluation, formal verification, or reproducible replication. Archive the episode unless substantive mathematical validation later revives it.
2026-08-11T20:43:07Z
The staleness check found no new evaluation, replication, or formal verification, so the system-specific validation gap remains unchanged. Keep the case on a long cadence and revisit only on substantive mathematical evidence.
2026-08-09T20:27:52Z
The staleness check and six additional HN points add no system-specific evidence; ProofCouncil remains an unverified workflow claim rather than corroborated open-math progress. Move to a longer validation cadence pending independent expert review, formal verification, or reproducible replication.
2026-08-07T19:37:43Z
The refreshed discussion is broader speculation about AI and open mathematics, not independent evaluation of ProofCouncil. Its system-specific claims remain unverified, so engagement-only updates should no longer trigger frequent review.
2026-08-05T16:35:14Z
No new system-specific validation is identifiable; broader evidence that AI can advance open mathematics still does not verify ProofCouncil’s workflow or claimed results. Keep this as a cold validation watch and ignore further engagement-only triggers until independent expert review, formal verification, or reproducible replication appears.
2026-08-05T15:26:04Z
The latest movement is only negligible engagement on already-priced evidence; broader AI open-math progress still does not validate ProofCouncil’s specific workflow or claimed results. Keep this as a cold system-specific validation watch pending independent expert review, formal verification, or reproducible replication.
2026-08-05T14:30:49Z
Independent coverage of AI progress on difficult open mathematics makes ProofCouncil’s broader capability claim more plausible, but it does not independently evaluate this workflow or verify its reported results. The case remains a system-specific validation watch rather than corroborated evidence of reliable open-problem progress.
2026-08-05T14:22:01Z
evidence attached: hn.story.49181519 — Independent coverage of AI making progress on historically difficult open mathematics materially bears on the open-math capability hypothesis.
2026-08-02T12:21:53Z
The latest trigger adds no substantive evidence and leaves the core validation gap unchanged. Keep ProofCouncil as a cold watch, but reconsider only if independent expert review, formal verification, or reproducible replication appears.
2026-08-02T03:21:05Z
The nominal attachment again adds no substantive evidence, leaving the case dependent on first-party claims and one anecdotal run. Keep it as a cold validation watch and defer reconsideration until independent mathematical review, formal verification, or reproducible replication appears.
2026-08-01T19:22:42Z
The purported new attachment adds no identifiable evidence, so the case still hinges on first-party claims and one anecdotal run without independent mathematical validation. Keep it as a cold validation watch and suppress further engagement-only repricing.
2026-08-01T16:22:38Z
The trigger contains no substantive new evidence, so ProofCouncil remains dependent on first-party claims and one anecdotal run rather than independent mathematical validation. Revisit only when expert review, formal verification, or reproducible replication appears.
2026-08-01T14:27:48Z
The nominal attachment adds no substantive evidence, leaving the claimed open-problem progress without independent expert validation or reproducible replication. This remains a cold validation watch; further engagement-only triggers should not prompt repricing.
2026-08-01T13:22:24Z
The nominal attachment contains no identifiable new evidence and does not alter the validation gap: the case still rests on the authors’ claims and one anecdotal user report. Ignore further engagement-only triggers until independent mathematical review, formal verification, or reproducible replication appears.
2026-08-01T08:22:32Z
The nominal attachment provides no substantive new evidence; ProofCouncil still rests on its authors’ claims and one unverified user report. Keep it as a cold validation watch and revisit only if independent expert review, formal verification, or reproducible replication emerges.
2026-08-01T05:21:11Z
The nominal attachment adds no identifiable independent review, replication, or formal verification, so the case remains an unresolved validation watch. Repetitive engagement-only triggers add no meaning and should not prompt frequent review.
2026-08-01T02:21:47Z
The purported attachment contains no identifiable new evidence, leaving the case dependent on the authors’ claims and one unverified user report. Defer further review until independent mathematical validation, formal verification, or reproducible replication appears.
2026-08-01T00:23:08Z
No genuinely new evidence is identifiable; the case still depends on the authors’ claims and one anecdotal, unverified user run. Reprice only when independent mathematical review, formal verification, or reproducible replication appears.
2026-07-31T23:22:51Z
No substantive new evidence has appeared; the only measurable change is negligible engagement, while independent expert review, replication, and formal verification remain absent. Keep this as a validation watch and stop repricing engagement-only updates.
2026-07-31T21:24:59Z
The nominal new attachment provides no identifiable independent review, replication, or formal verification beyond the paper and already-priced anecdotal run. The case remains a validation watch, and repeated engagement-only triggers should be ignored until substantive mathematical evidence appears.
2026-07-31T20:25:04Z
The nominal attachment adds no identifiable independent review, replication, or formal verification beyond the paper and already-priced anecdotal run. The case remains a low-temperature validation watch; engagement-only triggers no longer merit hourly reconsideration.
2026-07-31T19:25:39Z
No substantive new evidence is identifiable: the case still rests on the authors’ paper and one unverified user report. Further engagement or duplicate attachments should not change its meaning without independent expert review, reproducible results, or formal verification.
2026-07-31T18:25:45Z
The trigger adds no substantive evidence beyond the paper and already-priced anecdotal run; independent expert review, reproducible results, and formal verification remain absent. Repeated engagement updates should not advance the case without validation.
2026-07-31T17:27:57Z
The attachment adds no identifiable independent review, replication, or formal verification beyond the paper and anecdotal user run already priced in. This remains a low-temperature validation watch, with further engagement alone unable to advance it.
2026-07-31T16:28:25Z
The latest trigger adds no identifiable independent review, replication, or formal verification; it is repetitive amplification of the same anecdotal run. ProofCouncil remains a low-temperature validation watch rather than evidence of reliable progress on open mathematics.
2026-07-31T15:27:35Z
The newly attached material still supplies no independent mathematical review, reproducible result, or formal verification beyond the authors’ claims and one user’s anecdotal report. The case remains a validation watch rather than evidence that ProofCouncil reliably advances open mathematics.
2026-07-31T14:26:24Z
No new independent mathematical review or reproducible validation has appeared; the attached user report remains anecdotal amplification already reflected in the case. Keep watching for expert verification, formalized results, or independent replications rather than engagement.
2026-07-31T13:22:10Z
An external user report moves ProofCouncil beyond a paper-only claim by showing the workflow being run on open problems, but its alleged discoveries remain anecdotal and expert validation is still absent. The case now merits monitoring for independent mathematical review rather than promotion on engagement or self-reported results.
2026-07-31T13:21:37Z
evidence attached: reddit.post.1vbq62x — This is a concrete user-run report of an LLM agent searching literature and producing claimed mathematical discoveries, though expert validation remains absent.
2026-07-29T10:23:34Z
grounded: novel/low — No intersection found. The wiki_hits and radar_hits are both empty — Scott's own wikis contain no pages on ProofCouncil, mathematical reasoning agents, or the E
2026-07-29T10:22:19Z
origin walked (codex/luna, conf 0.98): anchor hn.story.49095351 -> echo.paper.d2efa02954 by Johannes Schmitt, Tim Gehrunger, Jasper Dekoninck, Gergely Bérczi, Uri Kreitner, Liam Price, and David Holmes
2026-07-29T10:21:20Z
case created — Single arXiv paper with low engagement; plausible research-agent episode but needs more evidence to confirm it's developing.