research-agents
band: hotmomentum: stable
score: 1.0
Episodes (35)
Trajectory notes
- 2026-09-25T12:29:53Z: paper2agent-executable-paper-agents closed (absorbed) — Zou's group has independently arrived at Scott's own doctrine — bound an artifact (paper + codebase) and compile it into agent-callable MCP tools gated by generated tests — making a Nature-published dated receipt for his c
- 2026-09-10T18:01:19Z: applied-compute-ari-research-agent closed (faded) — Applied Compute’s claimed failure-analysis-to-next-experiment workflow converges with Scott’s Compounding Test and agentic diagnostic loop: findings become input to subsequent work rather than just reports. For now this is ano
- 2026-09-04T17:29:49Z: gauntlet-verification-agent-harness closed (faded) — The Gauntlet combines positions Scott already holds in Skills and Workflows, Evaluation-Driven Development, Agent Receipts, and Architecture, Not Vibes: specialist modules surrounded by executable gates, audit traces, and str
- 2026-08-30T22:30:05Z: openai-rosalind-workbench closed (faded) — OpenAI is a consequential party productising Scott’s existing position that durable agent capability comes from a reusable harness combining specialised tool orchestration, governed workflows, human review, and provenance-bearing outpu
- 2026-08-30T18:32:14Z: praxist-parallel-research-agents closed (faded) — Scott already holds and implements this pattern in Discovery Accelerator and Synthetic Futures Factory, while the radar already tracks research agents and multi-agent orchestration. Praxist currently adds only an unvalidated pro
- 2026-08-29T19:38:01Z: flare-milp-reformulation-verification closed (faded) — Flare applies Scott’s verification-loop and deterministic-verifier position to a consequential new domain: an LLM proposes mathematical reasoning while a symbolic proof system supplies the machine-checkable gate for MILP eq
- 2026-08-29T15:32:32Z: bixbench3-biology-agent-workflows closed (faded) — BixBench3 independently operationalizes Scott’s position that agents should be judged through realistic, long-horizon work and observable artifact-level criteria rather than synthetic tasks or subjective completion claims. Its
- 2026-08-28T19:35:35Z: iluvatar-agent-built-codex-micro closed (faded) — The claimed research-to-working-artifact loop converges with Scott’s Cognition Supply Chain and agentic-coding position that useful autonomy comes from structured exploration, iteration and verification rather than one-shot gene
- 2026-08-24T19:55:02Z: claude-riemann-zeta-bound closed (faded) — The radar already tracks the same expert-review-dependent AI-mathematics pattern in ProofCouncil, Claude Fable, and other claimed proofs; Scott’s Challenger, Never Arbiter and Provenance-Coupled Work frameworks already require that the
- 2026-08-21T15:34:50Z: prime-intellect-autonomous-research-evals closed (faded) — Scott’s Evaluation-Driven Development and trace-backed agent-comparison pages already require repeatable, decision-useful evaluation rather than accepting benchmark claims at face value, while the radar’s agent-evaluati