2026-10-11 17:10 UTC

research-agents

band: hotmomentum: stable score: 1.0
temperature history

Episodes (35)

Expert review will determine whether Tencent’s Hyra agent and Hy3 model materially enabled a valid proof settling the optimal exponent relating sumsets and difference sets.
seedconvergesscott: medium
Independent evaluation will determine whether ProofCouncil's LLM-agent workflow can make reliable, meaningful progress on open mathematical problems beyond standard benchmark exercises.
expirednovelscott: low
Expert review will determine whether Claude materially contributed to a valid improvement from 41.6% to 67.2% in the proved lower bound on the proportion of Riemann zeta zeros satisfying the Riemann hypothesis.
expiredknownscott: medium
Independent technical evidence will determine whether Discovered Materials’ AI-agent workflow can identify experimentally credible materials for semiconductor thermal-management applications.
expiredknownscott: low
Independent use will determine whether Mole reliably enforces research spending limits, links claims to verified source quotations, and preserves a meaningful privacy boundary for local data.
expiredknownscott: medium
Independent evaluations will determine whether Prime Intellect's autonomous-research measurement framework produces reproducible, decision-useful comparisons of research agents.
expiredknownscott: low
Independent use will determine whether Huawei Noah’s released ScienceFlow agent can reliably execute practical long-horizon machine-learning research workflows.
expiredconvergesscott: medium
Expert verification will determine whether Levent Alpöge and Ava Howell, with material assistance from Claude, validly discovered an elliptic curve of rank 30.
corroboratedconvergesscott: medium
Expert review will determine whether GPT Sol materially assisted a valid resolution of the reported matrix-theory conjecture.
resolvedknownscott: low
Independent use will determine whether The Gauntlet’s specialist skills, executable verification, and automatic safeguards provide a practical structured harness for research and engineering agents.
expiredknownscott: low
Iluvatar Labs claims its AI research agent synthesized technical literature and produced a working open-source Codex Micro alternative costing about $40, demonstrating that research agents can complete nontrivial hardware-engineering projects.
expiredconvergesscott: medium
BixBench3’s authors claim frontier AI agents can reproduce roughly 48% of real computational-biology research workflows, providing a realistic measure of scientific-agent capability beyond synthetic tasks.
expiredconvergesscott: medium
Flare’s authors claim their LLM-based theorem-proving system can verify mixed-integer-program reformulations, potentially making equivalence checking more automated and rigorous than informal mathematical review.
expiredconvergesscott: medium
Sapient claims Praxist provides a usable autonomous R&D system that coordinates parallel research agents into substantive, reviewable research outputs.
expiredknownscott: low
OpenAI claims Rosalind Workbench lets individual scientists coordinate AI-assisted investigation, analysis, and research outputs through a reusable workflow rather than assembling a bespoke research-agent stack.
expiredconvergesscott: high
OpenAI claims its internal agents have reached “automated research intern” capability, with 3.1 agent-workdays of effort per human workday, and are progressing toward an automated AI researcher by March 2028, potentially shifting frontier-model R&D toward agent-executed research.
corroboratedconvergesscott: high
Terence Tao says the methods discussed in Buckmaster and Alpöge’s reported AI-assisted mathematics program could extend to Navier–Stokes despite enormous technical difficulties, potentially opening a route toward resolving the regularity problem.
corroboratedconvergesscott: high
Applied Compute presents Ari as its in-house AI research agent, potentially providing a concrete example of agent-executed research within a model-training and serving company.
expiredconvergesscott: low
Harvey presents post-trained RLM agents for end-to-end M&A diligence, potentially extending professional agents from isolated legal tasks to an integrated diligence workflow.
seedconvergesscott: medium
OpenAI reportedly claims a mathematical breakthrough on a Millennium Prize problem, potentially establishing substantive progress on a major open problem through AI-assisted research without yet establishing a complete solution.
acceleratingconvergesscott: high
Aniket Wathore claims the released Ramanujan workbench combines parallel multi-model research agents with isolated worktrees, literature provenance, and deterministic SymPy/Z3 checks, enabling inspectable computational-mathematics workflows rather than unchecked model consensus.
seedknownscott: low
Shengtao Guo, Ethan X. Fang, and Junwei Lu claim Odin discovered a proof giving a dimension-independent bound for the Komlós signing problem and square-root Beck–Fiala discrepancy, potentially establishing a major mathematical discovery by an AI research agent.
seednovelscott: low
InternLM claims its released Intern-S2-397B combines vision-language pretraining with multitask and long-horizon agent reinforcement learning to improve scientific reasoning and sustained agent work, potentially expanding open-model options for research workflows.
watchingnovelscott: low
Deep Dog 2's creator claims its released supervisor–subagent research package ranked fifth overall and first among open-source agents on DeepResearch Bench, with a separately estimated $0.25–$0.60-per-task configuration that could lower the cost of cited research reports.
seedknownscott: low
Martin Bertran Lopez and Aaron Roth claim successful ML research-agent strategies retain performance when compressed to as few as 16 tokens while overfit gains disappear, making compression a practical diagnostic for benchmark generalization.
seedconvergesscott: medium
Panel creator greentfrapp claims the released research workspace lets Claude Code create custom viewers alongside shared files, PDFs, and executable notebooks, reducing interface switching and bespoke visualization work during agent-assisted research.
corroboratedconvergesscott: medium
HN user kbr- claims their published AI-assisted, formalized proof solves a 12-year-old mathematics problem, potentially establishing a new machine-checkable research result if the formal statement and proof substantiate the claimed solution.
seedknownscott: low
John Sous and coauthors claim expert repairs and regrading reveal near-saturation of retained physics benchmark questions by frontier models, undermining low leaderboard scores as evidence of weak closed-form physics capability.
watchingconvergesscott: medium
Tong Zheng and coauthors claim Dream-RSI uses historical discovery trees to cheaply refine exploration policies around an unchanged coding agent, reducing discovery costs while maintaining or improving results in algorithm, mathematical-optimization, and GPU-kernel tasks.
watchingconvergesscott: high
James Zou and coauthors claim Paper2Agent converts papers, code, and data into MCP-backed agents that apply published methods to fresh datasets, potentially making research reproduction and reuse accessible through conversational tools.
resolvedconvergesscott: medium
AGORA's maintainer claims its released alpha lets independently operated agents publish, challenge, and reproduce research through an attributable append-only ledger while keeping credentials and private memory local, enabling auditable cross-owner research without treating consensus as scientific validity.
seedknownscott: low
James Zou and Harrison Zhang claim their 37,000-agent virtual biotech analyzed roughly 50,000 clinical trials in under a week and identified drug-success signals and a retrospectively concordant cancer-treatment strategy, potentially making large-scale agent orchestration useful for drug-research prioritization.
watchingconvergesscott: medium
Reuters reports that Anthropic is establishing a biology lab where Claude would direct laboratory robots with limited human intervention for preclinical drug discovery, extending its research agents into physical experimentation.
corroboratedconvergesscott: medium
Dan Abramov claims his published AI-assisted Lean proof resolves Conway's refinement conjecture for omnific integers, potentially establishing a new mathematical result through a nonexpert-led agent workflow despite outstanding expert verification.
seedconvergesscott: medium
Historian Benjamin Breen claims Opus 5.5, driven through his embedding-search-plus-agent workflow over the GLOBALISE VOC archive, surfaced a previously-unnoticed 1615 Dutch eyewitness record of dodo hunting and corrected a mistranslated 1638 red-rail reference — expert validation of the finds and replication of the workflow would establish frontier agents as producers of novel historical knowledge beyond math and code, while debunking or prior art would mark another inflated capability claim.
watchingconvergesscott: high

Trajectory notes