2026-10-11 16:38 UTC

Harvey presents post-trained RLM agents for end-to-end M&A diligence, potentially extending professional agents from isolated legal tasks to an integrated diligence workflow.

state: seedheat: lowuncertainty: highconvergesscott: mediumresearch-agents enterprise-agents long-horizon-orchestrationHarvey

What is this?

Harvey published “Post-Training RLM Agents for End-to-End M&A Diligence,” dated September 8, 2026, with authors including Niko Grupen, Julio Pereyra, and Gabe Pereyra. In the supplied X snippet, Harvey says it partnered with Baseten to post-train recursive language model (RLM) agents and claims that model-harness co-optimization meaningfully improves performance in long-horizon environments. The claim is progress toward end-to-end M&A diligence, not demonstrated completion: the supplied article excerpt contains no experimental results, performance figures, or training details. A separate Harvey snippet describes extending Legal Agent Bench to M&A diligence to evaluate end-to-end legal work and support open-model training and agent research.

Why it matters to Scott

Harvey’s claimed model–harness co-optimization for M&A diligence converges with Scott’s Model-Plus-Harness Benchmark Unit position, opening a concrete publishing opportunity around a professional-agent builder treating capability as a joint systems property rather than weights alone. The supplied radar hits do not track this same development, but absent results or architectural details, this is convergence in stated approach—not validation of Scott’s long-running architecture or a demonstrated challenge to his preference for bounded AI judgments over end-to-end autonomy.
ip:concept.model-plus-harness-benchmark-unitradar:concept.recursive-agentsradar:concept.long-running-orchestrationradar:concept.agent-evaluationradar:concept.enterprise-agents
queries asked of Scott's wikis
  • model-harness co-optimization and agent post-training
  • recursive language models and long-horizon agent orchestration
  • agent evaluation in realistic end-to-end task environments
  • professional agents moving from isolated tasks to integrated workflows
  • document-intensive research agents and multi-document diligence

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 771h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-09 13:27 (minted)⭐ origin echo-reconstructedThe linked first-party post is titled “Post-Training RLM Agents for End-to-End M&A Diligence”; the supplied observation includes no results
Harvey on blog (echo) · attributed from hn.story.49625790 · published time unknown
—
09-09 12:56first on hacker news · published · lag ?Post-Training RLM Agents for End-to-End M&A Diligence
yarapavan
—
09-09 12:56amplified on hacker news 👑hn.story.49625790
yarapavan
peak 4 · 0 comments · 98% of case engagement
09-09 13:21our radar first saw it · lag ?discovery anchor: hn.story.49625790—
pace: p11 vs 519 stories at the 720h mark (now 771h old) — behind addom-local-coding-harness (0.5x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnPost-Training RLM Agents for End-to-End M&A Diligenceyarapavan40
🟧 echo.blog ⭐The linked first-party post is titled “Post-Training RLM Agents for End-to-End M&A Diligence”; the supplied observation includes no results Harvey——

Interpretation history

Decision trace