2026-10-11 18:03 UTC

Independent evaluations will determine whether increasing an LLM research agent’s number of web searches improves answer quality more reliably than switching search providers.

state: expiredheat: lowuncertainty: highconvergesscott: highweb-search agent-evals retrievalOpenRouter

What is this?

OpenRouter published benchmark results from live production runs across BrowseComp, DeepSearchQA, WideSearch, and HLE, arguing that giving an LLM research agent more searches improves answer quality more than changing search engines. A separate LiveNewsBench ablation found consistent gains across evaluated LLMs when the search budget increased from one to seven, supporting the value of additional search actions. However, the supplied independent snippets do not directly compare increased search budgets with switching providers, so OpenRouter’s stronger relative claim remains unverified here.

Why it matters to Scott

OpenRouter’s production results converge with Scott’s inference-time-scaling and model-plus-harness positions: search budget and orchestration may matter more than provider choice. The unresolved provider-versus-budget comparison directly fits his active provider-side execution benchmark and trace-backed comparison method, creating a concrete independent-validation and publishing opportunity rather than merely another retrieval example.
ip:concept.inference-time-scalingip:concept.model-plus-harness-benchmark-unitdev:project.remote-execdev:concept.trace-backed-agent-comparisonradar:concept.openrouterradar:concept.agent-evaluationradar:concept.agent-benchmarksradar:concept.inference-economics
queries asked of Scott's wikis
  • agent search-budget scaling
  • retrieval breadth versus provider quality
  • web research agent evaluation harnesses
  • adaptive stopping for iterative search
  • search cost quality latency tradeoffs
  • multi-hop retrieval orchestration

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnLLM Web Search Benchmarks: More Searches Beat a Better Search EngineDGAP10
🟧 echo.blog ⭐The OpenRouter blog presents its own benchmark results from live production runs across BrowseComp, DeepSearchQA, WideSearch, and HLE. Its cAyush Patel——

Interpretation history

Decision trace