2026-10-11 16:38 UTC

Deep Dog 2's creator claims its released supervisor–subagent research package ranked fifth overall and first among open-source agents on DeepResearch Bench, with a separately estimated $0.25–$0.60-per-task configuration that could lower the cost of cited research reports.

state: seedheat: lowuncertainty: highknownscott: lowresearch-agents agent-harnesses inference-economicsbeneadieDeep Dog 2

What is this?

The case describes Deep Dog 2, attributed to creator beneadie, as a released Python research-agent package using supervisor–subagent delegation and citation validation, with a claimed historical fifth-place overall and first-place open-source ranking on DeepResearch Bench. The supplied web snippets establish that DeepResearch Bench, by Mingxuan Du and colleagues, evaluates research agents on 100 expert-crafted tasks across 22 fields, assessing report quality and citation effectiveness and accuracy. None of the snippets mentions Deep Dog 2 or its creator, so they do not independently verify the release, architecture, ranking, or estimated $0.25–$0.60 per-task cost; the case also does not establish that the separately priced configuration achieved the claimed ranking.

Why it matters to Scott

The architectural position is already held in Scott’s Micro-Agents Architecture; Deep Dog 2’s reported delegation and citation validation currently add another example, not an established extension of that framework or his Proof-of-read citation gate. No supplied radar hit tracks this specific release, but the creator’s unverified ranking and separately estimated cost do not establish a quality–cost improvement that would change Scott’s builds or arguments; that would require a same-configuration, trace-backed comparison.
ip:framework.micro-agents-architecturedev:concept.proof-of-read-citation-gatedev:concept.trace-backed-agent-comparisonradar:concept.research-agentsradar:concept.inference-economicsradar:apodex-frontieragent-research-harnessradar:mole-budgeted-research-agent
queries asked of Scott's wikis
  • supervisor subagent delegation research harnesses
  • citation validation source provenance research reports
  • agent inference economics cost quality tradeoffs
  • research agent benchmarks reproducible evaluation
  • open-source research pipelines knowledge wiki ingestion

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 671h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-13 17:34 (minted)⭐ origin echo-reconstructedDeep Dog 2 ships as a Python research package with structured delegation and citation validation; its README reports a historical fifth-plac
beneadie on github (echo) · attributed from hn.story.49686186 · published time unknown
—
09-13 17:14first on hacker news · published · lag ?DeepDog 2: Outperforms Gemini and OpenAI deep research at a fraction of the cost
beneadie01
—
09-13 17:14amplified on hacker news 👑hn.story.49686186
beneadie01
peak 2 · 1 comments · 101% of case engagement
09-13 17:22our radar first saw it · lag ?discovery anchor: hn.story.49686186—
pace: p32 vs 1032 stories at the 336h mark (now 671h old) — ahead of addom-local-coding-harness (1.5x), behind agentsec-static-config-auditing (0.8x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnDeepDog 2: Outperforms Gemini and OpenAI deep research at a fraction of the cost
Retrieved article excerpt

Open article · Retrieved 2026-09-13T17:23:31.132494+00:00

# Deep Dog 2 — research engine

Turn a question into a cited Markdown report. A supervisor plans the research, delegates work to platform agents, reviews their findings, and writes the final report.

Deep Dog 2 installs directly from GitHub as a **Python package** in your existing project. Call `run_research()` with just a question to use the defaults, or pass a per-run configuration to choose models, search, agent types and research budgets. See the [quickstart](https://github.com/beneadie/deep_dog_2#quickstart) for pip and uv installation, or [clone the source](https://github.com/beneadie/deep_dog_2#work-from-a-source-checkout) to edit the engine and run the repository examples.

## Official benchmark results

At the time of publication, Deep Dog 2 ranked **5th overall** and **1st among open-source research agents** on the [DeepResearch Bench](https://huggingface.co/spaces/muset-ai/DeepResearch-Bench-Leaderboard). The published run used a relatively economical profile: a 15-minute research window, a 20-iteration supervisor cap, at most 3 Exa searches per sub-agent, DeepSeek V4 Pro as supervisor, and DeepSeek V4 Flash for sub-agents.

[DeepResearch Bench rankings showing Deep Dog 2 in fifth place](https://github.com/beneadie/deep_dog_2/blob/main/assets/deep-dog-2-ranking.png)

- **Role-specific models** — use different models for the supervisor, initial draft, and platform sub-agents.
- **Reflection and delegation** — the supervisor plans research, delegates distinct questions, evaluates findings, and decides whether more work is needed.
- **Platform specialization** — Web is enabled by default; Reddit, Substack, General, PubMed, Arxiv, and SEC agents can be enabled through configuration.
- **Evidence-first reports** — findings are collected into a source registry, then final inline citations and the `## Sources` section are validated before the report is returned.
- **Operational control** — time limits, iteration caps, search budgets, fallback chains, and output modes can be tuned for local experiments or more economical deployments.

## Features

- **Supervisor + sub-agent architecture** — configurable platform agents in [config.py](https://github.com/beneadie/deep_dog_2/blob/main/deep_research/config.py); `ResearchWeb` is enabled by default, with Reddit, Substack, General, PubMed, Arxiv and SEC agents available through configuration
- **Multiple search providers** — Tavily and/or Exa (`WEB_SEARCH_ENGINE` at `deep_research/config.py:658`)
- **Provider-agnostic models** — DeepSeek, MiMo, Meta Muse, Gemini, OpenAI, GLM, and OpenRouter-hosted models via a single `get_model()` factory (`deep_research/config.py:883`)
- **Model fallback chains** — per-role fallback lists (`SUBAGENT_MODEL_FALLBACK_CHAIN`, `SUPERVISOR_MODEL_FALLBACK_CHAIN`, `DRAFT_REPORT_MODEL_FALLBACK_CHAIN`)
- **Cited reports** — Markdown report, source metadata and optional research trace; applications can save the returned results
- **LangGraph execution** — recursion limit, timeouts, and observability logging

## How It Works

Deep Dog 2 retains the draft-first, iterative refinement idea from Deep Dog 1, while making reflection, delegation, and source handling explicit:

```
User question
     |
     v
clarify_with_user → write_research_brief → write_draft_report
                                               |
                                               v
                                    supervisor research loop
                              ┌───────────────┼────────────────┐
                              │               │                │
                           reflect        delegate       conclude research
                         (think_tool)   (parallel agents)
                                              |
                                              v
                               Web / Reddit / Substack / ...
                                              |
                                              v
                                     findings + sources
                                              |
                                              └── repeat until complete
                                                        |
                                                        v
                              final report → citation validation → output
```

1. **Scope the question.** The input is converted into a structured research brief.
2. **Create a scaffold.** An initial draft establishes a useful report structure before live research begins. It is not treated as evidence.
3. **Reflect and delegate.** The supervisor uses internal reflection to identify gaps and delegates focused, non-overlapping research tasks to platform agents.
4. **Research in parallel.** Agents search, read, save, and compress findings using the tools available for their platform. Their results are returned with source metadata and citations.
5. **Evaluate.** The supervisor reviews the findings and may request another round.
6. **Finalize the report.** The final writer combines the brief, draft, and findings. Citation checks validate the relationship between inline citations and the final sources list.

Supervisor reflection is used for planning and control; it is not copied into the final research report. Optional subtopic evaluation and parallel subtopic reports can run after the main report when enabled in `deep_research/config.py`.

The supported prompt family is `OPEN`. Older configurations using `LEGACY` must switch to `OPEN`.

## Official Benchmark Results

At the time of publication, Deep Dog 2 ranked **5th overall** and **1st among open-source research agents** on the [DeepResearch Bench](https://huggingface.co/spaces/muset-ai/DeepResearch-Bench-Leaderboard) benchmark. These results were achieved with a relatively economical configuration: a maximum of 3 Exa searches per sub-agent, 15 minutes of research time, 20 total iterations, and DeepSeek V4 Pro and DeepSeek V4 Flash as the supervisor and sub-agent base models.

| Metric | Score |
| --- | --- |
| Overall | **0.5432** |
| Comprehensiveness | 0.5468 |
| Insight | 0.5532 |
| Instruction Following | 0.5426 |
| Readability | 0.5105 |

These results are a historical reproducibility profile, not a promise about current defaults. The current package defaults to DeepSeek V4 Flash for all model roles and uses a different supervisor iteration default. Benchmark rankings and scores may change as the leaderboard changes.

### Estimated cost comparison

The following is a pricing-based estimate for one research task. It is not a controlled cost benchmark: the systems use different architectures, search providers and token budgets.

| System and configuration | Estimated cost per task | Basis |
| --- | --- | --- |
| **Deep Dog 2** — DeepSeek V4 Flash (off-peak) + Exa; 15-minute maximum, 20 supervisor iterations, at most 3 searches per sub-agent | **$0.25–$0.60** | Approximately **$0.20–$0.40** in DeepSeek inference and **$0.05–$0.20** in Exa usage for this configuration |
| **Gemini Deep Research** — Google’s published typical-task estimate | **$1–$3** | Google estimates about 80 search queries, 250k input tokens and 60k output tokens for a moderate task |

Deep Dog 2’s range is an estimate based on the [DeepSeek V4 pricing (off-peak)](https://api-docs.deepseek.com/quick_start/pricing/) and [Exa pricing](https://exa.ai/pricing), using off-peak DeepSeek rates. Actual cost varies with prompt length, model output, cache hits, the number of delegated agents and how quickly the supervisor concludes. Exa currently advertises **$20 in sign-up credits and $10 in credits each month**; [Tavily](https://docs.tavily.com/documentation/api-credits) provides **1,000 free credits per month**, equivalent to $8 at its $0.008 per-credit pay-as-you-go rate.

Costs are easy to change by changing the configuration: shorten the research window, lower the supervisor or sub-agent iteration limits, reduce searches per sub-agent, or use fewer agents. [Nemotron 3.5 Lightning](https://openrouter.ai/nvidia/nemotron-3.5-lightning) has also worked successfully as a sub-agent model in Deep Dog 2 and costs approximately half as much as DeepSeek V4 Flash in the tested setup. Google’s Gemini figures are its own published estimates; see the [Gemini Deep Research pricing section](https://ai.google.dev/gemini-api/docs/deep-research#estimated-costs) for the assumptions behind them.

For a detailed explanation of the reflection and delegation methods used, see the engineering article [Deep Dog 2: How Reflection and Structured Delegation Improve Supervisor–Subagent Research Systems](https://beneadie01.substack.com/p/deep-dog-2-how-reflection-delegation).

## Quickstart

Requires **Python 3.11+**, **Git**, a DeepSeek API key and an Exa API key. In your own project directory, use your existing Python environment or create one below.

Create and activate an environment if you do not already have one

Choose standard Python tooling:

```
# Standard Python tooling
python -m venv .venv
```

Or, using [uv](https://docs.astral.sh/uv/):

```
uv venv --python 3.11 --seed
```

`--seed` includes pip so either installer below works in this environment.

Activate the environment:

```
# macOS / Linux
source .venv/bin/activate
```

```
# Windows PowerShell
.venv\Scripts\Activate.ps1
```

**Install the package directly from GitHub.** Choose pip or uv:

```
# pip
python -m pip install "git+https://github.com/beneadie/deep_dog_2.git"
```

Or, with uv:

```
uv pip install "git+https://github.com/beneadie/deep_dog_2.git"
```

The installer fetches the repository, builds the package and installs its dependencies. You do not need a local source checkout or a PyPI release. This works because the repository includes Python packaging metadata in [pyproject.toml](https://github.com/beneadie/deep_dog_2/blob/main/pyproject.toml); a Git repository needs a Python package build configuration to support this kind of install. The installed distribution is named `deep-dog-2`; the Python import is `deep_research`.

**Add your keys** to a `.env` file in your project directory:

```
DEEPSEEK_API_KEY=your-deepseek-key
EXA_API_KEY=your-exa-key
```

Keep `.env` out of version control by adding it to your project's `.gitignore`. If those keys are already set as environment variables, no `.env` file is needed; existing environment variables take precedence.

**Create `simple_research.py` in your project** with the following code. The two settings near the top are optional examples: change the model or maximum research time, and leave the rest of the configuration at its defaults:

```
import asyncio
from dotenv import load_dotenv

load_dotenv()

from deep_research.integration import RunConfig, run_research

MODEL = "deepseek-v4-flash"
MAX_MINUTES = 10

config = RunConfig(
    supervisor_model_fallback_chain=[MODEL],
    subagent_model_fallback_chain=[MODEL],
    draft_report_model_fallback_chain=[MODEL],
    research_time_max_minutes=MAX_MINUTES,
)

result = asyncio.run(
    run_research(
        "How is geothermal energy developing in Europe?",
        config=config,
    )
)
print(result.status)
print(result.final_report)
```

Run it with your project environment activated:

```
python simple_research.py
```

The example prints live progress, the run status and the Markdown report. It does not save a report file by default. Only DeepSeek and Exa credentials are needed here: the example uses DeepSeek V4 Flash for all roles and Exa for Web research. The maximum research window is set to 10 minutes above; other settings remain at their library defaults. Final writing can finish after the research window.

### Updating the installed package

An installation uses a snapshot of the repository; it does **not** update automatically when new commits are published. To fetch and reinstall the current default-branch version, use the matching installer in your project environment:

```
# pip
python -m pip install --upgrade --force-reinstall
beneadie0121
🟧 echo.github ⭐Deep Dog 2 ships as a Python research package with structured delegation and citation validation; its README reports a historical fifth-placbeneadie——

Interpretation history

Decision trace