Retrieved article excerpt
Open article · Retrieved 2026-09-13T17:23:31.132494+00:00
# Deep Dog 2 — research engine
Turn a question into a cited Markdown report. A supervisor plans the research, delegates work to platform agents, reviews their findings, and writes the final report.
Deep Dog 2 installs directly from GitHub as a **Python package** in your existing project. Call `run_research()` with just a question to use the defaults, or pass a per-run configuration to choose models, search, agent types and research budgets. See the [quickstart](https://github.com/beneadie/deep_dog_2#quickstart) for pip and uv installation, or [clone the source](https://github.com/beneadie/deep_dog_2#work-from-a-source-checkout) to edit the engine and run the repository examples.
## Official benchmark results
At the time of publication, Deep Dog 2 ranked **5th overall** and **1st among open-source research agents** on the [DeepResearch Bench](https://huggingface.co/spaces/muset-ai/DeepResearch-Bench-Leaderboard). The published run used a relatively economical profile: a 15-minute research window, a 20-iteration supervisor cap, at most 3 Exa searches per sub-agent, DeepSeek V4 Pro as supervisor, and DeepSeek V4 Flash for sub-agents.
[DeepResearch Bench rankings showing Deep Dog 2 in fifth place](https://github.com/beneadie/deep_dog_2/blob/main/assets/deep-dog-2-ranking.png)
- **Role-specific models** — use different models for the supervisor, initial draft, and platform sub-agents.
- **Reflection and delegation** — the supervisor plans research, delegates distinct questions, evaluates findings, and decides whether more work is needed.
- **Platform specialization** — Web is enabled by default; Reddit, Substack, General, PubMed, Arxiv, and SEC agents can be enabled through configuration.
- **Evidence-first reports** — findings are collected into a source registry, then final inline citations and the `## Sources` section are validated before the report is returned.
- **Operational control** — time limits, iteration caps, search budgets, fallback chains, and output modes can be tuned for local experiments or more economical deployments.
## Features
- **Supervisor + sub-agent architecture** — configurable platform agents in [config.py](https://github.com/beneadie/deep_dog_2/blob/main/deep_research/config.py); `ResearchWeb` is enabled by default, with Reddit, Substack, General, PubMed, Arxiv and SEC agents available through configuration
- **Multiple search providers** — Tavily and/or Exa (`WEB_SEARCH_ENGINE` at `deep_research/config.py:658`)
- **Provider-agnostic models** — DeepSeek, MiMo, Meta Muse, Gemini, OpenAI, GLM, and OpenRouter-hosted models via a single `get_model()` factory (`deep_research/config.py:883`)
- **Model fallback chains** — per-role fallback lists (`SUBAGENT_MODEL_FALLBACK_CHAIN`, `SUPERVISOR_MODEL_FALLBACK_CHAIN`, `DRAFT_REPORT_MODEL_FALLBACK_CHAIN`)
- **Cited reports** — Markdown report, source metadata and optional research trace; applications can save the returned results
- **LangGraph execution** — recursion limit, timeouts, and observability logging
## How It Works
Deep Dog 2 retains the draft-first, iterative refinement idea from Deep Dog 1, while making reflection, delegation, and source handling explicit:
```
User question
|
v
clarify_with_user → write_research_brief → write_draft_report
|
v
supervisor research loop
┌───────────────┼────────────────┐
│ │ │
reflect delegate conclude research
(think_tool) (parallel agents)
|
v
Web / Reddit / Substack / ...
|
v
findings + sources
|
└── repeat until complete
|
v
final report → citation validation → output
```
1. **Scope the question.** The input is converted into a structured research brief.
2. **Create a scaffold.** An initial draft establishes a useful report structure before live research begins. It is not treated as evidence.
3. **Reflect and delegate.** The supervisor uses internal reflection to identify gaps and delegates focused, non-overlapping research tasks to platform agents.
4. **Research in parallel.** Agents search, read, save, and compress findings using the tools available for their platform. Their results are returned with source metadata and citations.
5. **Evaluate.** The supervisor reviews the findings and may request another round.
6. **Finalize the report.** The final writer combines the brief, draft, and findings. Citation checks validate the relationship between inline citations and the final sources list.
Supervisor reflection is used for planning and control; it is not copied into the final research report. Optional subtopic evaluation and parallel subtopic reports can run after the main report when enabled in `deep_research/config.py`.
The supported prompt family is `OPEN`. Older configurations using `LEGACY` must switch to `OPEN`.
## Official Benchmark Results
At the time of publication, Deep Dog 2 ranked **5th overall** and **1st among open-source research agents** on the [DeepResearch Bench](https://huggingface.co/spaces/muset-ai/DeepResearch-Bench-Leaderboard) benchmark. These results were achieved with a relatively economical configuration: a maximum of 3 Exa searches per sub-agent, 15 minutes of research time, 20 total iterations, and DeepSeek V4 Pro and DeepSeek V4 Flash as the supervisor and sub-agent base models.
| Metric | Score |
| --- | --- |
| Overall | **0.5432** |
| Comprehensiveness | 0.5468 |
| Insight | 0.5532 |
| Instruction Following | 0.5426 |
| Readability | 0.5105 |
These results are a historical reproducibility profile, not a promise about current defaults. The current package defaults to DeepSeek V4 Flash for all model roles and uses a different supervisor iteration default. Benchmark rankings and scores may change as the leaderboard changes.
### Estimated cost comparison
The following is a pricing-based estimate for one research task. It is not a controlled cost benchmark: the systems use different architectures, search providers and token budgets.
| System and configuration | Estimated cost per task | Basis |
| --- | --- | --- |
| **Deep Dog 2** — DeepSeek V4 Flash (off-peak) + Exa; 15-minute maximum, 20 supervisor iterations, at most 3 searches per sub-agent | **$0.25–$0.60** | Approximately **$0.20–$0.40** in DeepSeek inference and **$0.05–$0.20** in Exa usage for this configuration |
| **Gemini Deep Research** — Google’s published typical-task estimate | **$1–$3** | Google estimates about 80 search queries, 250k input tokens and 60k output tokens for a moderate task |
Deep Dog 2’s range is an estimate based on the [DeepSeek V4 pricing (off-peak)](https://api-docs.deepseek.com/quick_start/pricing/) and [Exa pricing](https://exa.ai/pricing), using off-peak DeepSeek rates. Actual cost varies with prompt length, model output, cache hits, the number of delegated agents and how quickly the supervisor concludes. Exa currently advertises **$20 in sign-up credits and $10 in credits each month**; [Tavily](https://docs.tavily.com/documentation/api-credits) provides **1,000 free credits per month**, equivalent to $8 at its $0.008 per-credit pay-as-you-go rate.
Costs are easy to change by changing the configuration: shorten the research window, lower the supervisor or sub-agent iteration limits, reduce searches per sub-agent, or use fewer agents. [Nemotron 3.5 Lightning](https://openrouter.ai/nvidia/nemotron-3.5-lightning) has also worked successfully as a sub-agent model in Deep Dog 2 and costs approximately half as much as DeepSeek V4 Flash in the tested setup. Google’s Gemini figures are its own published estimates; see the [Gemini Deep Research pricing section](https://ai.google.dev/gemini-api/docs/deep-research#estimated-costs) for the assumptions behind them.
For a detailed explanation of the reflection and delegation methods used, see the engineering article [Deep Dog 2: How Reflection and Structured Delegation Improve Supervisor–Subagent Research Systems](https://beneadie01.substack.com/p/deep-dog-2-how-reflection-delegation).
## Quickstart
Requires **Python 3.11+**, **Git**, a DeepSeek API key and an Exa API key. In your own project directory, use your existing Python environment or create one below.
Create and activate an environment if you do not already have one
Choose standard Python tooling:
```
# Standard Python tooling
python -m venv .venv
```
Or, using [uv](https://docs.astral.sh/uv/):
```
uv venv --python 3.11 --seed
```
`--seed` includes pip so either installer below works in this environment.
Activate the environment:
```
# macOS / Linux
source .venv/bin/activate
```
```
# Windows PowerShell
.venv\Scripts\Activate.ps1
```
**Install the package directly from GitHub.** Choose pip or uv:
```
# pip
python -m pip install "git+https://github.com/beneadie/deep_dog_2.git"
```
Or, with uv:
```
uv pip install "git+https://github.com/beneadie/deep_dog_2.git"
```
The installer fetches the repository, builds the package and installs its dependencies. You do not need a local source checkout or a PyPI release. This works because the repository includes Python packaging metadata in [pyproject.toml](https://github.com/beneadie/deep_dog_2/blob/main/pyproject.toml); a Git repository needs a Python package build configuration to support this kind of install. The installed distribution is named `deep-dog-2`; the Python import is `deep_research`.
**Add your keys** to a `.env` file in your project directory:
```
DEEPSEEK_API_KEY=your-deepseek-key
EXA_API_KEY=your-exa-key
```
Keep `.env` out of version control by adding it to your project's `.gitignore`. If those keys are already set as environment variables, no `.env` file is needed; existing environment variables take precedence.
**Create `simple_research.py` in your project** with the following code. The two settings near the top are optional examples: change the model or maximum research time, and leave the rest of the configuration at its defaults:
```
import asyncio
from dotenv import load_dotenv
load_dotenv()
from deep_research.integration import RunConfig, run_research
MODEL = "deepseek-v4-flash"
MAX_MINUTES = 10
config = RunConfig(
supervisor_model_fallback_chain=[MODEL],
subagent_model_fallback_chain=[MODEL],
draft_report_model_fallback_chain=[MODEL],
research_time_max_minutes=MAX_MINUTES,
)
result = asyncio.run(
run_research(
"How is geothermal energy developing in Europe?",
config=config,
)
)
print(result.status)
print(result.final_report)
```
Run it with your project environment activated:
```
python simple_research.py
```
The example prints live progress, the run status and the Markdown report. It does not save a report file by default. Only DeepSeek and Exa credentials are needed here: the example uses DeepSeek V4 Flash for all roles and Exa for Web research. The maximum research window is set to 10 minutes above; other settings remain at their library defaults. Final writing can finish after the research window.
### Updating the installed package
An installation uses a snapshot of the repository; it does **not** update automatically when new commits are published. To fetch and reinstall the current default-branch version, use the matching installer in your project environment:
```
# pip
python -m pip install --upgrade --force-reinstall