2026-10-11 16:37 UTC

AllSpark Research claims its released Qwen-derived Iris-mini and Iris-pro search agents achieve 82.2% and 88.6% BrowseComp accuracy with a history-discarding harness, potentially advancing self-hostable search through combined model training and context management.

state: seedheat: mediumuncertainty: mediumconvergesscott: mediumsearch-agents open-models agent-harnessesAllSpark Research

What is this?

AllSpark Research presents Iris-mini and Iris-pro as open-weight web-search agents post-trained from Qwen3.6-35B-A3B and Qwen3.5-397B-A17B, respectively, with model download links in its GitHub repository. The repository reports BrowseComp accuracy of 82.2% and 88.6% using a context-management setting called “discard-all,” versus 64.7% and 72.6% without context management; secondary coverage describes training that alternates supervised fine-tuning and reinforcement learning against live search. These are author-reported results, not independent validation, and the repository explicitly notes that competing systems' scores use their own context-management settings. The supplied snippets do not establish the precise history-discarding mechanism, completeness of the released evaluation harness, or practical self-hosting costs.

Why it matters to Scott

AllSpark’s reported BrowseComp gains from changing context policy converge with Scott’s Context Engineering and Model-Plus-Harness Benchmark Unit positions, providing a concrete search-agent comparison worth testing against his agent-authored compaction and Ask’s lossy history handling—not merely another open-model launch. The supplied radar hits track related harness and context-management evaluations, but not Iris itself; the undisclosed discard mechanism, unverified results and uncertain harness completeness limit any claim that this validates Scott’s result-preserving structural forgetting or offers a practical self-hosted replacement.
ip:framework.context-engineeringip:concept.model-plus-harness-benchmark-unitdev:concept.agent-authored-context-compactiondev:project.askradar:ship-harness-benchradar:agentic-context-management-paperradar:concept.context-managementradar:concept.agent-search
queries asked of Scott's wikis
  • agent harness versus model capability evaluation
  • context management history discard versus persistent agent memory
  • self-hosted search agents local inference economics
  • multi-hop web research evidence gathering projects
  • agent training supervised fine-tuning reinforcement learning tool use
  • benchmark comparability scaffold controls context budgets

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 671h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-13 17:37 (minted)⭐ origin echo-reconstructedAllSpark releases Iris-mini and Iris-pro weights and their evaluation harness, reports benchmark gains from discard-all context management,
AllSpark Research on github (echo) · attributed from hn.story.49686115 · published time unknown
—
09-13 17:06first on hacker news · published · lag ?Iris-mini – open search agent on Qwen3.6-35B-A3B
iamsyr
—
09-13 17:06amplified on hacker news 👑hn.story.49686115
iamsyr
peak 2 · 0 comments · 98% of case engagement
09-13 17:22our radar first saw it · lag ?discovery anchor: hn.story.49686115—
pace: p9 vs 1032 stories at the 336h mark (now 671h old) — behind addom-local-coding-harness (0.5x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnIris-mini – open search agent on Qwen3.6-35B-A3B
Retrieved article excerpt

Open article · Retrieved 2026-09-13T17:23:32.060341+00:00

---

🤗 [**Hugging Face**](https://huggingface.co/collections/AllSpark-Research/iris)  |  
💻 [**GitHub**](https://github.com/AllSpark-Research/Iris)  |  
🔬 [**AllSpark Research**](https://github.com/AllSpark-Research)

# Iris

**Climbing to the Search Frontier.**

Iris-mini (35B-A3B) and Iris-pro (397B-A17B) are open-weight search agents post-trained from the
Qwen3.5/3.6 series. A capable search agent has to decide what to search, how to read what comes
back, when to keep going, and when the evidence it has gathered is enough. Iris is trained for
exactly that loop.

## Performance

|  |  |
| --- | --- |
|  |  |
|  |  |

**30–35B**

| Model | Size | BrowseComp | BrowseComp-ZH | DeepSearchQA | HLE |
| --- | --- | --- | --- | --- | --- |
| MiroThinker-1.7-mini | 30B | 67.9 | 72.3 | – | 36.4 |
| FORT-Searcher | 30B | 72.2 | 75.0 | – | – |
| Apodex-1.0-mini | 35B | 71.5 | 80.6 | 82.2 | 46.8 |
| Nex-N2-mini | 35B | 74.1 | 79.6r | 87.2r | 37.1r |
| Agents-A1 | 35B | 75.5 | – | – | 47.6 |
| XYZ-Aquila-mini | 35B | 78.8 | 82.9 | **89.5** | 51.1 |
| **Iris-mini** | 35B | **82.2** | **84.8** | 86.9 | **52.3** |

**~400B**

| Model | Size | BrowseComp | BrowseComp-ZH | DeepSearchQA | HLE |
| --- | --- | --- | --- | --- | --- |
| MiroThinker-1.7 | 397B | 74.0 | 75.3 | – | 42.9 |
| Apodex-1.0 | 397B | 75.5 | 82.6 | 84.6 | 49.0 |
| Nex-N2-Pro | 397B | 83.7 | 79.6r | 92.3r | 50.0r |
| XYZ-Aquila-pro | 397B | 84.8 | **85.1** | 92.5 | 53.3 |
| **Iris-pro** | 397B | **88.6** | **85.1** | **92.9** | **56.4** |

DeepSearchQA is scored with F1, the rest with accuracy; HLE uses the text-only subset. Iris numbers
use the `discard-all` context-management setting; baselines come from their public reports, each
under its own context management. r reproduced by the XYZ-Aquila team.

## Models

| Model | Base | Params (total / active) | Context | Download |
| --- | --- | --- | --- | --- |
| **Iris-mini** | Qwen3.6-35B-A3B | 35B / 3B | 256K | [🤗 Iris-mini](https://huggingface.co/AllSpark-Research/Iris-mini) |
| **Iris-pro** | Qwen3.5-397B-A17B | 397B / 17B | 256K | [🤗 Iris-pro](https://huggingface.co/AllSpark-Research/Iris-pro) |

## Context Management

Long-horizon search runs out of context before a hard question is resolved, so every serious system
carries some mechanism for this. It is worth enough that a single published number belongs to the
agent and its harness together, which is why we report every benchmark in both regimes, under one
tool set, one context limit, and one judge.

| Setting | BrowseComp | BrowseComp-ZH | DeepSearchQA | HLE |
| --- | --- | --- | --- | --- |
| **Iris-mini** |  |  |  |  |
| w/o | 64.7 | 72.3 | 81.0 | 43.2 |
| retry | – | 83.0 | 89.1 | 52.0 |
| discard-all | 82.2 | 84.8 | 86.9 | 52.3 |
| discard-all + retry | **85.9** | **85.1** | **89.9** | **52.4** |
| **Iris-pro** |  |  |  |  |
| w/o | 72.6 | 76.8 | 86.4 | 50.8 |
| retry | – | 84.1 | 92.3 | **56.6** |
| discard-all | 88.6 | **85.1** | 92.9 | 56.4 |
| discard-all + retry | **90.3** | **85.1** | **93.4** | **56.6** |

`discard-all` clears the accumulated tool history and restarts from the question once the running
context crosses a threshold. `retry` restarts an episode that ended without a parseable answer, carrying forward a short
summary of what was already ruled out. We report `discard-all` as the headline setting even where adding
`retry` scores higher.

## Evaluation

[`Iris-Harness/`](https://github.com/AllSpark-Research/Iris/blob/main/Iris-Harness) is the harness behind every number above: the agent loop, the two
tools, the context-management strategies, the four benchmarks and the graders. It runs against any
OpenAI-compatible endpoint.

```
cd Iris-Harness && uv sync
uv run python data/prepare_data.py
bash scripts/run_eval.sh --base-url http://127.0.0.1:21234/v1 --llm-config iris-mini \
  --benchmarks "browsecomp:0:1" --context-discard-threshold 131072
```

## Acknowledgements

Iris is built on open-source work, and we are grateful to the teams behind it:

- [**MiroThinker**](https://github.com/MiroMindAI/MiroThinker)
- [**Relax**](https://github.com/redai-studio/Relax)
- [**ms-swift**](https://github.com/modelscope/ms-swift)
- [**slime**](https://github.com/THUDM/slime)

---

The data construction and training pipelines are coming soon.

Questions or collaboration: reach us at
[email protected].
iamsyr20
🟧 echo.github ⭐AllSpark releases Iris-mini and Iris-pro weights and their evaluation harness, reports benchmark gains from discard-all context management, AllSpark Research——

Interpretation history

Decision trace