2026-10-11 17:11 UTC

GVS5H's authors claim their training-free shared-filesystem orchestration raises Qwen3.8-27B from 69.2% to 92.4% pass@1 on 100 hard LiveCodeBench problems versus Fable 5's 90.4%, potentially achieving frontier-level benchmark accuracy with self-hostable weights through harness design rather than training.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-harnesses local-inference inference-economics coding-agentsslee-persis

What is this?

GVS5H is a project published under the GitHub account slee-persis whose authors describe training-free orchestration: fresh instances of the same model coordinate through a shared filesystem containing a plan, notes, and a current solution. The repository reports tests across nine models on the 100 latest hard LiveCodeBench problems, claiming locally served open-weight Qwen3.8-27B improves from 69.2% to 92.4% pass@1 against Claude Fable 5's 90.4%, while some other models show no improvement or regress. These are author-reported results; the supplied snippets establish neither independent replication nor broader coding parity, and do not explain how fractional percentages were calculated over 100 problems. The explicit 19%-of-cost result concerns orchestrated GPT-5.6-Terra, not Qwen, and the snippets do not establish Qwen's local hardware requirements or total inference costs.

Why it matters to Scott

GVS5H’s fresh-worker/shared-filesystem design converges with Scott’s Markdown OS, Shared Blackboard and Long-Running Agents architectures; its claimed benchmark uplift extends his Model-Plus-Harness Benchmark Unit position into a concrete experiment worth reproducing with trace-backed comparisons, rather than merely repeating the architectural pattern. The supplied radar pages track related harness comparisons, not GVS5H itself, but author-only results, unexplained fractional scoring and missing Qwen compute costs make this a validation opportunity—not established frontier parity or local-inference savings.
ip:concept.model-plus-harness-benchmark-unitip:framework.markdown-osip:framework.long-running-agentsip:concept.shared-blackboarddev:concept.trace-backed-agent-comparisonradar:frontierharness-17x-cost-variationradar:ship-harness-benchradar:concept.agent-harnessesradar:concept.shared-stateradar:concept.coding-agent-benchmarks
queries asked of Scott's wikis
  • coding harness design versus model capability
  • filesystem agent memory shared plans persistent context
  • fresh context agents task decomposition orchestration
  • local open-weight inference economics hardware costs
  • coding benchmark evaluation pass@1 compute budgets

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

09-12 13:25 (minted)⭐ origin echo-reconstructedThe repository introduces ledger-based zero-shot self-orchestration using fresh model instances and a shared filesystem containing a plan, n
slee-persis on github (echo) · attributed from hn.story.49671996 · published time unknown
—
09-12 13:10first on hacker news · published · lag ?GVS5H: Five Qwen3.8-27B Models Match Claude Fable 5 on LiveCodeBench Hard
OakNinja
—
09-12 13:10amplified on hacker news 👑hn.story.49671996
OakNinja
peak 6 · 0 comments · 38% of case engagement
09-16 17:28amplified on hacker newshn.story.49730257
pggues
peak 5 · 1 comments · 38% of case engagement
09-17 13:46amplified on hacker newshn.story.49740680
anon373839
peak 2 · 0 comments · 12% of case engagement
09-23 13:34amplified on hacker newshn.story.49815925
kkm
peak 1 · 0 comments · 7% of case engagement
09-24 14:30amplified on hacker newshn.story.49831159
ibobev
peak 1 · 0 comments · 7% of case engagement
09-12 13:20our radar first saw it · lag ?discovery anchor: hn.story.49671996—

Evidence (6) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnGVS5H: Five Qwen3.8-27B Models Match Claude Fable 5 on LiveCodeBench Hard
Retrieved article excerpt

Open article · Retrieved 2026-09-12T13:22:04.593813+00:00

# GVS5H: Five Qwen3.8-27B Models Match Claude Fable 5 on LiveCodeBench Hard

## Results

[Manager vs single call, four models — LCB-100, 5 passes, 128k max tokens, reasoning ON](https://github.com/slee-persis/GVS5H/blob/master/assets/manager_vs_single_call_four_models.png)

[What one pass costs — LCB-100, 5 passes, single call vs manager, against Fable 5](https://github.com/slee-persis/GVS5H/blob/master/assets/what_one_pass_costs.png)

> **Abstract.** Frontier coding performance is typically bought with larger proprietary models at high cost. We introduce ledger-based zero-shot self-orchestration, a training-free method in which fresh instances of one model decompose problems and coordinate through a shared filesystem holding a plan, notes and current solution. Across nine open and closed-weight models on the 100 latest *hard* LiveCodeBench problems, the method yields gains of up to 23.2 percentage points on pinned backends and offers two routes to frontier-level accuracy. Orchestrated GPT-5.6-Terra reaches 88.0% pass@1 against Fable 5's 90.4% at 19% of the cost, and locally served, open-weight Qwen3.8-27B rises from 69.2% to 92.4%, slightly exceeding Fable 5. Gains are not universal: some models are unchanged or worse. Transcript analysis attributes the gain to decomposition and persistent context. Inference-time organization can approach frontier coding accuracy at a fraction of the cost, or slightly exceed it on self-hostable weights.
>
> — [the paper](https://github.com/slee-persis/GVS5H/blob/master/paper/iclr2027_conference.pdf)

## Running the code

Needs [uv](https://docs.astral.sh/uv/) and an API key for the model you want to test.

```
cd codebase/v2-current
export OPENAI_API_KEY=...

LCB_RELEASE=release_v6 \
ESCALATION_CLOUD_MAX_TOKENS=128000 \
ESCALATION_CLOUD_TIMEOUT=7200 \
MULTIAGENT_MODEL=openai:gpt-5.6-terra \
uv run --no-project --python 3.12 --with 'datasets<4' --with numpy --with anthropic \
  python escalation/run_bench.py --engine multiagent --only lcb --lcb 100 --parallel 8
```

- `--engine multiagent` runs the manager; `--engine single` is the one-call baseline.
- Other models: `anthropic:<model>`, `dashscope:<model>`, `openrouter:<model>`, each with its own `*_API_KEY`.
- The pass@1 score prints at the end. Results are written to `runs/results.json`, workspaces to `runs/ws/`.

## License

Code is under the [MIT License](https://github.com/slee-persis/GVS5H/blob/master/LICENSE). The paper, figures and run data are under
[CC BY 4.0](https://github.com/slee-persis/GVS5H/blob/master/LICENSE-CC-BY-4.0). The LiveCodeBench fork, the benchmark problem statements
and the LaTeX template files keep their own licenses. See [NOTICE.md](https://github.com/slee-persis/GVS5H/blob/master/NOTICE.md) for
which license covers which path.
OakNinja60
🟧 echo.github ⭐The repository introduces ledger-based zero-shot self-orchestration using fresh model instances and a shared filesystem containing a plan, nslee-persis——
🟧 hnQwen3.8 models outperforms Fable 5 Using GVS5H Harnesspggues51
🟧 hnGVS5H: Five Qwen3.8 Models Match Claude Fable 5 on LiveCodeBench Hardanon37383920
🟧 hnUnlocking parallel test-time scaling for long-horizon agentskkm10
🟧 hnUnlocking parallel test-time scaling for long-horizon agentsibobev10

Interpretation history

Decision trace