2026-10-11 17:12 UTC

Independent runs will determine whether dspy-factorio’s combined RLM and GEPA approach enables agents to make sustained progress on Factorio’s long-horizon tasks.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-harnesses rlm gepa long-horizon-agentsukituki

What is this?

dspy-factorio is presented as an open implementation for teaching agents to play Factorio by combining RLM with GEPA, a DSPy optimizer that uses natural-language reflection and evolutionary search to improve prompts from execution feedback. Factorio provides a long-horizon test environment in which agents must form sub-objectives and can be evaluated across independent runs using production scores and technology milestones. The supplied snippets do not directly document dspy-factorio’s results, ukituki’s role, or independent replication, so the claim that this specific combination enables sustained progress remains to be established.

Why it matters to Scott

The implementation independently combines execution-feedback-driven harness evolution with a long-horizon environment, converging with Scott’s Reflexive Agent Design, Replay-Driven Design Evolution, and model-plus-harness evaluation positions. Independent Factorio runs could provide useful evidence about whether those loops produce measurable sustained progress, but no results or replication are yet supplied, limiting the present significance.
ip:framework.reflexive-agent-designip:framework.replay-driven-design-evolutionip:framework.long-running-agentsip:concept.model-plus-harness-benchmark-unitip:concept.measurable-convergenceradar:concept.agent-harnessesradar:concept.long-horizon-agentsradar:concept.agent-evaluationradar:prime-agent-harness-validation
queries asked of Scott's wikis
  • long-horizon agent progress and trajectory management
  • reflective prompt optimization from execution traces
  • agent harness evaluation with independent runs
  • Factorio or simulation environments for coding agents
  • RLM recursive reasoning and context management
  • DSPy GEPA evolutionary prompt optimization

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Teaching agents how to play Factorio (using RLM, GEPA)ukituki10
🟧 echo.github ⭐An open implementation for teaching agents to play Factorio using RLM and GEPA.ukituki——

Interpretation history

Decision trace