2026-10-11 16:36 UTC

Firelex claims his released Jeff-Code β€” a 0.8B decision model inside Pi's agent loop that takes routine steps itself and routes Qwen 3.8-27B's thinking β€” cuts time per coding task to 0.68x at an unchanged pass rate across 1,242 paired benchmark tasks, and independent replication or adoption makes small-decision-model-in-the-loop a standard coding-agent acceleration.

state: seedheat: lowuncertainty: mediumconvergesscott: highagentic-coding inference-economics local-inference

What is this?

Jeff is a set of three tiny open-weight decision models (0.8B and 2B Qwen3.5, Gemma4-E2B) released 28–29 September 2026 by a developer under the GitHub handle firelex: each answers a closed 'which of these options fits?' question in a single forward pass (~22 ms on an RTX PRO 6000, ~28 ms on an M4 Max), descends from the MIT-licensed AutoJev recipe, and was trained on one workstation GPU in roughly two to 3.5 hours; the README's own Doom/Frogger/Pac-Man harness shows the less-accurate 0.8B playing repeated loops better than the 2B, i.e. one-step benchmark scores don't certify loop performance. The supplied snippets do NOT corroborate the case's specific claim β€” no snippet mentions a 'Jeff-Code' variant inside Pi's coding-agent loop, the step-delegation/thinking-routing of Qwen 3.8-27B, or the 0.68x/1,242-paired-task result β€” so that rests entirely on the case's own evidence title and should be marked unverified. The surrounding territory is real and crowded: InternLM shipped competing Intern-Decision 0.8B/2B/4B decision models two days before Jeff, hosted decision model Jev (83.0%) is the narrow accuracy benchmark Jeff's 2B edges at 83.1%, Qwen 3.8-27B is a validated open coding model at ~$0.15/M effective agent-loop input cost, and Inherent Labs' Faraday independently demonstrates the small-orchestrator-routes-big-executor pattern at 27B scale.

Why it matters to Scott

Independently arrives at economics Scott already argues β€” the Model Barbell's cheap end and the Scout-and-Senior claim that agents overpay frontier prices for non-judgment work β€” but at exactly the 'generic small-model/large-model routing' seam his own Scout–Senior Split page explicitly parks as adjacent-not-a-sufficient-instance, now carrying 1,242-paired-task receipts at unchanged pass rate: both a dated-receipts publishing opportunity and a measured refinement of the barbell's boundary. It also lands on active work β€” Jeff descends from the AutoJev recipe and edges Jev at 83.1% (dev:project.jev), and a ~22ms one-pass local decision model is a weekend-scale insert into the ask loop's cheap/smart chain on his gamepc stack β€” with the grounding's caveat that the supplied snippets do not corroborate the Jeff-Code result itself, so the replication watch is the load-bearing condition.
ip:concept.model-barbellip:framework.scout-senior-splitip:source.the-scout-and-the-senior-ebookdev:project.jevdev:project.askradar:concept.model-routingradar:concept.local-inferenceradar:concept.inference-economicsradar:concept.small-language-modelsradar:system-one-harness-typed-action-loopradar:holstered-skill-routing-hookradar:openai-decisions-apiradar:qwen38-dflash2-long-context-speedup
queries asked of Scott's wikis
  • agent loop small model router step delegation harness
  • local inference latency cost economics coding agent
  • model cascade routing cheap fast model guards expensive model
  • synthetic data distillation fine-tune tiny classifier
  • paired benchmark evaluation unchanged pass rate speedup
  • compounding errors long-horizon agent loop reliability

Measured heat

now 0 pts/hpeak 1 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 171h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

10-04 13:00⭐ origin echo-reconstructedOriginal README announcement: "A coding agent where a 0.8B decision model works alongside Qwen 3.8-27B. Coding tasks 47% faster (32% less ti
firelex (GitHub user; Hugging Face handle mstrasser) β€” same person who posted to HN on github (echo) Β· attributed from hn.story.49968868
β€”
10-05 18:51first on hacker news Β· published Β· +29.9hJeff-Code: a 0.8B model makes Qwen 3.8-27B coding 47% faster, same pass rate
firelex
β€”
10-05 18:51amplified on hacker news πŸ‘‘hn.story.49968868
firelex
peak 2 Β· 0 comments Β· 98% of case engagement
10-05 20:21our radar first saw it Β· +31.4hdiscovery anchor: hn.story.49968868β€”
pace: p22 vs 1188 stories at the 168h mark (now 171h old) β€” ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (2) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnJeff-Code: a 0.8B model makes Qwen 3.8-27B coding 47% faster, same pass rate
Retrieved article excerpt

Open article Β· Retrieved 2026-10-05T20:39:45.318088+00:00

[Jeff](https://jeffhub.ai)

# Jeff-Code

A coding agent where a 0.8B decision model works alongside Qwen 3.8-27B.  
**Coding tasks 47% faster (32% less time) on average, at the same pass rate.**

[**jeffhub.ai**](https://jeffhub.ai) Β·
[Jeff](https://github.com/firelex/jeff) Β·
[code adapter](https://huggingface.co/mstrasser/jeff-adapter-code) Β·
[code-router adapter](https://huggingface.co/mstrasser/jeff-adapter-code-router) Β·
[Pi's README](https://github.com/firelex/jeff-code/blob/main/README-pi.md)

Jeff-Code is a fork of [Pi](https://pi.dev), the coding agent by Mario Zechner and the Pi contributors. It puts
[Jeff](https://github.com/firelex/jeff), a 0.8B decision model, inside Pi's agent loop. Around every Qwen turn, Jeff
makes two quick decisions (about 0.2 s each):

1. **Can I take the next step myself?** When Jeff is confident, it picks the tool and its argument and runs it. It
   takes information steps (read a file, list a folder, search the code, check which tools are installed), and it can
   also run the tests or a build, repeat Qwen's last command and, with the run-approval setting the evaluation used,
   run a script Qwen wrote or install a missing package. Writing and editing files always stay with Qwen. Jeff can
   take several steps in a row, and Qwen then starts its turn with the results already in front of it. Whenever Jeff
   is unsure, it hands over to Qwen.
2. **Does Qwen need to think hard on this turn?** Thinking stays off unless Jeff's probability that the turn needs
   full thinking reaches 0.6. In the evaluation, about three quarters of Qwen's turns ran with thinking off.

Each decision is a small LoRA adapter on the same Jeff v1.3 base.

## Results

Run side by side in paired blocks: each task ran under every setting at the same time, on the same Qwen server, and
every comparison is paired by task. The baseline is Qwen 3.8-27B alone in the same build with every Jeff feature
switched off, thinking at full on every turn and no thinking limit, which is how plain Pi runs it.

|  | Jeff-Code (threshold 0.6) vs Qwen alone |
| --- | --- |
| **Pass rate** | 62.4% vs 62.8%; paired difference βˆ’0.2 points (95% interval βˆ’2.6 to +2.1), 1,242 paired tasks |
| **Time per task, on average** | **0.68Γ—** (0.64-0.72; geometric mean of per-task time ratios), median 0.70Γ— |
| **Total time, all tasks combined** | 0.86Γ— (0.80-0.93) |
| **Per benchmark** | SWE-bench Verified 0.63Γ—, SWE-rebench 0.66Γ—, Terminal-Bench Pro 0.64Γ—, Harbor Index 0.71Γ—; no clear speed-up on Terminal-Bench 2.0 (0.96Γ—) or SkillsBench (0.91Γ—) |
| **Thinking off on every turn instead** | faster still, but βˆ’7.6 points (βˆ’10.6 to βˆ’4.5); βˆ’13.5 on Terminal-Bench 2.0 |
| **A less cautious Jeff (threshold 0.7)** | βˆ’4.5 points |

- **Why total time drops less than time per task:** in about 5% of tasks, Jeff-Code runs more than 30 minutes longer
  than Qwen alone, because it keeps going where Qwen alone gives up after a few minutes. The good news is, sometimes that pays off: in
  those tasks Jeff-Code solved 26 to Qwen's 24.
- **Only tasks Jeff never saw in training.** SWE-bench Verified ran in full (500 tasks; none of its repositories were
  used for training). For the benchmarks we also trained on, the tasks were split and every held-out task was run
  (a few pairs hit by repeated infrastructure failures are left out, see Exclusions): Terminal-Bench 2.0 (40 tasks,
  3 attempts each), SWE-rebench (189, two rounds), Terminal-Bench Pro (100, two rounds), SkillsBench (44) and Harbor
  Index (41). Terminal-Bench (original) and Terminal-Bench Science also ran, but Qwen
  solves almost none of their tasks in any setting, so they are left out of the pooled numbers.
- **Exclusions:** task pairs hit by an infrastructure failure (out of memory, a stalled session, a test environment
  that would not start) were run once more; 27 pairs that failed again are left out for both sides. About 10 long
  re-runs were still running when these numbers were taken. One Terminal-Bench 2.0 task, pytorch-model-recovery, is
  left out of every comparison: a harness bug stopped the baseline sessions before they began.
- **Thinking-off comparison:** it had the same safeguards and thinking limit as Jeff-Code (below); the only difference
  is Jeff's decisions.

### Per benchmark

| Benchmark | Paired tasks | Qwen alone | Jeff-Code | Difference, points (95% interval) | Time per task |
| --- | --- | --- | --- | --- | --- |
| SWE-bench Verified | 486 | 70.6% | 70.8% | +0.4 (βˆ’3.5 to +4.3) | 0.63Γ— (0.57-0.69) |
| SWE-rebench (2 rounds) | 370 | 58.9% | 58.1% | βˆ’0.8 (βˆ’5.7 to +3.5) | 0.66Γ— (0.61-0.73) |
| Terminal-Bench Pro (2 rounds) | 195 | 61.2% | 62.8% | +1.5 (βˆ’4.6 to +7.7) | 0.64Γ— (0.55-0.76) |
| Terminal-Bench 2.0 (40 tasks, 3 attempts each) | 108 | 75.9% | 70.0% | βˆ’4.6 (βˆ’12.1 to +3.7) | 0.96Γ— (0.78-1.16) |
| SkillsBench | 42 | 28.6% | 31.0% | +2.4 (βˆ’11.9 to +16.7) | 0.91Γ— (0.68-1.20) |
| Harbor Index | 41 | 12.2% | 9.8% | βˆ’2.4 (βˆ’12.2 to +7.3) | 0.71Γ— (0.51-0.99) |
| **All six, pooled** | **1,242** | **62.8%** | **62.4%** | **βˆ’0.2 (βˆ’2.6 to +2.1)** | **0.68Γ— (0.64-0.72)** |

No benchmark shows a clear pass-rate difference: every interval includes zero. SWE-rebench and Terminal-Bench Pro ran twice; their two rounds are combined, with intervals computed over both. Time per task is the geometric mean of
the per-task time ratios (Jeff-Code's time divided by Qwen alone's), below 1 is faster. The pass rates count every
finished session; the difference counts only tasks finished in both settings, so it is not exactly the gap between the
two pass rates. Terminal-Bench (original) and Terminal-Bench Science also ran, but Qwen alone and Jeff-Code both solve
0% of their tasks, so they are left out.

Full report: [results/imitation/eval-tonight.md](https://github.com/firelex/jeff-code/blob/main/results/imitation/eval-tonight.md).

## How the adapters were trained

Jeff predicts what Qwen would do next.

- **Steps (`jeff-adapter-code`):** each label is the step Qwen actually took next in recorded Qwen sessions, built by
  code, with no other model involved.
- **Thinking (`jeff-adapter-code-router`):** each recorded Qwen turn at full thinking was asked again with thinking
  off, then low, then medium. The label is the cheapest level whose action was as good as the original, or full
  thinking if none was. "As good" is decided by code wherever possible (the same kind of step on the same target);
  otherwise Qwen3.8-Max, with thinking off, judges whether the cheaper step would serve the task just as well. These
  strict labels make the router cautious, and that is what keeps the pass rate.

## What changed from Pi

- **Thinking per turn:** Pi switches thinking on or off for the whole session. Jeff-Code sets Qwen's thinking level
  for each turn, either fixed or decided by Jeff.
- **Loop guard:** catches repeated or near-identical actions up to six steps back (a file write only counts as progress
  if it changes the file). A caught repeat is thrown away and the turn is asked again with full thinking; near-identical
  outputs or two failed commands in a row send the next turn to full thinking.
- **Runaway cut-off:** if Qwen's thinking or text keeps repeating itself, the reply is stopped and asked again with
  full thinking.
- **Thinking limit:** at 8,000 thinking tokens, Qwen answers from what it has thought so far. This also rescues replies
  that would otherwise hit the 32K output limit, which ends a Pi session.
- **Jeff steps:** before each Qwen turn, Jeff-Code builds a menu of concrete next steps from what is already known, and
  Jeff takes them when it is confident.
- **Jeff server pool and trace logging:** one Jeff server per GPU behind one address, and every Qwen request and Jeff
  decision is logged as JSON Lines.

The code is in [packages/coding-agent/src/core/jeff-first/](https://github.com/firelex/jeff-code/blob/main/packages/coding-agent/src/core/jeff-first); the
evaluation and training tools are in [tools/jeff-first/](https://github.com/firelex/jeff-code/blob/main/tools/jeff-first).

## Running it

You need Qwen 3.8-27B behind an OpenAI-compatible server (for example vLLM) and a Jeff server with the two adapters.

**1. Jeff server.** Install [Jeff](https://github.com/firelex/jeff), then download the base and both adapters (the
folder names are the adapter names Jeff-Code asks for):

```
hf download mstrasser/jeff-base --revision v1.3 --local-dir jeff-base-v1.3
hf download mstrasser/jeff-adapter-code --revision v1.3 --local-dir adapters/jeff-step
hf download mstrasser/jeff-adapter-code-router --revision v1.3 --local-dir adapters/jeff-router
JEFF_CHECKPOINT=jeff-base-v1.3 JEFF_ADAPTERS=adapters JEFF_DEVICE=cuda JEFF_HOST=0.0.0.0 PORT=8920 \
  JEFF_CPU_THREADS=8 python tools/jeff-first/jeff_serve.py      # with the Python of the Jeff environment
```

`jeff_serve.py` starts Jeff's own server and adds the endpoint that fits long prompts into Jeff's 8,192 tokens. For
many parallel sessions, run one server per GPU behind [tools/jeff-first/jeff\_pool.py](https://github.com/firelex/jeff-code/blob/main/tools/jeff-first/jeff_pool.py).
GGUF versions for llama.cpp are on Hugging Face (`-gguf`); run the router on the Q8\_0 base, since about 6% of its
decisions change at Q4\_K\_M.

**2. Jeff-Code with Qwen.** Build the repository (`npm install && npm run build`) and add Qwen to `~/.jeff/agent/models.json`
as an OpenAI-compatible model with `"reasoning": true`, `"compat": {"thinkingFormat": "qwen-chat-template"}` and
`"maxTokens": 32768` (see [the model docs](https://github.com/firelex/jeff-code/blob/main/packages/coding-agent/docs/models.md)).

**3. Switch Jeff on** with the settings the evaluation used:

```
export JEFF_FIRST_MODE=jeff
export JEFF_FIRST_JEFF_URL=http://localhost:8920
export JEFF_FIRST_JEFF_STEP_ADAPTER=jeff-step
export JEFF_FIRST_JEFF_STEP_THRESHOLD=0.40               # Jeff's top option needs 0.40, otherwise Qwen takes over
export JEFF_FIRST_THINKING_ROUTER=jeff-off-unless:jeff-router:0.6
export JEFF_FIRST_THINKING_LIMIT=8000
export JEFF_FIRST_OUTPUT_TRIM=off                      # required; leave off
export JEFF_FIRST_RUN_APPROVAL=all                       # all, seen or never: may Jeff run scripts Qwen wrote and install packages
export JEFF_FIRST_DRIVER_BUILD=qwen3.8-27b               # the exact Qwen build, written to the trace
export JEFF_FIRST_TRACE_FILE=$HOME/jeff-code/trace.jsonl # its folder must exist
export JEFF_FIRST_TASK_ID=my-project
./jeff-test.sh
```

Every setting is required in `jeff` mode, and a missing or invalid one stops Jeff-Code with an error starting with
`JeffFirst:`. Unset `JEFF_FIRST_MODE` for plain Pi.

**Status:** research code. It was built and measured inside benchmark containers (Harbor); interactive use works, but
has had far less testing.

## Licence and credit

Jeff-Code is a fork of [Pi](https://github.com/earendil-works/pi) and keeps Pi's MIT licence ([LICENSE](https://github.com/firelex/jeff-code/blob/main/LICENSE)).
All credit for the agent itself goes to Pi's authors; Pi's own README is in [README-pi.md](https://github.com/firelex/jeff-code/blob/main/README-pi.md). We'd happily
upstream whatever Pi wants to take. The Jeff adapters are Apache 2.0.
firelex20
🟧 echo.github ⭐Original README announcement: "A coding agent where a 0.8B decision model works alongside Qwen 3.8-27B. Coding tasks 47% faster (32% less tifirelex (GitHub user; Hugging Face handle mstrasser) β€” same person who posted to HNβ€”β€”

Interpretation history

Decision trace