2026-10-11 17:12 UTC

Burrito Core’s maintainer claims its released GPT-OSS training and inference stack restores reliable tool calling and refusal behavior while sustaining fast 128K-context inference on a single RTX 3090, potentially making GPT-OSS more practical for local agents.

state: expiredheat: lowuncertainty: highconvergesscott: mediumlocal-inference inference-harnesses tool-calling open-modelsiamskeoleskeoleBurrito Core

What is this?

Burrito Core is presented as a GPT-OSS training, evaluation, and inference stack maintained by iamskeole/skeole. Its maintainer claims that an inference harness restores reliable tool calling and refusal behavior while supporting 128K-context inference on one RTX 3090, backed by 320,192 evaluation runs across eight seeds, 3.49B tokens, and 1,062 GPU-hours. The supplied search results are unrelated to the software, so the release, benchmark methodology, performance, and behavioral improvements remain uncorroborated beyond the case’s own titles and summary.

Why it matters to Scott

The claim directly converges with Scott’s position that agent capability belongs to the model-plus-harness unit, and with his Ask compatibility layer for repairing inconsistent tool-call formats. If independently reproduced, the single-3090 long-context stack could be tested against his active Ollama/gamepc substrate and create a dated-receipts opportunity; for now, the claimed evaluation and performance gains remain uncorroborated.
ip:concept.model-plus-harness-benchmark-unitip:concept.evaluation-driven-developmentdev:project.askdev:concept.multi-format-tool-call-parsingdev:concept.hardware-aware-local-inferencedev:project.gamepcradar:concept.agent-harnessesradar:concept.tool-callingradar:concept.local-inferenceradar:stencil-harness-coding-improvementradar:vllm-silent-tool-parser-failuresradar:ctx-cliff-local-inference-benchmark
queries asked of Scott's wikis
  • local long-context inference economics on consumer GPUs
  • tool-calling reliability in agent harnesses
  • harness fixes versus model capability
  • open-weight models for local coding agents
  • behavioral evals for refusals and tool use
  • GPT-OSS deployment and agent integration

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditDon't trust me bro: 3.49B tokens, 320,192 evals, 8 seeds, at batch size 1 over 1,062 GPU hours on a single RTX 3090. And an inference harness that fixes gpt-oss.
LocalLLaMA
skeole028
🟧 hnShow HN: Don't trust me bro: fixing GPT-OSS (3.49B tokens, 1k GPU hours, 1x3090)iamskeole20
🟧 echo.github ⭐The author’s public write-up says: “320,192 runs. 8 seeds ... 1,062 GPU hours on a single RTX 3090.” It identifies burrito-evals as the compiamskeole——

Interpretation history

Decision trace