Burrito Core is presented as a GPT-OSS training, evaluation, and inference stack maintained by iamskeole/skeole. Its maintainer claims that an inference harness restores reliable tool calling and refusal behavior while supporting 128K-context inference on one RTX 3090, backed by 320,192 evaluation runs across eight seeds, 3.49B tokens, and 1,062 GPU-hours. The supplied search results are unrelated to the software, so the release, benchmark methodology, performance, and behavioral improvements remain uncorroborated beyond the case’s own titles and summary.
The claim directly converges with Scott’s position that agent capability belongs to the model-plus-harness unit, and with his Ask compatibility layer for repairing inconsistent tool-call formats. If independently reproduced, the single-3090 long-context stack could be tested against his active Ollama/gamepc substrate and create a dated-receipts opportunity; for now, the claimed evaluation and performance gains remain uncorroborated.
ip:concept.model-plus-harness-benchmark-unitip:concept.evaluation-driven-developmentdev:project.askdev:concept.multi-format-tool-call-parsingdev:concept.hardware-aware-local-inferencedev:project.gamepcradar:concept.agent-harnessesradar:concept.tool-callingradar:concept.local-inferenceradar:stencil-harness-coding-improvementradar:vllm-silent-tool-parser-failuresradar:ctx-cliff-local-inference-benchmark
queries asked of Scott's wikis
- local long-context inference economics on consumer GPUs
- tool-calling reliability in agent harnesses
- harness fixes versus model capability
- open-weight models for local coding agents
- behavioral evals for refusals and tool use
- GPT-OSS deployment and agent integration
2026-09-08T00:25:52Z
The refreshed discussion adds no technical results beyond the already-known template incorporation, and no concrete validation milestone is pending. Retire active monitoring as faded, not disproved; independent reliability tests or measured single-3090 long-context performance would justify reopening.
2026-09-06T00:22:57Z
This staleness check adds no substantive evidence: the downstream template incorporation remains a narrow adoption signal, not confirmation of reliability or long-context performance. Keep the stack as an optional local-agent experiment rather than a validated improvement, and reduce polling frequency.
2026-09-03T23:28:03Z
The refreshed discussion is repetitive amplification and skepticism, with no new implementation results, independent reproduction, or benchmark audit. The case remains a testable maintainer-reported harness whose core reliability and single-3090 performance claims are unvalidated.
2026-09-02T22:45:16Z
A downstream user says they incorporated Burrito Core’s fixes into an existing GPT-OSS Jinja template, adding a small implementation/adoption signal. This still does not independently reproduce the claimed reliability gains, refusal behavior, or 128K single-3090 performance, so the case remains unvalidated.
2026-09-01T22:23:35Z
The refreshed discussion adds no independent reproduction, benchmark audit, or technical counterevidence; it is further repetitive reaction to the maintainer’s existing claims. The case remains a released, testable harness with potentially useful local-agent implications, but not a validated advance.
2026-09-01T20:56:00Z
The refreshed comments remain repetitive skepticism about presentation and model choice, with no independent reproduction, benchmark audit, or technical counterevidence. The case still means a testable maintainer-reported harness release rather than a validated advance.
2026-09-01T19:02:55Z
The refreshed discussion adds an adjacent agent-side parser idea but no independent reproduction, benchmark audit, or substantive counterevidence. The case remains a testable maintainer-reported implementation awaiting external receipts.
2026-09-01T16:55:13Z
No independent reproduction or new technical evidence arrived; the added discussion is mostly skepticism about clarity and model choice rather than validation or substantive counterevidence. The released stack remains testable and relevant, but the episode has cooled pending receipts from another implementer.
2026-09-01T16:48:25Z
grounded: converges/medium — The claim directly converges with Scott’s position that agent capability belongs to the model-plus-harness unit, and with his Ask compatibility layer for repair
2026-09-01T16:46:22Z
origin walked (codex/luna, conf 0.97): anchor hn.story.49523381 -> echo.github.8fb39667c6 by iamskeole
2026-09-01T16:44:59Z
case created — The Reddit and Hacker News observations trace to the same substantial evaluation and released implementation, so they form one moving episode.