LocalLLaMA user brainchillzZ reports that Gufo's headline GitHub benchmark โ 70.56 tok/s single-user Qwen3.8-27B Q4 on Strix Halo โ holds only under a degenerate prompt; Gufo correcting or clearly qualifying the benchmark methodology, or the community validating the number under representative prompts, resolves whether the project's marquee claim survives scrutiny or costs it adoption.
state: resolvedheat: lowuncertainty: lowconvergesscott: mediumlocal-inference benchmark-integritygufo-orgbrainchillzZ
What is this?
Gufo (gufo-org/gufo) is a newly trending open-source, all-in-one inference engine purpose-built for AMD's Strix Halo APUs, advertising headline figures like 'up to 70.56 tok/s' single-user token generation for Qwen3.8-27B Q4 alongside large prefill and 8-user concurrency gains. Per the case hypothesis, LocalLLaMA user brainchillzZ reports that 70.56 figure only reproduces under a degenerate prompt โ but none of the supplied snippets contain brainchillzZ's post or any Gufo response, so the critique itself is unverified here. What the snippets do establish: the number is marketed with an 'up to' qualifier, a Framework-forum test plan explicitly notes Gufo's benchmarks don't cover the single-user agentic case (70โ130k context, tool calls, temp 1.0), and independent Strix Halo measurements put realistic 27B-class single-user decode at roughly 20โ40 tok/s at working context, well below the headline. Whether the claim survives or costs Gufo adoption hinges on a methodology correction/qualification or community replication under representative prompts.
Why it matters to Scott
Converges with his benchmarking-the-wrong-unit and model-plus-harness canon: an external controlled test is challenging a trending Strix Halo engine's 'up to' headline for holding only under a degenerate prompt, and the Framework test plan's observation that Gufo skips the agentic case is his wrong-unit argument arriving independently from the community. Medium rather than high because it confirms a position he already holds instead of challenging anything load-bearing โ the value is dated-receipt material for the wrong-unit concept plus a calibration point (independent ~20โ40 tok/s realistic decode vs the 70.56 marketed figure) for the Strix Halo engine-comparison thread his Ollama/hardware-aware-local-inference notes and the radar's inference-benchmarking cluster already track, including whether the number is speculative-decoding acceptance inflation on unrepresentative prompts.
ip:concept.benchmarking-the-wrong-unitip:concept.model-plus-harness-benchmark-unitdev:concept.hardware-aware-local-inferencedev:technology.ollamaradar:concept.inference-benchmarkingradar:concept.benchmark-integrityradar:picchio-llama-cpp-bottleneck-diagnosticsradar:llamacpp-fork-fragmentationradar:atlas-local-inference-engineradar:aa-agentperf-local-benchmark
queries asked of Scott's wikis
- speculative decoding MTP n-gram acceptance inflation on degenerate benchmark prompts
- benchmark integrity: 'up to' marketing numbers vs representative agentic workloads
- Strix Halo inference engine comparison: llama.cpp vs forks vs purpose-built engines decode tok/s
- local coding agent inference requirements: long context, tool calls, time-to-first-token
- unified memory APU local LLM serving: real-world decode benchmarks and memory-bandwidth limits
- llama.cpp tg/pp benchmark methodology criticism and prompt representativeness
Measured heat
no measured readings yet โ the hourly heat pass fills this in
How the heat travelled
Evidence (5) โ โญ canonical anchor
Interpretation history
2026-10-03T08:46:35Z
The dispute is settled in substance: representative-workload measurements (~40 tok/s agentic) and independent 20โ40 tok/s realistic-decode data proved brainchillzZ's critique, the mechanism (spec-decode acceptance on greedy output) became thread consensus, and Gufo's silence plus an expanding Windows port show the inflated headline costs no adoption. The episode closes as a settled receipt for best-case benchmark marketing rather than an open question.
2026-10-03T08:24:40Z
evidence attached: reddit.post.1wwhu1a โ Unofficial Windows pre-package widens gufo's adoption surface and adds a representative-workload data point (~40 tok/s on agentic tasks) bearing on whether the disputed headline number survives real-prompt scrutiny.
2026-10-03T08:24:40Z
evidence attached: hn.story.49941824 โ shared external link with case evidence
2026-10-03T07:06:04Z
The critique's substance is now corroborated โ brainchillzZ's test, independent ~20โ40 tok/s realistic-decode measurements, and the Framework test plan's agentic-gap note all point the same way โ but the episode is resolving socially toward normalization: the thread is drifting down (15โ9, ratio 0.78โ0.66) with replies shrugging 'everybody advertises best case,' Gufo has not responded, and the new 35B-A3B headline numbers continue the same marketing pattern on a different model without touching the disputed 27B figure.
2026-10-03T06:30:14Z
evidence attached: reddit.post.1wwfw06 โ New headline Strix Halo numbers (3095 tok/s prefill, 190 tok/s decode) from the same gufo project fall directly under the open benchmark-methodology dispute.
2026-10-02T00:10:09Z
origin walked (opencode/cheap-glm, conf 0.85): anchor reddit.post.1wvbmi6 -> echo.github.c56aa392ac by gufo-org
2026-10-01T23:42:00Z
grounded: converges/medium โ Converges with his benchmarking-the-wrong-unit and model-plus-harness canon: an external controlled test is challenging a trending Strix Halo engine's 'up to' h
2026-10-01T23:33:13Z
case created โ A controlled user test directly challenging an open inference engine's headline marketing number is a bounded, resolvable benchmark-honesty episode, not routine chatter.
Decision trace
- 10-03 18:46resolveThe dispute is settled in substance: representative-workload measurements (~40 tok/s agentic) and independent 20โ40 tok/s realistic-decode data proved brainchillzZ's critique, the mechanism (spec
- 10-03 18:24attachUnofficial Windows pre-package widens gufo's adoption surface and adds a representative-workload data point (~40 tok/s on agentic tasks) bearing on whether the disputed headline number survives r
- 10-03 18:24attachshared external link with case evidence
- 10-03 18:23propose_attachUnofficial Windows pre-package widens gufo's adoption surface and adds a representative-workload data point (~40 tok/s on agentic tasks) bearing on whether the disputed headline number survives r
- 10-03 18:21propose_attachshared external link with case evidence
- 10-03 17:06repriceThe critique's substance is now corroborated โ brainchillzZ's test, independent ~20โ40 tok/s realistic-decode measurements, and the Framework test plan's agentic-gap note all point the
- 10-03 16:30attachNew headline Strix Halo numbers (3095 tok/s prefill, 190 tok/s decode) from the same gufo project fall directly under the open benchmark-methodology dispute.
- 10-03 16:25propose_attachNew headline Strix Halo numbers (3095 tok/s prefill, 190 tok/s decode) from the same gufo project fall directly under the open benchmark-methodology dispute.
- 10-02 16:21sensor_dirtycomment_update
- 10-02 10:10promote_anchororigin walk conf 0.85
- 10-02 09:42groundConverges with his benchmarking-the-wrong-unit and model-plus-harness canon: an external controlled test is challenging a trending Strix Halo engine's 'up to' headline for holding only
- 10-02 09:33createA controlled user test directly challenging an open inference engine's headline marketing number is a bounded, resolvable benchmark-honesty episode, not routine chatter.