2026-10-11 18:04 UTC

LocalLLaMA user brainchillzZ reports that Gufo's headline GitHub benchmark โ€” 70.56 tok/s single-user Qwen3.8-27B Q4 on Strix Halo โ€” holds only under a degenerate prompt; Gufo correcting or clearly qualifying the benchmark methodology, or the community validating the number under representative prompts, resolves whether the project's marquee claim survives scrutiny or costs it adoption.

state: resolvedheat: lowuncertainty: lowconvergesscott: mediumlocal-inference benchmark-integritygufo-orgbrainchillzZ

What is this?

Gufo (gufo-org/gufo) is a newly trending open-source, all-in-one inference engine purpose-built for AMD's Strix Halo APUs, advertising headline figures like 'up to 70.56 tok/s' single-user token generation for Qwen3.8-27B Q4 alongside large prefill and 8-user concurrency gains. Per the case hypothesis, LocalLLaMA user brainchillzZ reports that 70.56 figure only reproduces under a degenerate prompt โ€” but none of the supplied snippets contain brainchillzZ's post or any Gufo response, so the critique itself is unverified here. What the snippets do establish: the number is marketed with an 'up to' qualifier, a Framework-forum test plan explicitly notes Gufo's benchmarks don't cover the single-user agentic case (70โ€“130k context, tool calls, temp 1.0), and independent Strix Halo measurements put realistic 27B-class single-user decode at roughly 20โ€“40 tok/s at working context, well below the headline. Whether the claim survives or costs Gufo adoption hinges on a methodology correction/qualification or community replication under representative prompts.

Why it matters to Scott

Converges with his benchmarking-the-wrong-unit and model-plus-harness canon: an external controlled test is challenging a trending Strix Halo engine's 'up to' headline for holding only under a degenerate prompt, and the Framework test plan's observation that Gufo skips the agentic case is his wrong-unit argument arriving independently from the community. Medium rather than high because it confirms a position he already holds instead of challenging anything load-bearing โ€” the value is dated-receipt material for the wrong-unit concept plus a calibration point (independent ~20โ€“40 tok/s realistic decode vs the 70.56 marketed figure) for the Strix Halo engine-comparison thread his Ollama/hardware-aware-local-inference notes and the radar's inference-benchmarking cluster already track, including whether the number is speculative-decoding acceptance inflation on unrepresentative prompts.
ip:concept.benchmarking-the-wrong-unitip:concept.model-plus-harness-benchmark-unitdev:concept.hardware-aware-local-inferencedev:technology.ollamaradar:concept.inference-benchmarkingradar:concept.benchmark-integrityradar:picchio-llama-cpp-bottleneck-diagnosticsradar:llamacpp-fork-fragmentationradar:atlas-local-inference-engineradar:aa-agentperf-local-benchmark
queries asked of Scott's wikis
  • speculative decoding MTP n-gram acceptance inflation on degenerate benchmark prompts
  • benchmark integrity: 'up to' marketing numbers vs representative agentic workloads
  • Strix Halo inference engine comparison: llama.cpp vs forks vs purpose-built engines decode tok/s
  • local coding agent inference requirements: long context, tool calls, time-to-first-token
  • unified memory APU local LLM serving: real-world decode benchmarks and memory-bandwidth limits
  • llama.cpp tg/pp benchmark methodology criticism and prompt representativeness

Measured heat

no measured readings yet โ€” the hourly heat pass fills this in

How the heat travelled

08-10 14:00โญ origin echo-reconstructedThe gufo-org/gufo repo (Strix Halo inference engine) is the primary artifact publishing the claim the Reddit post examines. Repo description
gufo-org on github (echo) ยท attributed from reddit.post.1wvbmi6
โ€”
10-01 21:18first on r/LocalLLaMA ยท published ยท +1255.3hGufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.
brainchillzZ
โ€”
10-03 06:26first on hacker news ยท published ยท +1288.4hGufo-Qwen3.6-35B-A3B-Q6dense โ€“ 3095tok/s prefill; 190 tok/s decode on Strix Halo
nubela
โ€”
10-01 21:18amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wvbmi6
brainchillzZ
peak 15 ยท 35 comments ยท 79% of case engagement
10-03 06:13amplified on r/LocalLLaMAreddit.post.1wwfw06
nubela
peak 8 ยท 3 comments ยท 17% of case engagement
10-03 06:26amplified on hacker newshn.story.49941824
nubela
peak 1 ยท 0 comments ยท 3% of case engagement
10-03 08:14amplified on r/LocalLLaMAreddit.post.1wwhu1a
hiImMate
peak 1 ยท 0 comments ยท 2% of case engagement
10-01 23:20our radar first saw it ยท +1257.3hdiscovery anchor: reddit.post.1wvbmi6โ€”

Evidence (5) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditGufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.
LocalLLaMA
brainchillzZ935
๐ŸŸง echo.github โญThe gufo-org/gufo repo (Strix Halo inference engine) is the primary artifact publishing the claim the Reddit post examines. Repo descriptiongufo-orgโ€”โ€”
๐ŸŸ  redditgufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halo
LocalLLaMA
nubela194
๐ŸŸง hnGufo-Qwen3.6-35B-A3B-Q6dense โ€“ 3095tok/s prefill; 190 tok/s decode on Strix Halonubela10
๐ŸŸ  redditgufo_windows pre-package for Strix Halo users
LocalLLaMA
hiImMate116

Interpretation history

Decision trace