2026-10-11 17:10 UTC

Independent use will determine whether Picchio reliably exposes llama.cpp layer placement and separates prefill from decode performance well enough to prevent misleading local-inference benchmarks.

state: expiredheat: lowuncertainty: highknownscott: mediumllama-cpp local-inference inference-benchmarking

What is this?

The supplied snippets establish that llama.cpp benchmarking should distinguish prompt evaluation (prefill) from token generation (decode), since the two stages can perform very differently, and that direct llama.cpp exposes controls relevant to hardware/layer placement. The evidence title suggests a tool or method called Picchio diagnosed a 14× slowdown for Qwen3.8-27B on an RTX 4070 Super, but none of the snippets identifies Picchio, who built it, how it exposes placement, or any independent validation. The web answer’s claim that independent use confirms Picchio’s reliability is therefore unsupported by the supplied results.

Why it matters to Scott

The methodological position is already held in Scott’s Hardware-aware local inference and Model-Plus-Harness Benchmark Unit pages: accelerator placement must be explicit, and performance claims must disclose the runtime harness rather than collapse distinct stages into one throughput number. Picchio could become practically useful for Scott’s gamepc benchmarking if independent use validates its placement diagnostics and prefill/decode separation, but the supplied evidence establishes neither, so this currently adds no validated finding.
dev:concept.hardware-aware-local-inferenceip:concept.model-plus-harness-benchmark-unitip:concept.observabilitydev:project.gamepcradar:concept.llama-cppradar:concept.local-inferenceradar:concept.model-evaluationradar:concept.benchmark-integrity
queries asked of Scott's wikis
  • prefill vs decode benchmark methodology
  • llama.cpp GPU layer placement visibility
  • local inference benchmark reproducibility
  • misleading tokens-per-second metrics
  • CPU GPU offload diagnostics
  • local model performance harnesses

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐I found why Qwen3.8-27B was 14× slower on my 4070 Super
LocalLLaMA
luckokkkk019

Interpretation history

Decision trace