2026-10-11 17:16 UTC

Independent replication will determine whether revision prompting reduces decoded tokens and serving cost by 2–10× in repeatedly updated structured-output workflows without reducing correctness.

state: expiredheat: lowuncertainty: highconvergesscott: mediuminference-economics llm-tooling

What is this?

Revision prompting is presented as a technique for repeatedly updated structured outputs that shifts work from relatively slow decoded output tokens to cheaper prefilled input tokens, with a claimed 2–10× reduction in decoded tokens and serving cost while preserving correctness. The supplied snippets establish that structured-output pipelines commonly use schema validation, constrained decoding, and retries, and that token economics vary by model and workload. However, they do not document an independent replication of revision prompting or substantiate the specific 2–10× and no-correctness-loss claims; no creator or research group is identified.

Why it matters to Scott

The proposed technique operationalizes Scott’s existing prefill-versus-decode economics in a structured-output revision loop, while its claimed 2–10× savings could materially shift the economic threshold between incremental repair and full regeneration. Independent correctness and cost benchmarks would therefore bear directly on his validation-gated extraction systems and Nuke and Regenerate doctrine, but the supplied evidence does not yet substantiate the claimed gains.
ip:concept.prefix-caching-economicsip:concept.regeneration-economicsip:framework.nuke-and-regeneratedev:concept.validation-gated-llm-extractionradar:concept.inference-economicsradar:concept.llm-toolingradar:concept.llm-serving
queries asked of Scott's wikis
  • prefill versus decode token economics
  • incremental revision versus full structured-output regeneration
  • schema validation and constrained decoding tradeoffs
  • agent state patching versus complete regeneration
  • structured-output retry loops and correctness
  • inference cost benchmarks for repeated workflows

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Revision Prompting: Trades slow (decoded) output tokens for cheap (prefilled) input tokens.
LocalLLaMA
Dry_Rabbit_1123147
🟧 hnRevision Prompting: improves industrial LLM processesidiliv10

Interpretation history

Decision trace