2026-10-11 17:11 UTC

Independent review and replication will determine whether explicit output-concision instructions reduce LLM inference cost while preserving task accuracy more reliably than input-prompt compression.

state: expiredheat: lowuncertainty: highconvergesscott: mediuminference-economics token-efficiency llm-evaluation

What is this?

A reported nine-model evaluation claims that explicit instructions to produce concise outputs can lower inference cost while preserving task accuracy, whereas compressing input prompts may fail to save money and can sometimes trigger much longer outputs. The supplied snippets support the mechanism—output-token constraints can directly limit generation, and one research-series snippet describes a dramatic “compression paradox”—but they do not identify the study’s authors, repository, benchmark design, or replication status. The earliest-artifact title is truncated, so the provenance and strength of the primary evidence remain unclear.

Why it matters to Scott

The reported result converges with Scott’s harness-level token discipline and evaluation-driven approach: output verbosity should be treated as a controllable, benchmarked cost variable rather than assuming input compression automatically reduces spend. It could affect defaults and evaluations in Ask and related context-compaction systems, but unclear provenance and absent independent replication limit it to a tentative implementation signal rather than a strong dated-receipts opportunity.
ip:concept.model-plus-harness-benchmark-unitip:concept.evaluation-driven-developmentip:concept.token-disciplinedev:project.askdev:concept.agent-authored-context-compactionradar:revision-prompting-token-savingsradar:rtk-coding-agent-cost-regressionradar:concept.token-economicsradar:concept.model-evaluation
queries asked of Scott's wikis
  • output-token economics versus context compression
  • verbosity controls in agent harnesses
  • accuracy-preserving prompt cost optimization
  • LLM evaluation for cost-quality tradeoffs
  • prompt compression causing output inflation
  • default concision policies for coding agents

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditDoes telling an LLM to "be concise" actually save you money? We measured it across 9 models. Compressing the output can save you money and keep accuracy, compressing the input prompt does not. [R]
MachineLearning
ibubbles346312
🟧 echo.github ⭐The earliest identified primary artifact is the repository's initial commit, titled “Cavewoman: How Large Language Models Behave Under LinguMorayo Adeyemi——
🟧 hnClaude has a "Concise" output styleball_of_lint11
🟠 redditSaw the other attempts at fixing "claudish", so I built my own plugin: Unclaudish
ClaudeAI
Global_Mess147518

Interpretation history

Decision trace