2026-10-11 18:03 UTC

Independent use will determine whether the released ctx-cliff benchmark reproducibly identifies context-length, VRAM-fit, and serving-configuration failure boundaries in local LLM deployments.

state: expiredheat: lowuncertainty: highknownscott: mediumlocal-inference llm-benchmarks inference-toolingcHunter789

What is this?

The case concerns `ctx-cliff.py`, a benchmark script attributed to cHunter789 that is intended to locate local-LLM decode-performance cliffs as context length changes, particularly around VRAM fit and serving configuration. The supplied web results support the underlying problem: practical context limits can fall well below advertised windows, KV-cache memory grows with context, and runtime settings can materially affect usable context and speed. However, the snippets provide no independent results for this specific script, so its reproducibility and ability to distinguish among these failure boundaries remain unestablished.

Why it matters to Scott

Scott already treats local inference as a hardware-aware, evaluation-driven systems problem, and `ctx-cliff.py` could directly test context/VRAM boundaries on his gamepc Ollama substrate. The radar already tracks the closely overlapping failure mode on `radar:ollama-silent-context-truncation`; this script is potentially useful new instrumentation, but without independent results it does not yet establish a new finding or challenge Scott’s position.
dev:project.gamepcdev:concept.hardware-aware-local-inferenceip:concept.evaluation-driven-developmentdev:technology.ollamaradar:ollama-silent-context-truncationradar:llama-cpp-mtp-autofit-memory
queries asked of Scott's wikis
  • local LLM performance-cliff benchmarking
  • context length versus KV-cache and VRAM limits
  • local inference benchmark reproducibility
  • llama.cpp and Ollama serving configuration
  • hardware-aware model and context selection
  • decode throughput collapse and memory spill

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditNew/Old benchmark that provides a lot of answers for local LLM
LocalLLaMA
Pablo_the_brave44
🟧 echo.other ⭐The primary artifact is the author's ctx-cliff.py script, whose docstring says it locates “the decode performance cliff as a function of concHunter789——
🟧 hnWhy your local LLM feels dumber than it isfelineflock417171
🟠 redditWhy your local LLM feels dumber than it is
LocalLLaMA
johnnyApplePRNG03

Interpretation history

Decision trace