Independent reproduction and Ollama’s response will determine whether Ollama can silently serve models with materially smaller context windows than configured or advertised and whether explicit runtime checks are required.
state: expiredheat: lowuncertainty: highconvergesscott: highlocal-inference context-windows llm-servingOllamaBigbonus
What is this?
A Show HN report attributed to Bigbonus alleges that Ollama silently served a model advertised or configured for a 40k-token context with only a 4k window, accompanied by a reproduction/checking artifact. Ollama’s documentation confirms a 4096-token default, supports overrides through `OLLAMA_CONTEXT_LENGTH` or `num_ctx`, and recommends checking allocation with `ollama ps`; another reported integration issue describes Ollama-backed models being restricted below their stated context sizes. The supplied snippets do not establish that Bigbonus’s specific 40k-to-4k result was independently reproduced, that an explicit 40k configuration was ignored, or that Ollama has responded.
Why it matters to Scott
The allegation converges with Scott’s production-systems position that configured model capabilities must be observable and validated at runtime rather than trusted from metadata. It directly bears on his shared Ollama endpoint and Ask’s `--ollama` path: if reproduced, silent context reduction could invalidate agent behavior and warrants a fail-fast context-allocation check, though the supplied evidence does not yet establish the failure.
dev:technology.ollamadev:project.gamepcdev:project.askip:framework.12-factor-agents-frameworkip:concept.agent-observabilityradar:concept.context-windowradar:concept.llm-servingradar:concept.local-inferenceradar:concept.long-context-inference
queries asked of Scott's wikis
- silent context truncation in local inference
- runtime capability checks for LLM serving
- advertised versus effective context windows
- Ollama harness configuration and observability
- fail-fast validation for model serving
- context-window defaults in coding agents
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-22T10:28:11Z
The allegation produced no independent reproduction, Ollama response, or further technical evidence within its observation horizon, so it remains unvalidated and has faded as an active episode.
2026-08-20T09:39:32Z
No independent reproduction, Ollama response, or new technical evidence has appeared; the unchanged observation leaves the allegation consequential but unvalidated and cools the case pending confirmation.
2026-08-20T09:33:31Z
grounded: converges/high — The allegation converges with Scott’s production-systems position that configured model capabilities must be observable and validated at runtime rather than tru
2026-08-20T09:31:16Z
case created — The linked first-party reproduction artifact alleges a silent correctness failure in a widely used local-inference runtime.
Decision trace
- 08-22 20:28expireThe allegation produced no independent reproduction, Ollama response, or further technical evidence within its observation horizon, so it remains unvalidated and has faded as an active episode.
- 08-22 20:28alert_silentThe only delta is elapsed staleness; there is no new consequential evidence to route, and repeating the original unvalidated allegation would add no value.
- 08-22 20:28alert_routeThe only delta is elapsed staleness; there is no new consequential evidence to route, and repeating the original unvalidated allegation would add no value.
- 08-20 19:39repriceNo independent reproduction, Ollama response, or new technical evidence has appeared; the unchanged observation leaves the allegation consequential but unvalidated and cools the case pending confirmat
- 08-20 19:39alert_silentThere is no new consequential delta beyond the already-routed allegation, so another alert would only repeat unvalidated information.
- 08-20 19:39alert_routeThere is no new consequential delta beyond the already-routed allegation, so another alert would only repeat unvalidated information.
- 08-20 19:36alert_shadowThe report is not yet independently validated or confirmed by Ollama, but it includes a concrete checking artifact and directly affects Scott’s Ollama-backed workflows. Knowing today enables a quick e
- 08-20 19:36alert_routeThe report is not yet independently validated or confirmed by Ollama, but it includes a concrete checking artifact and directly affects Scott’s Ollama-backed workflows. Knowing today enables a quick e
- 08-20 19:33groundThe allegation converges with Scott’s production-systems position that configured model capabilities must be observable and validated at runtime rather than trusted from metadata. It directly bears on
- 08-20 19:31createThe linked first-party reproduction artifact alleges a silent correctness failure in a widely used local-inference runtime.