2026-10-11 17:12 UTC

Independent reproduction and Ollama’s response will determine whether Ollama can silently serve models with materially smaller context windows than configured or advertised and whether explicit runtime checks are required.

state: expiredheat: lowuncertainty: highconvergesscott: highlocal-inference context-windows llm-servingOllamaBigbonus

What is this?

A Show HN report attributed to Bigbonus alleges that Ollama silently served a model advertised or configured for a 40k-token context with only a 4k window, accompanied by a reproduction/checking artifact. Ollama’s documentation confirms a 4096-token default, supports overrides through `OLLAMA_CONTEXT_LENGTH` or `num_ctx`, and recommends checking allocation with `ollama ps`; another reported integration issue describes Ollama-backed models being restricted below their stated context sizes. The supplied snippets do not establish that Bigbonus’s specific 40k-to-4k result was independently reproduced, that an explicit 40k configuration was ignored, or that Ollama has responded.

Why it matters to Scott

The allegation converges with Scott’s production-systems position that configured model capabilities must be observable and validated at runtime rather than trusted from metadata. It directly bears on his shared Ollama endpoint and Ask’s `--ollama` path: if reproduced, silent context reduction could invalidate agent behavior and warrants a fail-fast context-allocation check, though the supplied evidence does not yet establish the failure.
dev:technology.ollamadev:project.gamepcdev:project.askip:framework.12-factor-agents-frameworkip:concept.agent-observabilityradar:concept.context-windowradar:concept.llm-servingradar:concept.local-inferenceradar:concept.long-context-inference
queries asked of Scott's wikis
  • silent context truncation in local inference
  • runtime capability checks for LLM serving
  • advertised versus effective context windows
  • Ollama harness configuration and observability
  • fail-fast validation for model serving
  • context-window defaults in coding agents

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Ollama served my 40k-context model at 4k, silentlysatoshiakiyama10
🟧 echo.github ⭐A reproduction and checking artifact for the claim that Ollama silently served a 40k-context model with a 4k context window.Bigbonus——

Interpretation history

Decision trace