2026-10-11 17:19 UTC

Independent benchmarks and artifact review will determine whether the reported sub-2-bit 250M-parameter model can deliver useful local inference from a roughly 60MB deployment while using disk-backed compression for histories approaching 100 million tokens.

state: expiredheat: lowuncertainty: highknownscott: mediumlocal-inference model-compression long-context

What is this?

An unnamed developer reports creating SHADOW 250M Instruct from scratch, training it on 30B English-text tokens plus 0.7B instruction-tuning tokens, and compressing it into an approximately 60MB deployment at sub-2-bit precision. The case also claims disk-backed compression can support histories approaching 100 million tokens, but the supplied results only establish that quantization can reduce local-inference memory requirements and that performance depends heavily on workload and hardware. They do not independently benchmark or validate SHADOWโ€™s quality, deployment size, or long-history capability.

Why it matters to Scott

Scott already holds the relevant positions in `ip:concept.usable-mass-over-unusable-power` and `ip:framework.context-engineering`: deployable small models can outperform unusable power, while durable histories should be externalised and selectively rehydrated rather than treated as giant live context windows. The specific artifact is still worth testing against his `gamepc` local-model substrate and pointer-backed transcript compression work, but until its quality, 60MB footprint, and near-100M-token history claims are independently validated, it does not extend those positions.
ip:concept.usable-mass-over-unusable-powerip:framework.context-engineeringdev:project.gamepcdev:concept.hardware-aware-local-inferencedev:concept.pointer-backed-transcript-compressionip:concept.evaluation-driven-developmentradar:bonsai-extreme-quantizationradar:cachyllama-persistent-kv-cacheradar:concept.extreme-quantizationradar:concept.tiny-modelsradar:concept.long-contextradar:concept.model-evaluation
queries asked of Scott's wikis
  • sub-2-bit quantization quality tradeoffs
  • tiny local models and task-specific utility
  • local inference memory and privacy economics
  • disk-backed agent memory versus context windows
  • compressed histories and long-context retrieval
  • artifact-first evaluation of model claims

Measured heat

no measured readings yet โ€” the hourly heat pass fills this in

How the heat travelled

no chain yet โ€” the hourly chain pass fills this in

Evidence (3) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditI developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB [R]
MachineLearning
Final-Data-141035253
๐ŸŸง echo.other โญThe model card describes SHADOW 250M Instruct as built from scratch, trained on 30B English-text tokens plus 0.7B instruction-tuning tokens,NODEMIND (Sai Kiran Bathula)โ€”โ€”
๐ŸŸ  redditI developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB
LocalLLaMA
Final-Data-141027247

Interpretation history

Decision trace