An unnamed developer reports creating SHADOW 250M Instruct from scratch, training it on 30B English-text tokens plus 0.7B instruction-tuning tokens, and compressing it into an approximately 60MB deployment at sub-2-bit precision. The case also claims disk-backed compression can support histories approaching 100 million tokens, but the supplied results only establish that quantization can reduce local-inference memory requirements and that performance depends heavily on workload and hardware. They do not independently benchmark or validate SHADOWโs quality, deployment size, or long-history capability.
Scott already holds the relevant positions in `ip:concept.usable-mass-over-unusable-power` and `ip:framework.context-engineering`: deployable small models can outperform unusable power, while durable histories should be externalised and selectively rehydrated rather than treated as giant live context windows. The specific artifact is still worth testing against his `gamepc` local-model substrate and pointer-backed transcript compression work, but until its quality, 60MB footprint, and near-100M-token history claims are independently validated, it does not extend those positions.
ip:concept.usable-mass-over-unusable-powerip:framework.context-engineeringdev:project.gamepcdev:concept.hardware-aware-local-inferencedev:concept.pointer-backed-transcript-compressionip:concept.evaluation-driven-developmentradar:bonsai-extreme-quantizationradar:cachyllama-persistent-kv-cacheradar:concept.extreme-quantizationradar:concept.tiny-modelsradar:concept.long-contextradar:concept.model-evaluation
queries asked of Scott's wikis
- sub-2-bit quantization quality tradeoffs
- tiny local models and task-specific utility
- local inference memory and privacy economics
- disk-backed agent memory versus context windows
- compressed histories and long-context retrieval
- artifact-first evaluation of model claims
2026-08-26T18:40:43Z
Repeated reobservations have produced only engagement and recycled discussion, with no independent artifact review, benchmark reproduction, or archive-performance evidence. The implementation remains inspectable but unvalidated, and this attention window has faded without a reason to expect near-term resolution.
2026-08-24T18:26:56Z
The refreshed comments add no independent artifact inspection, benchmark reproduction, or demonstrated archive performance; they continue to recycle known use cases and the 2K live-context versus disk-backed retrieval distinction. The implementation remains inspectable but unvalidated.
2026-08-24T17:25:48Z
The refreshed discussion adds no independent artifact inspection, benchmark reproduction, or demonstrated archive performance; it continues to repeat known use cases and the distinction between a 2K live context and disk-backed retrieval. The implementation remains inspectable but unvalidated.
2026-08-24T15:25:31Z
The refreshed comments add no independent artifact inspection, benchmark reproduction, or demonstrated archive performance; discussion continues to recycle known use cases and the 2K-context versus disk-backed retrieval distinction. The case remains an inspectable but unvalidated tiny-model implementation.
2026-08-24T13:23:46Z
The refreshed comments add no independent artifact inspection, benchmark reproduction, or demonstrated archive performance; they repeat the already-understood distinction between a 2K live context and disk-backed retrieval. The case remains an inspectable but unvalidated tiny-model implementation.
2026-08-24T09:24:06Z
The refreshed discussion remains repetitive amplification of the known 2K live-context and disk-backed archive design. No independent artifact inspection, benchmark reproduction, or archive-performance evidence changes the speculative validation case.
2026-08-24T08:23:10Z
The comment refresh adds only repetitive appreciation and discussion of the already-understood 2K live context plus disk-backed retrieval design. The inspectable artifact remains unvalidated on model quality, footprint, throughput, and archive performance.
2026-08-24T07:27:39Z
The refreshed discussion only repeats appreciation and the already-incorporated distinction between a 2K live context and a disk-backed retrieval archive. No independent artifact review, benchmark reproduction, or demonstrated archive performance changes the case.
2026-08-24T06:22:59Z
A new artifact-specific critique sharpens the architecture distinction: the claimed 100M-token capability appears to be a disk-backed search-and-extraction archive feeding a 2K context, not an extreme context window. This narrows the long-history claim but provides no independent validation of model quality, footprint, throughput, or archive performance.
2026-08-24T05:23:43Z
The released artifact turns the report into an inspectable implementation rather than a claim-only episode, warranting watching status. Its quality, 60MB footprint, throughput, and disk-backed 100M-token history remain author-reported with no independent benchmark or reproduction.
2026-08-24T05:21:59Z
evidence attached: reddit.post.1vwt6m7 โ The released artifact and reported 60MB deployment directly support the open sub-2-bit local-model validation case, though this is author-reported rather than independent corroboration.
2026-08-23T20:31:44Z
The refreshed discussion remains appreciative and speculative, adding no artifact access, independent benchmark, or reproduction. The case still means a technically unusual but creator-originated local-model claim awaiting validation.
2026-08-23T14:28:45Z
The refreshed comments remain questions and appreciative amplification, with no artifact inspection, independent benchmark, or reproduction. The case still means a technically unusual but wholly creator-originated local-model claim awaiting validation.
2026-08-22T14:40:15Z
The refreshed discussion remains repetitive appreciation and questions, with no accessible artifact review, independent benchmark, or reproduction. The case still means a technically unusual but creator-originated local-model claim awaiting validation.
2026-08-22T12:27:28Z
The refreshed comments remain appreciative speculation rather than artifact inspection, benchmarking, or reproduction. The case still represents an unusual but wholly creator-originated local-model claim awaiting substantive validation.
2026-08-22T07:23:17Z
The refreshed discussion is only appreciative amplification and adds no artifact inspection, independent benchmark, or reproduction. The case remains a speculative but testable local-model claim whose unusual compression and disk-backed history design still require validation.
2026-08-22T05:28:56Z
The small engagement increase adds no substantive evidence; the deployment, performance, quality, and 100M-token archive claims remain creator-originated and independently unvalidated. The case still means a testable but speculative local-model artifact awaiting inspection or reproduction.
2026-08-22T05:27:33Z
grounded: known/medium โ Scott already holds the relevant positions in `ip:concept.usable-mass-over-unusable-power` and `ip:framework.context-engineering`: deployable small models can o
2026-08-22T05:25:23Z
origin walked (codex/luna, conf 0.96): anchor reddit.post.1vv2nkh -> echo.other.01df285f37 by NODEMIND (Sai Kiran Bathula)
2026-08-22T05:24:07Z
case created โ The report describes a bounded and technically unusual local-inference system, but no linked artifact or independent benchmark yet supports its performance and long-context claims.