ortegaalfredo, a llama.cpp modifier, publishes a fork (llama.cpp-NLTM) whose README says it adds hot-swappable knowledge injection into the per-layer-embedding (PLE) n-gram table of Qwen3.8 Flash-Next — a separate ~51B-parameter n-gram lookup table that ships with the model — claiming the table is rewritten on every prompt so parts of it can be patched at runtime without reloading model weights, at the cost of hard-to-control output. Surrounding snippets confirm the table itself is real infrastructure the local-LLM community actively manages: users tune where it resides (RAM/SSD offload of the per_layer_token_embd tensor in llama.cpp) and stack it as a drafting source for speculative decoding with >90% acceptance on code, and a result recorded in the case (Nicolodeva's Qwengram-0.8B) shows the table's contents can be frozen and transferred onto a 0.8B backbone for a 5.05% validation-perplexity gain. The specific runtime prompt-driven write claim, however, appears only in the author's own post and README — no supplied snippet shows independent reproduction, upstream llama.cpp endorsement, or Qwen confirmation of the mechanism.
Nicolodeva's quantified result (frozen ~51B-param PLE table transferred onto a 0.8B backbone for −5.05% validation perplexity with a trained reader) is the first independent evidence that a shipped model-internal table is a reusable knowledge substrate — independent experimenters are building Scott's frozen-backbone-plus-separately-owned-learning-layer structure, but in opaque n-gram vectors rather than his human-readable, Git-diffable Soft Weights, making this a live edge-case test of both his Soft Weights definition and the Frozen Model Paradox boundary (which ortegaalfredo's still-unverified runtime write would directly poke). If that write mechanism reproduces with controlled persistence/paraphrase tests, a PLE memory tier becomes a third entry in his per-corpus substrate rule for local agents on gamepc/Ollama-class stacks; until then it stays a medium-heat watch item rather than something to build against.
ip:concept.soft-weightsip:concept.frozen-model-paradoxip:framework.rag-wiki-substrate-ruledev:concept.hardware-aware-local-inferenceradar:sillage-compact-model-memoryradar:infinite-parameter-live-weight-adaptationradar:longcat-sparse-24gb-inferenceradar:concept.agent-memoryradar:concept.model-adaptation
queries asked of Scott's wikis
- frozen model weights inference-time knowledge injection Soft Weights position
- agent memory persistent writable knowledge store hot-swap
- local model RAG versus in-weights knowledge editing
- n-gram table RAM offload local inference economics
- exact-token retrieval paraphrase generalization memory write
- runtime weight patching interference output control
2026-09-26T23:44:58Z
Fourth consecutive velocity-spike trigger resolves to the identical Qwengram artifact — score and comment count literally unchanged (365/97), 1.8 pts/h vs ~98 peak, zero comment flow, and the 'accelerating' momentum label is again a thin-window artifact over a dead age-cohort — so nothing has landed on the case's named reheat triggers and meaning is unchanged. Expire is withheld only because txgsync's stated 'next weekend' retraining window is open right now; the case stays dormant at corroborated/low pending real evidence.
2026-09-26T13:38:22Z
Third consecutive velocity-spike trigger resolves to the same Qwengram post's cooling tail (~2.5 pts/h vs ~98 peak, momentum cooling; the 96th peer percentile reads a dead age-cohort, not renewed spread) — no new implementations, results, author responses, or platform spread since 09-25, and the magnitude-valve flag stays discounted because the second platform is a 2-point HN stub. Meaning is unchanged: PLE storage/transfer is corroborated but practically contested vs plain LoRA, ortegaalfredo's runtime write remains single-source, and the case goes dormant at low heat pending its named reheat triggers.
2026-09-26T10:37:10Z
The velocity spike is the same Qwengram post's tail with a minor late bounce (~5.5 pts/h vs ~98 peak; the 'accelerating' momentum label reads off a thin recent window — noise, not a second wave); no new implementations, results, or author responses since 09-25, so the pending-contradiction reason for holding medium has lapsed and heat cools to low per this case's own stated policy. Meaning is unchanged: PLE transfer/storage is corroborated but practically contested vs plain LoRA, while ortegaalfredo's prompt-driven runtime write remains single-source.
2026-09-26T03:28:51Z
The storage-half result (Qwengram-0.8B, −5.05% perplexity) now carries a credible in-thread counterpoint: TokenRingAI reports their own tests on Qwen 3.5 4B and 35B where plain LoRA on the frozen backbone beats PLE+Adapter+LoRA, and Middle_Bullfrog flags the missing LoRA control that could confound Nicolodeva's gain — so the PLE table's transferability stands as a measurement but its practical value over a plain LoRA becomes the live question. The velocity spike is the same Qwengram post cooling from its peak, not new periphery: no new implementations or platforms since 09-25 (HN item still an empty stub), so the magnitude-valve flag stays discounted.
2026-09-25T13:58:01Z
grounded: converges/medium — Nicolodeva's quantified result (frozen ~51B-param PLE table transferred onto a 0.8B backbone for −5.05% validation perplexity with a trained reader) is the firs
2026-09-25T13:50:03Z
First independent implementation result: Nicolodeva's Qwengram-0.8B freezes Qwen Flash-Next's ~51B-param PLE table onto a 0.8B backbone with a small trained reader/gate for 5.05% lower validation perplexity, demonstrating the PLE table is a reusable, transferable knowledge substrate — this corroborates the storage half of the hypothesis while ortegaalfredo's prompt-driven runtime patching, persistence and output control remain unverified.
2026-09-25T13:25:17Z
evidence attached: reddit.post.1wpvep4 — Independent practitioner freezes Qwen3.8 Flash-Next's 51B-param PLE n-gram memory into a 0.8B model with a trained reader/gate for 5% lower perplexity — evidence the PLE memory is a reusable, transferable knowledge substrate, directly bearing on the PLE case's hypothesis.
2026-09-21T21:54:29Z
The HN attachment supplies only a PLE title, not technical content or corroboration of writable knowledge. Despite the magnitude-valve flag, the supplied cross-platform evidence is a two-point, comment-free listing alongside previously assessed Reddit speculation—not broad uptake of this implementation—so attention remains low.
2026-09-21T21:22:54Z
evidence attached: hn.story.49793198 — This provides useful technical context for the open PLE case by explaining the mechanism behind per-layer embeddings.
2026-09-19T13:31:05Z
The new thread proposes runtime backpropagation confined to the n-gram table, a distinct adaptation mechanism rather than validation of prompt-driven hot-swapping. It adds no demonstrated result or independent reproduction; the periphery remains same-platform speculative discussion, not expanding implementation evidence.
2026-09-19T13:21:56Z
evidence attached: reddit.post.1wkldv3 — This directly explores the same hypothesis that writable in-memory n-gram tables could provide low-cost persistent knowledge in local models.
2026-09-13T22:28:40Z
The new thread adds an independent report of an ongoing small-model retraining experiment using Qwen’s n-gram table, but no results and no reproduction of prompt-driven memory writes. This broadens experimental interest without validating hot-swappable knowledge or changing Scott’s practical options.
2026-09-13T22:21:38Z
evidence attached: reddit.post.1wfkb2t — The discussion directly bears on n-gram and per-layer-embedding approaches to persistent knowledge in local models.
2026-09-12T00:28:42Z
The attached thread is a theoretical proposal, not a second implementation or validation of writable n-gram memory. Its comments sharpen the unresolved questions about learned-vector compatibility and exact-token retrieval, but provide neither demonstrated failure nor corroboration.
2026-09-12T00:22:13Z
evidence attached: reddit.post.1wdwmwh — The proposal adds a concrete writable n-gram layer that is directly relevant to whether local models can gain hot-swappable persistent knowledge.
2026-09-09T18:30:30Z
The refreshed discussion adds no technical evidence separating prompt-driven table mutation from useful, durable knowledge retrieval. The author-reported experiment remains relevant but unvalidated; further routine comment checks offer little value without a controlled write/read result or independent reproduction.
2026-09-07T17:45:15Z
The refreshed comments add aspirations for real-time learning and swappable expertise, not evidence that table mutation yields retrievable, durable knowledge. This remains an author-reported experiment; neither the surrounding local-model activity nor the use of llama.cpp establishes independent validation or upstream endorsement.
2026-09-05T17:29:37Z
The staleness review adds no technical evidence: prompt-driven table mutation remains an author-reported experiment, not demonstrated durable knowledge or reliable behavioral adaptation. Keep it on a slow watch for reproducible write/read tests and persistence results rather than treating speculative discussion as validation.
2026-09-03T16:47:26Z
The refreshed discussion remains speculative enthusiasm rather than independent reproduction or new technical evidence. The experiment still lacks validation of persistence, capacity, and reliable behavioral influence.
2026-09-03T15:56:22Z
Refreshed discussion remains speculative amplification rather than independent reproduction or new implementation evidence. The prototype is still interesting but unvalidated on persistence, capacity, and reliable output influence.
2026-09-03T13:30:31Z
Refreshed comments remain speculative enthusiasm about training, codebase memory, and expert implants; they add no independent reproduction or evidence on persistence, capacity, or output control. The case remains an inspectable but uncorroborated local-memory experiment.
2026-09-03T12:30:23Z
The added attention is repetitive amplification, not independent validation of persistence, capacity, or behavioral control; the inspectable prototype remains an intriguing but uncorroborated local-memory experiment.
2026-09-03T12:27:10Z
grounded: converges/medium — If verified, this is a new local implementation route toward Scott’s frozen-model and Soft Weights position: useful knowledge can be changed at inference time w
2026-09-03T12:24:14Z
case created — The first-party implementation report describes a distinct model-internal memory mechanism with concrete behavior and acknowledged limitations.