2026-10-11 16:37 UTC

llama.cpp modifier ortegaalfredo claims Qwen’s in-memory PLE n-gram table can be patched from prompts without reloading model weights, potentially providing local models with a low-cost form of hot-swappable persistent knowledge despite limited output control.

state: corroboratedheat: lowuncertainty: highconvergesscott: mediumagent-memory local-inference model-adaptationortegaalfredollama.cppQwen

What is this?

ortegaalfredo, a llama.cpp modifier, publishes a fork (llama.cpp-NLTM) whose README says it adds hot-swappable knowledge injection into the per-layer-embedding (PLE) n-gram table of Qwen3.8 Flash-Next — a separate ~51B-parameter n-gram lookup table that ships with the model — claiming the table is rewritten on every prompt so parts of it can be patched at runtime without reloading model weights, at the cost of hard-to-control output. Surrounding snippets confirm the table itself is real infrastructure the local-LLM community actively manages: users tune where it resides (RAM/SSD offload of the per_layer_token_embd tensor in llama.cpp) and stack it as a drafting source for speculative decoding with >90% acceptance on code, and a result recorded in the case (Nicolodeva's Qwengram-0.8B) shows the table's contents can be frozen and transferred onto a 0.8B backbone for a 5.05% validation-perplexity gain. The specific runtime prompt-driven write claim, however, appears only in the author's own post and README — no supplied snippet shows independent reproduction, upstream llama.cpp endorsement, or Qwen confirmation of the mechanism.

Why it matters to Scott

Nicolodeva's quantified result (frozen ~51B-param PLE table transferred onto a 0.8B backbone for −5.05% validation perplexity with a trained reader) is the first independent evidence that a shipped model-internal table is a reusable knowledge substrate — independent experimenters are building Scott's frozen-backbone-plus-separately-owned-learning-layer structure, but in opaque n-gram vectors rather than his human-readable, Git-diffable Soft Weights, making this a live edge-case test of both his Soft Weights definition and the Frozen Model Paradox boundary (which ortegaalfredo's still-unverified runtime write would directly poke). If that write mechanism reproduces with controlled persistence/paraphrase tests, a PLE memory tier becomes a third entry in his per-corpus substrate rule for local agents on gamepc/Ollama-class stacks; until then it stays a medium-heat watch item rather than something to build against.
ip:concept.soft-weightsip:concept.frozen-model-paradoxip:framework.rag-wiki-substrate-ruledev:concept.hardware-aware-local-inferenceradar:sillage-compact-model-memoryradar:infinite-parameter-live-weight-adaptationradar:longcat-sparse-24gb-inferenceradar:concept.agent-memoryradar:concept.model-adaptation
queries asked of Scott's wikis
  • frozen model weights inference-time knowledge injection Soft Weights position
  • agent memory persistent writable knowledge store hot-swap
  • local model RAG versus in-weights knowledge editing
  • n-gram table RAM offload local inference economics
  • exact-token retrieval paraphrase generalization memory write
  • runtime weight patching interference output control

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 916h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-03 11:40⭐ origin directly observedQwen-3.8-Next-Flash Ngram Hot-Swappable Knowledge Injector for llama.cpp
ortegaalfredo on r/LocalLLaMA
—
09-11 23:45first on r/LocalLLaMA · published · +204.1hLearning/RSI through ngrams?
RapidRaid
—
09-21 20:49first on hacker news · published · +441.2hPer-Layer Embeddings (PLE)
Bluestein
—
09-03 11:40amplified on r/LocalLLaMAreddit.post.1w64y26
ortegaalfredo
peak 238 · 61 comments · 34% of case engagement
09-11 23:45amplified on r/LocalLLaMAreddit.post.1wdwmwh
RapidRaid
peak 15 · 11 comments · 3% of case engagement
09-13 21:29amplified on r/LocalLLaMAreddit.post.1wfkb2t
OuterKey
peak 55 · 21 comments · 9% of case engagement
09-19 12:48amplified on r/LocalLLaMAreddit.post.1wkldv3
Mrinohk
peak 1 · 5 comments · 1% of case engagement
09-21 20:49amplified on hacker newshn.story.49793198
Bluestein
peak 2 · 0 comments · 0% of case engagement
09-25 12:46amplified on r/LocalLLaMA 👑reddit.post.1wpvep4
Nicolodeva
peak 377 · 100 comments · 54% of case engagement
09-03 12:20our radar first saw it · +0.7hdiscovery anchor: reddit.post.1w64y26—
pace: p90 vs 519 stories at the 720h mark (now 916h old) — ahead of openai-agents-api (1.1x), behind sanotts-microcontroller-tts (1.0x)

Evidence (6) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Qwen-3.8-Next-Flash Ngram Hot-Swappable Knowledge Injector for llama.cpp
LocalLLaMA
ortegaalfredo23860
🟠 redditLearning/RSI through ngrams?
LocalLLaMA
RapidRaid1511
🟠 redditAre there any experimental small models (9B or less) which are out or currently being trained using engram/n-grams?
LocalLLaMA
OuterKey5221
🟠 redditRe-writable n-grams?
LocalLLaMA
Mrinohk15
🟧 hnPer-Layer Embeddings (PLE)Bluestein20
🟠 redditQwengram-0.8B: I transferred Qwen3.8 Flash-Next’s n-gram memory into Qwen3.5-0.8B — 5.05% lower validation perplexity
LocalLLaMA
Nicolodeva376100

Interpretation history

Decision trace