2026-10-11 18:01 UTC

Rakuen Software's aimee project claims released vLLM plugins (Qwen 3.8, Gemma 4) and a preprint give local transformer and Mamba-class models native, context-free access to a self-learning external knowledge store without retraining; replication of the preprint and real plugin adoption resolve whether native non-context memory is a practical local-LLM layer.

state: seedheat: lowuncertainty: mediumconvergesscott: highagent-memory local-inference model-architectureRakuen Software

What is this?

Rakuen Software's aimee project claims to give local LLMs โ€” transformer models like Qwen 3.8 and Gemma 4, plus Mamba-class architectures โ€” native, context-free access to a self-learning external knowledge store via released vLLM plugins and an accompanying preprint, with no retraining required. The referenced substrate is real and current: Qwen 3.8 27B is a released dense multimodal model, Gemma 4 is Google DeepMind's April 2026 open model family, and vLLM supports both Qwen and Mamba architectures, so the plugin targets are technically coherent. However, the supplied search returned no trace of Rakuen Software, the aimee project, the plugins, or the preprint itself โ€” nothing here verifies the claims, only the plausibility of the underlying stack. One practical caveat from the snippets: vLLM has open incompatibilities around hybrid Mamba/GDN attention (batch-invariant support blocked entirely), which is precisely the serving layer a Mamba-class memory plugin would sit on.

Why it matters to Scott

Rakuen has independently shipped, as released plugins plus a preprint, the position Scott's canon already argues โ€” learning lives in an external, self-learning knowledge store because the frozen weights can't retain it โ€” but claims the piece his doctrine lacks: context-free native access from the vLLM serving layer, which would bypass the attention budget his context-engineering work treats as the binding constraint and offer his wiki substrate an integration path beside the MCP tool surface. That makes the artifacts worth reading both as dated receipts and as a rival architecture, with the caveat that nothing external verifies Rakuen or the preprint, so plugin-code replication is the crux.
ip:concept.frozen-model-paradoxip:concept.scaffolding-hypothesisip:framework.wiki-is-the-kernelip:concept.attention-budgetip:concept.rag-as-sensordev:concept.llm-navigated-wikidev:project.mcp-ip-wikiradar:zero-mem-pi-retrievalradar:sillage-compact-model-memoryradar:llama-cpp-hot-swappable-ple-memoryradar:kimi-linear-local-validationradar:vllm-tenstorrent-plugin
queries asked of Scott's wikis
  • agent memory external knowledge store architecture
  • non-context memory vs RAG for agents
  • vLLM plugin development local inference tooling
  • Mamba state-space vs transformer architecture notes
  • local LLM knowledge layer without retraining
  • model-native retrieval vs context-window stuffing

Measured heat

now 0 pts/hpeak 4 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 266h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-30 14:00โญ origin echo-reconstructedPreprint "Shared Native Memory: Expanding Knowledge Without Retraining" (DOI 10.5281/zenodo.23077865, v1 created 2026-10-01T08:04:05Z). Abst
Jared Bailes (Rakuen Software / the Aimee project) on paper (echo) ยท attributed from reddit.post.1wzbnbr
โ€”
10-06 19:22first on r/LocalLLaMA ยท published ยท +149.4hNative memory for local LLMs,
One-Arugula1163
โ€”
10-06 19:22amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wzbnbr
One-Arugula1163
peak 6 ยท 4 comments ยท 101% of case engagement
10-06 20:20our radar first saw it ยท +150.3hdiscovery anchor: reddit.post.1wzbnbrโ€”
pace: p38 vs 1188 stories at the 168h mark (now 266h old) โ€” ahead of agentsec-static-config-auditing (1.2x), behind agent-trace-tampering (0.8x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditNative memory for local LLMs,
LocalLLaMA
One-Arugula116364
๐ŸŸง echo.paper โญPreprint "Shared Native Memory: Expanding Knowledge Without Retraining" (DOI 10.5281/zenodo.23077865, v1 created 2026-10-01T08:04:05Z). AbstJared Bailes (Rakuen Software / the Aimee project)โ€”โ€”

Interpretation history

Decision trace