2026-10-11 17:14 UTC

Skillmem's maintainers claim their released local memory layer reinforces coding procedures only with external evidence and reserves rule approval for owners, enabling reusable cross-session skills without automatically promoting agent-written memories into trusted instructions.

state: seedheat: mediumuncertainty: mediumknownscott: lowagent-memory coding-agents agentic-securityliza-studio

What is this?

The case describes Skillmem, attributed to liza-studio, as a released local procedural-memory layer for coding agents, with evidence titles pointing to a Show HN announcement and repository documentation for SQLite storage, reinforcement and decay, and Claude Code/Codex integration. Its central maintainer claim is that external evidence is required to reinforce procedures and only owners can approve rules, rather than letting agent-written memories automatically become trusted instructions. None of the supplied web snippets identifies Skillmem or its maintainers; they describe other memory-and-skills systems, so they do not independently establish this release, its implementation, or its claimed trust controls.

Why it matters to Scott

Skillmem’s claimed external-evidence reinforcement and owner-only approval repeat positions Scott already holds in CLAUDE.md Pattern, Verification Loops, and Recommendation–authority separation; the radar also tracks a closely related implementation claim in Backpass, though not Skillmem itself. The supplied material neither independently verifies these controls nor establishes a consequential new adoption or capability that would change Scott’s dev-wiki design, so this is presently another example of his existing position rather than a new publishing or implementation opportunity.
ip:concept.claude-md-patternip:concept.verification-loopsdev:concept.recommendation-authority-separationradar:backpass-evidence-gated-memory-editsradar:astrum-hsam-memory-provenanceradar:slowave-adaptive-local-memory
queries asked of Scott's wikis
  • procedural memory reusable coding workflows across sessions
  • agent-written memory promotion trusted instructions owner approval
  • evidence-gated learning verification reinforcement memory decay
  • persistent memory poisoning provenance trust boundaries
  • Claude Code Codex local memory integration

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 553h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-18 15:35 (minted)⭐ origin echo-reconstructedThe repository documents SQLite-backed procedural memory, evidence-based reinforcement and decay, Claude Code and Codex integration, and ver
liza-studio on github (echo) · attributed from hn.story.49755605 · published time unknown
—
09-18 15:16first on hacker news · published · lag ?Show HN: Skillmem – local memory for coding agents that stores how, not what
mrPetrukovich
—
09-18 15:16amplified on hacker news 👑hn.story.49755605
mrPetrukovich
peak 2 · 0 comments · 98% of case engagement
09-18 15:20our radar first saw it · lag ?discovery anchor: hn.story.49755605—
pace: p9 vs 1032 stories at the 336h mark (now 553h old) — behind addom-local-coding-harness (0.5x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Skillmem – local memory for coding agents that stores how, not what
Retrieved article excerpt

Open article · Retrieved 2026-09-18T15:23:02.707844+00:00

# skillmem

[CI](https://github.com/liza-studio/skillmem/actions/workflows/ci.yml)

**Self-improving skills for Claude Code and Codex — your agents learn, recall, reinforce, and forget.**

[skillmem demo: a Russian query finds an English skill, unused skills decay](https://github.com/liza-studio/skillmem/blob/main/docs/demo.gif)

Strength has to be earned — saying a skill helped is not evidence, a passing test is:

[skillmem: self-report does not raise strength, a passing test does, and rare rules can be pinned](https://github.com/liza-studio/skillmem/blob/main/docs/demo-evidence.svg)

Generated from a real run: `scripts/demo.sh --record | python3 scripts/cast_to_svg.py > docs/demo-evidence.svg`.

skillmem gives Claude Code and the Codex CLI a local, persistent skill & memory layer. After every non-trivial task the agent can record *how it was done* as a skill; before the next task it recalls the relevant ones; skills that keep proving useful get stronger, and skills nobody uses fade away — the way human memory works.

- **$0 per write and per read** — no LLM calls, no cloud, no API keys. Plain SQLite on your disk.
- **Bilingual hybrid search, fully local** — FTS5 BM25 + Snowball stemming (EN/RU) matches inflected forms within a language; the multilingual ONNX embedder is what lets a Russian query find an English skill, so install the `semantic` extra if you work across both. All on CPU, offline.
- **Ebbinghaus strength model, earned not claimed** — strength rises only on evidence from outside the agent's own judgement, falls after a failure, and fades on a schedule when unused; dead skills are swept to a backed-up archive (never deleted). Rules that are rare by nature can be pinned out of decay.
- **Provenance, and trust the owner grants** — every memory records where it came from (`owner` / `agent` / `imported` / `derived`), and only the owner approves one as a rule (`skillmem trust <slug>`). Anything unapproved — an imported pack, a summary of a transcript that quoted a web page, a rule an agent was talked into saving — is injected inside a marked block that says it is data, not instructions. Editing an approved memory drops the approval with it.
- **Tamper-evident history** — every edit is appended to a SHA256 hash-chain; `skillmem verify` detects any after-the-fact tampering.
- **Deep Claude Code integration** — hooks on five events + 10 MCP tools installed with one command.
- **One memory, several agents** — Claude Code and Codex share a single database, and every
  record carries the agent that wrote it, taken from the MCP handshake, so authorship stays
  readable when they learn side by side.
- **Cross-platform** — macOS (launchd), Windows (schtasks), Linux (systemd user timers, cron fallback).
- **No vendor lock** — `export-all` dumps everything to plain markdown with YAML frontmatter; re-importing the dump yields the same records. One destination per database: the exporter prunes its own stale files via a manifest and will not judge another database's.

## Why

Agents repeat their mistakes because each session starts from zero. Existing "memory" tools store facts; skillmem stores *procedures* — trigger, steps, outcome, lessons — and ranks them by how often they actually helped. The write path costs nothing, so the agent can afford to learn from every task.

## What 0.10.0 changed

Memory that an agent writes is not the same thing as a rule you set, and until 0.10.0 this
project treated them the same. An external text — a README, a web page — reaches a transcript,
a model distils it into a note, and the note comes back in the next session under a heading
that reads like your own rules. A document could also talk an agent into saving a rule through
`mem_learn`, and that rule looked exactly like one you wrote.

Now provenance is a field, trust is an act, and the summariser that reads your transcripts runs
with **no tools at all** (`--tools ""` plus `--strict-mcp-config`; a CLI that does not understand
those flags gets no recap rather than an uncaged one). The full list — including the migration
and what it does and does not approve on upgrade — is in the [CHANGELOG](https://github.com/liza-studio/skillmem/blob/main/CHANGELOG.md).

The seven releases before it, in one line each, because they were all about the same hook:
0.9.3 stopped the Stop hook recursing into itself (one machine spawned 4083 summary sessions in
a day); 0.9.4 put a rate limit on it and stopped a failing model buying a call per turn; 0.9.5
fixed four silent defects, including recall being dead for notebook edits; 0.9.6 stopped a slow
summary overwriting a fresher one; 0.9.7 added `skillmem recap` and `skillmem hooks-status`;
0.9.8 stopped a skipped turn reading a 59 MB transcript first; 0.9.9 made publishing a summary
compare-and-swap. **Anyone on 0.9.0–0.9.2 should upgrade** — those versions contain the
recursion.

## How it differs

The memory products in this space — Mem0, Zep, Letta, LangMem, Cognee — are built mostly for
conversational and user memory, entity graphs, or agent-managed context, and most of them offer
a hosted tier. skillmem is narrower on purpose and different on four axes:

|  | skillmem |
| --- | --- |
| **What it stores** | procedures — trigger, steps, outcome, lessons — not facts about a user |
| **What it forgets** | actively: unused skills decay on an Ebbinghaus schedule and are archived; rare-but-critical rules are pinned out of it |
| **Where strength comes from** | outside evidence only — a passing test, an accepted diff, your confirmation. An agent saying "that helped" moves recency, never strength, so it cannot promote its own mistake. `reinforce` is not idempotent: a retried confirmation counts again (evidence ids are a later release) |
| **Who is trusted** | you. Provenance is recorded, approval is yours to give, and unapproved memory arrives framed as data |
| **Where it runs** | your disk. SQLite + FTS5 + a local ONNX embedding model. No API key, no cloud, no Docker, no graph database |
| **How it reaches the agent** | hooks on five events (SessionStart, UserPromptSubmit, PreToolUse, Stop, SessionEnd) — recall happens whether or not the agent thinks to ask, plus 10 MCP tools when it does |

Retrieval quality is measured, not asserted: **hit@5 0.871 / MRR 0.622** on the full LongMemEval
oracle set, hybrid retrieval, k=5, CPU only, reproducible from this repo — see
[Benchmarks](https://github.com/liza-studio/skillmem#benchmarks) for the per-type table and the reporting rules we hold ourselves to.

## Quickstart

macOS / Linux:

```
bash install.sh                 # installs python + uv if needed, venv, symlinks
```

Windows (PowerShell):

```
powershell -ExecutionPolicy Bypass -File install.ps1
```

Or from a checkout:

```
uv venv && uv pip install -e '.[semantic]'
source .venv/bin/activate       # or prefix the commands below with `uv run`
skillmem init --claude-code     # wires MCP server + hooks into Claude Code
skillmem init --codex           # wires the MCP server into the Codex CLI
skillmem init --all-agents      # ...or all six at once (see below)
skillmem doctor                 # health check: DB, schema, semantic status
```

Flags combine in one run — the agents then share one database.

### All six agents

| Flag | Agent | Config it writes |
| --- | --- | --- |
| `--claude-code` | Claude Code | `~/.claude.json` + hooks in `~/.claude/settings.json` |
| `--codex` | Codex CLI | `~/.codex/config.toml` |
| `--cursor` | Cursor | `~/.cursor/mcp.json` |
| `--windsurf` | Windsurf | `~/.codeium/windsurf/mcp_config.json` |
| `--gemini` | Gemini CLI | `~/.gemini/settings.json` |
| `--opencode` | opencode | `~/.config/opencode/opencode.json` |

Every entry is idempotent and backed up before it is touched; a config that
does not parse is left alone rather than overwritten. Each agent is stamped
with `SKILLMEM_AGENT`, so in a shared database "who learned this" stays
answerable. `skillmem uninstall` removes all of them (`--no-editors` to keep
the editor entries).

`init --claude-code` registers the MCP server in `~/.claude.json` and the hooks in `~/.claude/settings.json` (idempotent, with backups). Use `--hooks minimal` for no hooks at all (only the `skillmem trust` deny rule below), or `--hooks none` for MCP only. Hand-written memory files are imported with `skillmem migrate --source <dir>`; there is no per-turn import hook.

### Codex CLI

```
skillmem init --codex
```

Appends an `[mcp_servers.skillmem]` table to `~/.codex/config.toml` and marks the entry with
`SKILLMEM_AGENT=codex`. The tag is belt-and-braces: with no tag set, the server takes the
author's name from the agent's own MCP handshake, so attribution is right in a shared
database whichever way skillmem was installed.
The file is appended to, never rewritten: your own settings and comments stay where you put
them, the result is parsed before it is written, and invalid TOML is refused rather than
overwritten. `skillmem uninstall` removes the table again and leaves the rest of the file intact.

Codex reads `AGENTS.md` for project rules; if you keep yours in `CLAUDE.md`, point Codex at it
with `project_doc_fallback_filenames = ["CLAUDE.md"]` in the same config file — then both agents
follow one set of rules and one memory.

### As a plugin

The repo is also a plugin, in two flavours, both pointing at the same `skillmem-mcp` binary:

- **Agent Plugins** (`plugin.json` + `mcp.json` at the repo root) — what the Codex CLI installs from a
  marketplace. `mcp.json` needs both its `$schema` and `"type": "stdio"`, and the command must be a bare
  executable name rather than an absolute path — Codex's parser ignores the file otherwise, with no error.
  `codex mcp list` listing the server is the check that it parsed.
- **Claude Code** (`.claude-plugin/` + `hooks/hooks.json`) — MCP server *and* all six hooks in one install.

Either way the package itself must be on PATH (`pip install skillmem`); the plugin wires the server, not the runtime. An MCP Registry manifest (`server.json`) is in the repo as well:

```
/plugin marketplace add liza-studio/skillmem
/plugin install skillmem@liza-studio
```

The plugin requires the skillmem Python package on PATH and replaces `skillmem init --claude-code`'s wiring — use one or the other, not both (see [docs/PUBLISHING.md](https://github.com/liza-studio/skillmem/blob/main/docs/PUBLISHING.md)).

### Claude Desktop (chat app)

The MCP server also works in the Claude Desktop chat app — add to
`claude_desktop_config.json` (Settings → Developer → Edit Config):

```
{
  "mcpServers": {
    "skillmem": { "command": "skillmem-mcp" }
  }
}
```

You get all 10 `mem_*` tools on demand (search, learn, recall, reinforce…).
The automatic hooks (auto-recall on every prompt, session recap) are a
Claude Code mechanism and do not run in the chat app.

## How it works

```
 learn ──▶ recall ──▶ reinforce ──▶ decay
   │          │            │           │
   │          │            │           └─ daily job: unused skills lose strength;
   │          │            │              fully faded ones are archived (backed up)
   │          │            └─ strength +0.15 on outside evidence; ×0.7 after a failure
   │          └─ hybrid BM25 + vector search, strength-weighted ranking
   └─ after a hard task: trigger / steps / outcome / lessons
```

1. **learn** — after a task that took real debugging, the agent calls `mem_learn` with a slug, trigger, steps, outcome, and lessons.
2. **recall** — before the next task, `mem_recall` (or the automatic hooks) surfaces the most relevant skills, fusing lexical and semantic signals via Reciprocal Rank Fusion.
3. **reinforce** — when a recalled skill is confirmed by something outside the agent's own judgement (a test that passed, a diff that was accepted, the user saying so), `mem_reinforce` raises its strength, so proven skills rank higher next time. The agent calling its own skill useful is recorded but not rewarded; a task that failed after applying a skill lowers it. 
mrPetrukovich20
🟧 echo.github ⭐The repository documents SQLite-backed procedural memory, evidence-based reinforcement and decay, Claude Code and Codex integration, and verliza-studio——

Interpretation history

Decision trace