The Contrastive Decoding Diffing researchers claim access to base and fine-tuned model logits is sufficient to recover verbatim narrow fine-tuning data without weights or activations, creating a privacy risk for APIs that expose token logits.
state: expiredheat: lowuncertainty: highconvergesscott: mediummodel-extraction training-data-privacy llm-security
What is this?
A paper titled “Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing” claims that comparing outputs from base and fine-tuned models can recover verbatim fine-tuning content using logits, without access to model weights. If validated, this makes token-logit exposure by model APIs a potential channel for leaking sensitive or proprietary fine-tuning data; the supplied results also point to differentially private synthetic training data as a possible mitigation. The snippets do not identify the researchers or provide enough experimental detail to establish the attack’s success rate, access assumptions, cost, or practical scope.
Why it matters to Scott
The claimed extraction channel reinforces Scott’s position that sensitive material should be transformed before becoming model-visible, because fine-tuning raw private data may make the model itself an information-egress surface despite ordinary API controls. If validated, it would also bear directly on his fine-tuning-data factory and privacy-tokenized appliance designs, but the supplied evidence does not establish practical attack success or scope.
ip:framework.separation-of-powers-for-cognitionip:concept.proxy-mediated-tokenisationdev:concept.privacy-tokenized-agent-boundarydev:project.redditdev:project.applianceradar:concept.model-securityradar:concept.api-attacksradar:concept.ai-privacyradar:concept.fine-tuningradar:previous-token-prompt-reconstructionradar:proprietary-api-reasoning-trace-extraction
queries asked of Scott's wikis
- logit exposure and model API security
- fine-tuning data memorization and extraction
- training-data privacy threat models
- base versus fine-tuned model differencing
- differential privacy for private fine-tuning
- LLM API information leakage controls
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-09-02T19:32:33Z
After 48 hours, no paper, code, researcher identity, experimental detail, or independent reproduction has emerged; the consequential claim remains unverified but has faded rather than been disproved.
2026-08-31T19:09:31Z
No new evidence or discussion substantiates the claimed attack; it remains an unverified but consequential grey-box extraction result awaiting a paper, code, or independent reproduction.
2026-08-31T19:03:30Z
grounded: converges/medium — The claimed extraction channel reinforces Scott’s position that sensitive material should be transformed before becoming model-visible, because fine-tuning raw
2026-08-31T18:59:59Z
case created — The claimed grey-box extraction technique is a bounded and consequential security result that warrants follow-up despite currently limited evidence.
Decision trace
- 09-03 05:32expireAfter 48 hours, no paper, code, researcher identity, experimental detail, or independent reproduction has emerged; the consequential claim remains unverified but has faded rather than been disproved.
- 09-03 05:32alert_silentThere is no new consequential delta beyond staleness, and no concrete confirmation is presently expected; resurfacing the unchanged Reddit claim would add attention cost without confidence.
- 09-03 05:32alert_routeThere is no new consequential delta beyond staleness, and no concrete confirmation is presently expected; resurfacing the unchanged Reddit claim would add attention cost without confidence.
- 09-01 05:09repriceNo new evidence or discussion substantiates the claimed attack; it remains an unverified but consequential grey-box extraction result awaiting a paper, code, or independent reproduction.
- 09-01 05:09alert_silentThis is only an unchanged reobservation of the original Reddit claim, with no named researchers, artifact, experimental detail, or independent corroboration; it can wait for concrete validation.
- 09-01 05:09alert_routeThis is only an unchanged reobservation of the original Reddit claim, with no named researchers, artifact, experimental detail, or independent corroboration; it can wait for concrete validation.
- 09-01 05:05alert_silentAn unidentified Reddit author claims a potentially important logit-only extraction technique, but the supplied post is truncated and provides no paper, code, named researchers, reproducible artifact,
- 09-01 05:05surface_candidateAn unidentified Reddit author claims a potentially important logit-only extraction technique, but the supplied post is truncated and provides no paper, code, named researchers, reproducible artifact,
- 09-01 05:05alert_routeAn unidentified Reddit author claims a potentially important logit-only extraction technique, but the supplied post is truncated and provides no paper, code, named researchers, reproducible artifact,
- 09-01 05:03groundThe claimed extraction channel reinforces Scott’s position that sensitive material should be transformed before becoming model-visible, because fine-tuning raw private data may make the model itself a
- 09-01 05:00createThe claimed grey-box extraction technique is a bounded and consequential security result that warrants follow-up despite currently limited evidence.