2026-10-11 18:50 UTC

The Contrastive Decoding Diffing researchers claim access to base and fine-tuned model logits is sufficient to recover verbatim narrow fine-tuning data without weights or activations, creating a privacy risk for APIs that expose token logits.

state: expiredheat: lowuncertainty: highconvergesscott: mediummodel-extraction training-data-privacy llm-security

What is this?

A paper titled “Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing” claims that comparing outputs from base and fine-tuned models can recover verbatim fine-tuning content using logits, without access to model weights. If validated, this makes token-logit exposure by model APIs a potential channel for leaking sensitive or proprietary fine-tuning data; the supplied results also point to differentially private synthetic training data as a possible mitigation. The snippets do not identify the researchers or provide enough experimental detail to establish the attack’s success rate, access assumptions, cost, or practical scope.

Why it matters to Scott

The claimed extraction channel reinforces Scott’s position that sensitive material should be transformed before becoming model-visible, because fine-tuning raw private data may make the model itself an information-egress surface despite ordinary API controls. If validated, it would also bear directly on his fine-tuning-data factory and privacy-tokenized appliance designs, but the supplied evidence does not establish practical attack success or scope.
ip:framework.separation-of-powers-for-cognitionip:concept.proxy-mediated-tokenisationdev:concept.privacy-tokenized-agent-boundarydev:project.redditdev:project.applianceradar:concept.model-securityradar:concept.api-attacksradar:concept.ai-privacyradar:concept.fine-tuningradar:previous-token-prompt-reconstructionradar:proprietary-api-reasoning-trace-extraction
queries asked of Scott's wikis
  • logit exposure and model API security
  • fine-tuning data memorization and extraction
  • training-data privacy threat models
  • base versus fine-tuned model differencing
  • differential privacy for private fine-tuning
  • LLM API information leakage controls

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Contrastive Decoding Diffing (CDD): recovering verbatim finetuning data from logits alone, no weight access needed[R]
MachineLearning
CebulkaZapiekana4511

Interpretation history

Decision trace