2026-10-11 17:11 UTC

Independent replication will determine whether the published attack can recover meaningful hidden reasoning traces from proprietary LLM APIs that do not expose chain-of-thought.

state: expiredheat: lowuncertainty: highknownscott: highllm-security reasoning-traces api-attacks

What is this?

A research paper reports an architectural vulnerability in proprietary LLM APIs that return hidden reasoning traces to clients as encrypted blocks: those blocks could allegedly be replayed across sessions, users, and sometimes sibling models. The researchers say a weaker model could be jailbroken and used as a decryption oracle to recover a stronger model’s hidden reasoning, creating risks including model distillation and behavioral analysis. The supplied snippets describe the authors’ findings and secondary discussion, but do not establish an independent replication; one secondary source says providers implemented mitigations and the attacks were no longer reproducible by August 2026.

Why it matters to Scott

The radar already tracks this exact replication question in `radar:proprietary-api-reasoning-trace-extraction`. It directly bears on Scott’s verification-boundary analysis and provider-bound reasoning-continuity design: successful cross-session, cross-account, or cross-model replay would expose limits in the integrity and isolation assumptions around encrypted reasoning state and could require stricter replay scoping or stripping in agent harnesses.
ip:concept.verification-boundaryip:source.the-unverified-conversation-why-llms-can-t-trust-their-own-history-ebookip:concept.reasoning-paradoxdev:concept.provider-bound-reasoning-continuityradar:proprietary-api-reasoning-trace-extractionradar:concept.llm-securityradar:concept.llm-apisradar:concept.model-distillation
queries asked of Scott's wikis
  • client-side state trust boundaries in LLM APIs
  • hidden reasoning traces and chain-of-thought security
  • agent logging and caching of encrypted reasoning blocks
  • cross-model replay and decryption-oracle attacks
  • reasoning-trace leakage as a model-distillation vector
  • security assumptions for proprietary model APIs

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnStealing Reasoning Traces from Proprietary LLM APIsamai10
🟧 echo.blog ⭐Green’s original post reports his own experiments with encrypted reasoning blobs: replaying unmodified blocks across sessions, accounts, andMatthew Green——
🟧 hnStealing Reasoning Traces from Proprietary LLM APIsdoppp20
🟠 redditTrace-Inverter-4B: A Qwen 4B FT on Jackrongs datasets of synthetic traces to support {prompt, final answer} -> synthetic reasoning
LocalLLaMA
Signature9763

Interpretation history

Decision trace