2026-10-11 16:38 UTC

OpenAI says it disrupted a coordinated adversarial-distillation campaign — manipulation of model interactions, including replaying encrypted reasoning across conversations for decryption, peaking at 16,000 requests from 4,000+ users on July 24-25 within a 15,000+ user cluster fully disrupted by July 28 — attributing a core cluster to individuals associated with Moonshot AI; follow-on disclosures by other labs, a Moonshot response, or Frontier Model Forum-coordinated mitigations would establish bulk-query capability extraction as a recognized cross-lab security threat class.

state: watchingheat: mediumuncertainty: mediumconvergesscott: highmodel-distillation ai-security api-abuseOpenAIMoonshot AI
Surfaced 2026-10-03T15:55:11Z — OpenAI: 'We recently identified and disrupted a coordinated campaign designed to extract protected reasoning from our models' — 'not a vulne — The episode's news cycle has closed — ~97h old, zero current engagement, and none of the hypothesized follow-ons (Moonshot response, other-lab disclosures, coordinated FMF mitigations) surfaced — so its meaning shifts from 'breaking disclosure' to 'settled first-party receipt now watched for cross-lab recognition.' The disclosure text itself already satisfies the information-sharing prong (FMF sharing claimed; independent researchers' related attack paths confirmed), but the Moonshot attribution still rests solely on OpenAI.

What is this?

OpenAI — the US frontier AI lab behind ChatGPT, per the generic corporate coverage in the supplied snippets — has, per the case's captured first-party statement, identified and disrupted a coordinated campaign to extract 'protected reasoning' from its models: manipulation of model interactions including replaying encrypted reasoning across conversations for decryption, peaking at ~16,000 requests from 4,000+ users on July 24–25 within a 15,000+ user cluster OpenAI says was fully disrupted by July 28, with a core cluster attributed to individuals associated with Moonshot AI. The disclosure's framing is that this is abuse rather than a security flaw (title fragment: 'not a vulne…'). The supplied web results contain no independent coverage of this disclosure, no Moonshot response, and none of the hypothesized follow-ons (other-lab disclosures, Frontier Model Forum-coordinated mitigations), so the campaign details rest entirely on the captured first-party evidence object. A second captured item ('The AI industry has discovered intellectual property') appears to be commentary on the episode.

Why it matters to Scott

OpenAI's first-party account of attackers replaying encrypted reasoning across conversations for decryption lands exactly on the verification boundary Scott mapped in 'The Unverified Conversation' ebook and his provider-bound reasoning-continuity dev concept — first-party, dated receipts that the reasoning-state/conversation-history seam is under industrial-scale (15k-user) attack, directly usable for LeverageAI's security positioning and the ebook's receipts file. The Moonshot-linked attribution also sharpens his capability-symmetry argument: rented frontier reasoning is extraction bait, reinforcing that the durable moat is the private corpus, not the model.
ip:source.the-unverified-conversation-why-llms-can-t-trust-their-own-history-ebookip:concept.verification-boundarydev:concept.provider-bound-reasoning-continuityip:concept.reasoning-paradoxip:concept.capability-symmetrywork:project.leverageairadar:openai-kimi-distillation-disruptionradar:concept.reasoning-tracesradar:concept.model-distillationradar:concept.agent-securityradar:proprietary-api-reasoning-trace-extractionradar:moonshot-claude-routing-allegationradar:cross-lab-frontier-incident-investigation
queries asked of Scott's wikis
  • reasoning trace IP chain-of-thought confidentiality
  • distillation capability transfer open-weights catch-up
  • model moats frontier lab defensibility
  • API abuse detection rate limiting agent harness
  • cross-lab AI security coordination threat intelligence sharing
  • hidden chain-of-thought exposure encrypted reasoning

Measured heat

now 0 pts/hpeak 14 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 290h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-29 14:00⭐ origin echo-reconstructedOpenAI: 'We recently identified and disrupted a coordinated campaign designed to extract protected reasoning from our models' — 'not a vulne
OpenAI on blog (echo) · attributed from reddit.post.1wv0l7i
—
10-01 14:15first on r/artificial · published · +48.3hThe AI industry has discovered intellectual property
Haunting_Ganache_850
—
10-04 01:40first on hacker news · published · +107.7hDisrupting a coordinated model-distillation campaign
gmays
—
10-01 14:15amplified on r/artificial 👑reddit.post.1wv0l7i
Haunting_Ganache_850
peak 190 · 60 comments · 98% of case engagement
10-04 01:40amplified on hacker newshn.story.49949746
gmays
peak 3 · 0 comments · 2% of case engagement
09-30 19:40our radar first saw it · +29.7hdiscovery anchor: reddit.post.1wv0l7i—
10-03 15:42reached heat=high · +97.7h · via ledger——
pace: p76 vs 1188 stories at the 168h mark (now 290h old) — ahead of cactus-whistle-edge-asr (1.0x), behind ship-harness-bench (1.0x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditThe AI industry has discovered intellectual property
artificial
Retrieved article excerpt

Open article · Retrieved 2026-10-01T14:32:30.377984+00:00

September 30, 2026

[Security](https://openai.com/news/security/)

# Disrupting a coordinated model-distillation campaign

Loading…

Share

We recently identified and disrupted a coordinated campaign designed to extract protected reasoning from our models, with the earliest observed activity occurring in the first week of July. This activity is consistent with adversarial distillation: the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model. Protected reasoning is the model’s internal record for working through a task; extracting it can reveal information withheld from the final answer and help others reproduce the model’s capabilities.

The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations. Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service. This manipulation is not a vulnerability unique to OpenAI’s models, and we have shared information about it with industry partners through the Frontier Model Forum in order to strengthen collective defenses against adversarial distillation.

Before publishing, we investigated the scope and potential impact, deployed our own mitigations, and shared with and took feedback from researchers and industry partners to ensure protections against this type of attack are in place. Additional mitigation and investigation work is continuing. We believe sharing what we have learned now will help the broader ecosystem strengthen its defenses.

## What we observed

We saw operators attempt to extract protected reasoning in novel ways, including by copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe the hidden reasoning content.

[Independent security researchers⁠(opens in a new window)](https://arxiv.org/abs/2608.09867) also brought related cross-model and conversation-compaction vulnerabilities to our attention through responsible disclosure. We investigated their findings and confirmed that the attack paths they identified were real. Their work helped us understand the broader attack class and accelerate mitigations.

The activity began on July 1, initially at a low volume until we observed high-volume spikes on July 24 and 25 consisting of 16,000 requests[1](https://openai.com/index/disrupting-a-coordinated-model-distillation-campaign/#citation-bottom-1) using a relevant extraction pattern from over 4,000 users. Further investigation identified related prompt-pattern activity across a cluster of more than 15,000 users, which we fully disrupted by July 28.

The activity evolved over time, reinforcing that adversarial distillation is a broader security challenge that requires layered, adaptive defenses.

## Our assessment of attribution

It is unclear whether all operators we observed during the relevant time period originated from a single actor. However, we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.

## Why this matters

Adversarial distillation poses safety and national security risks. Extracted reasoning could be used to train another model without preserving the safeguards applied to the original model’s user-facing outputs. At scale, distillation can also accelerate the transfer of advanced capabilities without requiring the same investment in safety. These concerns become heightened as models gain capabilities in dual use domains.

This risk is not unique to OpenAI. As cited above, similar techniques may affect other advanced AI systems, making this a shared security challenge that requires coordination across the industry.

## How we responded

We mitigated this recent distillation campaign through a combination of account enforcement, technical controls, and partner coordination. We banned or restricted fraudulent accounts, strengthened signup and infrastructure controls, and expanded monitoring for related networks.

We also strengthened protections for hidden reasoning across users, workspaces, organizations, and model families. We closed a pathway that allowed someone who already possessed another user's encrypted reasoning to replay it and recover its contents, and added checks to detect and hold streamed output that might expose reasoning. When related activity moved through third-party services, we worked with those providers to identify and disrupt the accounts involved.

Finally, we shared relevant findings through the Frontier Model Forum and appropriate government information-sharing channels so that other frontier developers and public-sector partners could look for similar activity and strengthen their own defenses. Systems that support portable or replayable reasoning artifacts may face related risks.

## What comes next

We expect adversarial distillation attempts to become more sophisticated as frontier models improve and as actors look for cheaper ways to mimic their capabilities. Defending against this activity requires layered controls and continual adaptation.

This work is not finished. Partner-hosted deployments need the same protections as first-party services, and tool-output attacks require protections that examine more than ordinary visible text. We are continuing to improve tool defenses, classifier coverage, model refusals, and propagate relevant controls across cloud partners.

Our response will continue to focus on three areas: stronger technical protections against extraction, better detection and enforcement against coordinated campaigns, and deeper threat-information sharing across industry and government.

- [2026](https://openai.com/news/?tags=2026)

## Author

OpenAI

## Footnotes

1. 1

   These figures describe attempted, not necessarily successful, extractions.

## Keep reading

[View all](https://openai.com/news/)

Daybreak for Critical Infrastructure — cover

[Daybreak for Frontline Defenders

SecuritySep 3, 2026](https://openai.com/index/daybreak-for-frontline-defenders/)

Path to Astra — Clean square cover — Neutral Option 062 v1

[Path to Astra: critical capabilities and frontier safeguards

SafetySep 1, 2026](https://openai.com/index/path-to-astra/)

[The Hugging Face incident and the road ahead

SecurityAug 26, 2026](https://openai.com/index/hugging-face-incident-and-the-road-ahead/)
Haunting_Ganache_85019060
🟧 echo.blog ⭐OpenAI: 'We recently identified and disrupted a coordinated campaign designed to extract protected reasoning from our models' — 'not a vulneOpenAI——
🟧 hnDisrupting a coordinated model-distillation campaigngmays30

Interpretation history

Decision trace