— reading now
Flash

Story·Safety & Risk·2026-10-02 21:30

OpenAI Disrupts "Distillation" Extraction Campaign: 16,000 Requests Point at Moonshot-Linked Actors

On September 30, OpenAI disclosed that it had disrupted a coordinated campaign targeting its models' hidden reasoning: on July 24–25, more than 4,000 accounts fired some 16,000 extraction-pattern requests, with related activity touching 15,000+ users. OpenAI attributed the core cluster to "individuals associated with Moonshot AI" — while conceding it cannot confirm all operators came from a single actor.

What Happened

Timeline: the campaign began July 1 at low volume, then peaked on July 24–25 — more than 4,000 users sending some 16,000 extraction-pattern requests in two days. OpenAI traced related prompt-pattern activity spanning 15,000+ users. By July 28 the whole operation was fully disrupted. On September 30, OpenAI disclosed it publicly.

The technique: operators copied protected encrypted reasoning out of one conversation and, in a separate conversation, got the model to decrypt and transcribe the "unreadable ciphertext" — manipulating model interactions so hidden reasoning would reproduce in a form visible to the requester. OpenAI stressed three "didn'ts": the encryption was not broken, no database was compromised, and no stored user conversations were directly accessed.

OpenAI has formally named the behavior "adversarial distillation": the systematic, unauthorized use of one model's outputs or reasoning to train, reproduce, or improve another model. The cleanup included banning or restricting the accounts involved, tightening new-account signup verification, closing the replay pathway, and adding detection for streamed output that might expose reasoning. Findings were shared with peers and officials via the Frontier Model Forum and government channels.

Key Facts

  1. Disclosure & window: Disclosed by OpenAI on Sept 30; the extraction activity ran July 1–28.
  2. Scale: July 24–25 peak: 4,000+ users, ~16,000 requests; related prompt-pattern activity touched 15,000+ users (per OpenAI; the counts measure "attempts," not confirmed successes).
  3. Attribution: Core cluster attributed to "individuals associated with Moonshot AI" (the Beijing company behind Kimi). OpenAI also said: "It is unclear whether all operators we observed during the relevant time period originated from a single actor."
  4. Technique: Carrying encrypted reasoning across conversations and inducing the model to decrypt it; three "didn'ts" — encryption unbroken, database uncompromised, stored conversations untouched.
  5. Official line: Caroline Zier, who leads OpenAI's strategic national security policy work, told Bloomberg the concern is "violation of our terms of service, not open models or legitimate distillation"; OpenAI says "adversarial distillation poses safety and national security risks."
  6. Earlier allegation: Per CellCog, this is the second distillation allegation against Moonshot this month — Anthropic raised similar claims in a Sept 10 threat report; an August paper by independent researchers showed encrypted reasoning blocks are interchangeable across sessions, users, and models.

Context

OpenAI encrypts reasoning, and Anthropic tightened external testing. An August independent paper showed encrypted reasoning blocks are interchangeable across sessions, users, and even models.

Hudson Institute's Duesterberg says it shows Chinese firms "can't compete at all unless they steal our stuff." Other outlets note OpenAI published "attempt" counts, not success counts, and OpenAI says it cannot confirm all operators came from a single actor.

Why it matters

OpenAI formally named the behavior "adversarial distillation" and says it poses national-security risks; it is the second distillation allegation against Moonshot this month, after one from Anthropic.
Useful Tap if this story helped you

SourcesOpenAI (disclosure, 9/30/2026), The Hacker News, CellCog, RuntimeWire, aiweekly, Daily Caller. Compiled from public reporting; not investment advice.

Comments

  1. Loading comments…
Ask the cat