Az OpenAI azonosított és leállított egy koordinált kampányt, amelynek célja a modellek rejtett gondolkodási folyamatainak megszerzése volt. A támadók nem az adatbázisokat törték fel, hanem a modellekkel való interakciókat manipulálva próbálták megkerülni a korlátozásokat, hogy a kinyert adatokkal saját AI-rendszereiket fejlesszék.
A vizsgálat szerint a tevékenység egy jelentős része a Kimi nevű chatbotot fejlesztő Moonshot AI munkatársaihoz köthető. A támadók többek között azzal kísérleteztek, hogy az egyik beszélgetésből kimásolt titkosított gondolatmenetet egy másik beszélgetésben próbálták megfejtetni és leíratni a modellel.
Az OpenAI több mint 15 000 felhasználói fiókot tiltott le vagy korlátozott az ügy kapcsán. A vállalat megerősítette a biztonsági ellenőrzéseket, és megosztotta a tapasztalatokat az iparági partnerekkel, hogy megelőzzék a hasonló visszaéléseket.
Az eredeti szöveg (OpenAI)
We recently identified and disrupted a coordinated campaign designed to extract protected reasoning from our models, with the earliest observed activity occurring in the first week of July. This activity is consistent with adversarial distillation: the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model. Protected reasoning is the model’s internal record for working through a task; extracting it can reveal information withheld from the final answer and help others reproduce the model’s capabilities.
The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations. Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service. This manipulation is not a vulnerability unique to OpenAI’s models, and we have shared information about it with industry partners through the Frontier Model Forum in order to strengthen collective defenses against adversarial distillation.
Before publishing, we investigated the scope and potential impact, deployed our own mitigations, and shared with and took feedback from researchers and industry partners to ensure protections against this type of attack are in place. Additional mitigation and investigation work is continuing. We believe sharing what we have learned now will help the broader ecosystem strengthen its defenses.
We saw operators attempt to extract protected reasoning in novel ways, including by copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe the hidden reasoning content.
Independent security researchers(opens in a new window) also brought related cross-model and conversation-compaction vulnerabilities to our attention through responsible disclosure. We investigated their findings and confirmed that the attack paths they identified were real. Their work helped us understand the broader attack class and accelerate mitigations.
The activity began on July 1, initially at a low volume until we observed high-volume spikes on July 24 and 25 consisting of 16,000 requests1 using a relevant extraction pattern from over 4,000 users. Further investigation identified related prompt-pattern activity across a cluster of more than 15,000 users, which we fully disrupted by July 28.
The activity evolved over time, reinforcing that adversarial distillation is a broader security challenge that requires layered, adaptive defenses.
It is unclear whether all operators we observed during the relevant time period originated from a single actor. However, we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.
Adversarial distillation poses safety and national security risks. Extracted reasoning could be used to train another model without preserving the safeguards applied to the original model’s user-facing outputs. At scale, distillation can also accelerate the transfer of advanced capabilities without requiring the same investment in safety. These concerns become heightened as models gain capabilities in dual use domains.
This risk is not unique to OpenAI. As cited above, similar techniques may affect other advanced AI systems, making this a shared security challenge that requires coordination across the industry.
We mitigated this recent distillation campaign through a combination of account enforcement, technical controls, and partner coordination. We banned or restricted fraudulent accounts, strengthened signup and infrastructure controls, and expanded monitoring for related networks.
We also strengthened protections for hidden reasoning across users, workspaces, organizations, and model families. We closed a pathway that allowed someone who already possessed another user's encrypted reasoning to replay it and recover its conte