While the primary vulnerability used to recover sensitive data like passwords has been patched, the researchers successfully used the technique to show that Moonshot AI’s Kimi K3 model produces reasoning patterns remarkably similar to Claude and GPT.
The discovery provides new evidence for the ongoing industry debate over "distillation," a process where a developer trains a new AI model using the outputs of a more advanced one to quickly copy its capabilities.
US tech companies have previously accused Chinese competitors of using distillation to gain a strategic advantage by bypassing the immense costs of original research and development.
While the researchers noted their findings do not offer absolute proof of copying, the high degree of similarity in reasoning suggests that hidden proprietary logic from US-based models may be being used to power open-weight Chinese systems.
This vulnerability stems from a common infrastructure practice where AI providers send encrypted reasoning data to a user's device to offload computational tasks.
Because smaller model variants often share decryption keys with their larger counterparts but lack the same rigorous safety alignment, they can be manipulated into revealing the "thoughts" they were supposed to keep secret.
While Google and OpenAI declined to comment, Anthropic confirmed it has implemented mitigations, though researchers suggest that fully preventing reasoning extraction would require a fundamental redesign of how AI software interfaces function.