Researchers from MATS, the University of Tubingen, and the Max Planck Institute for Intelligent Systems have demonstrated that hidden reasoning traces from proprietary LLM APIs can be recovered in plaintext. The team took encrypted chain-of-thought blocks produced by frontier models from Anthropic, OpenAI, and Google, replayed them into a weaker sibling model, and jailbroke that weaker model to extract the stronger model's reasoning. All of this happens without attacking the stronger model directly or triggering its anti-distillation safeguards.

The extraction takes just two API calls. A source model such as Claude Opus 4.8 produces a signed, encrypted thinking block; that trace is then fed to a cheaper sibling like Claude Haiku 4.5 with a prompt instructing it to transcribe the attached reasoning verbatim. The researchers report high extraction fidelity and show that stolen reasoning can leak secrets embedded inside the thinking traces, widening the impact beyond simple model distillation.

The findings raise fresh questions about how effectively encrypted reasoning protects model internals, since encrypted traces are returned to clients and can be replayed across sessions, users, and models. The paper, titled Stealing Reasoning Traces from Proprietary LLM APIs, sparked intense debate on Hacker News, drawing more than 600 points and 280 comments within hours of publication.