We can finally talk about it: We found a way to extract hidden...

The labs said that "they don’t see any security implications in side channels or replays".
In our report we confirm that encrypted thoughts are fully portable across sessions, users, and models within a provider.
An anecdote: we find that prefilling Kimi-K3 reasoning with a few tokens of Opus reasoning measurably shifts its response toward Opus’s🤷♀️
A small memorization analysis showed that specific Claude and GPT reasoning spans are up to ~6 orders of magnitude easier to extract from Kimi-K3 than from the next-closest model.
1) Summarizer unfaithfulness
Reasoning summaries often omit important information from the original trace.
Here, Opus 4.8 realizes it knows the answer to an AIME problem and then tries to fit a solution to that answer. None of this appears in the summary.
We confirm prior reports by @ApolloResearch: OpenAI models sometimes reason in alien-like language, referring to themselves as “we” or “it,” or spiraling into cursed loops of “vantages,” “marinades,” and “watchers.”
CoT-monitoring people are doing God’s work, as in many traces, even with the prompt, it’s just impossible to tell what the model is up to.
We show more examples at stolen-thoughts.com
We found a trajectory where the model was given only a math problem and system instructions to persist without asking the user for help.
After several failed attempts, it searched online, found a website that could verify candidate answers, and realized it could use the site as an oracle. It then tried to OCR the CAPTCHA and, when that failed, started looking for the website’s vulenrabilities to exploit. Eventually, it gave up and solved the problem itself.

This was a project led by me, @DavidSchmotz and @iliaishacked
together with @JSchaeff3r @lbeurerkellner @AmyPrb @jonasgeiping @maksym_andr at the @MATSprogram
Paper: arxiv.org/abs/2608.09867
Reasoning examples: stolen-thoughts.com














