Stealing Reasoning Traces from Proprietary LLM APIs
We find that encrypted chain-of-thought blocks are interchangeable across sessions with most LLM providers. This allows attackers to replay a frontier model trace into a weaker, jailbroken sibling which recovers the hidden reasoning verbatim, enabling distillation, large-scale extraction of private data such as secrets and PII from agent logs, safety violations, and invisible prompt injection.
Alexander Panfilov 1,2,3,4 * David Schmotz 2,3,4 * Ilia Shumailov 5 * Luca Beurer-Kellner 6 Joachim Schaeffer 1 Ameya Prabhu 2,4,7 ‡ Jonas Geiping 2,3,4 ‡ Maksym Andriushchenko 2,3,4 ‡ 1 MATS Research 2 ELLIS Institute Tübingen 3 Max Planck Institute for Intelligent Systems 4 Tübingen AI Center 5 AI Sequrity Company 6 Snyk 7 University of Tübingen *Equal contribution, order decided by dice roll · ‡Equal supervision Aug 10, 2026 Read the Paper Website…
saved by
related reading
- Let’s talk about encrypted reasoningblog.cryptographyengineering.com
- Stolen Thoughtsstolen-thoughts.com
- [2608.09867] Stealing Reasoning Traces from Proprietary LLM APIsarxiv.org
- [2510.24941] Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thoughtarxiv.org
- Stealing Reasoning Traces from Proprietary LLM APIsarxiv.org
- Quantifying the Necessity of Chain of Thought through Opaque Serial Deptharxiv.org
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com
- [2603.07267] How to Steal Reasoning Without Reasoning Tracesarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- As Rocks May Think | Eric Jangevjang.com
- Learning to reason with LLMs | OpenAIopenai.com
- Prompt Injection as Role Confusionrole-confusion.github.io