flâneur

Stealing Reasoning Traces from Proprietary LLM APIs

research.snyk.io · 2,142 words · saved by 1 readers

We find that encrypted chain-of-thought blocks are interchangeable across sessions with most LLM providers. This allows attackers to replay a frontier model trace into a weaker, jailbroken sibling which recovers the hidden reasoning verbatim, enabling distillation, large-scale extraction of private data such as secrets and PII from agent logs, safety violations, and invisible prompt injection.

Alexander Panfilov 1,2,3,4 * David Schmotz 2,3,4 * Ilia Shumailov 5 * Luca Beurer-Kellner 6 Joachim Schaeffer 1 Ameya Prabhu 2,4,7 ‡ Jonas Geiping 2,3,4 ‡ Maksym Andriushchenko 2,3,4 ‡ 1 MATS Research 2 ELLIS Institute Tübingen 3 Max Planck Institute for Intelligent Systems 4 Tübingen AI Center 5 AI Sequrity Company 6 Snyk 7 University of Tübingen *Equal contribution, order decided by dice roll · ‡Equal supervision Aug 10, 2026 Read the Paper Website…

saved by

related reading