
Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov
Machine Learning Street Talk (MLST)
Frontier LLMs from providers like Anthropic, OpenAI, and Google encrypt internal reasoning traces, yet these "thought" blobs remain vulnerable to extraction and decoding. This architectural flaw allows attackers to replay reasoning traces across different user sessions and models, enabling privacy breaches, prompt injections, and unauthorized model distillation. By injecting these traces into smaller models, researchers observed significant shifts in output style and length, suggesting that reasoning is not as isolated as intended. While labs have been notified, the vulnerability highlights a critical trade-off between model monitorability and security. Addressing this requires architectural revisions, such as restricting trace access or preventing the replay of reasoning blobs, to mitigate the risks posed by these increasingly agentic and inscrutable AI systems.
Sign in to continue reading, translating and more.
Open full episode in Podwise