Stealing Reasoning Traces from Proprietary LLM APIs
Summary
Stealing Reasoning Traces from Proprietary LLM APIs reports that proprietary reasoning can be recovered from encrypted traces and replayed into weaker models to reveal plaintext reasoning. The paper demonstrates this across frontier models (Anthropic, OpenAI, Google), highlights privacy leaks including API keys and tokens, and discusses implications for model providers and users, along with jailbreak techniques and mitigation challenges.