Stealing Reasoning Traces from Proprietary LLM APIs
- Published
- Source
- arXiv
- Paper number
- 858
- Field
- AI / General
- arXiv ID
- 2608.09867
Key points
- It discovers a fundamental design flaw in encrypted reasoning blocks that makes them interchangeable across sessions, users, and models.
- It demonstrates a cross-model attack that turns the reasoning of a strong model into something that can be decoded by a weaker model against OpenAI, Anthropic, and Google.
- It recovers real sensitive information from public developer logs, including 62 API keys, 33 passwords, and 30 emails.
- It also confirms an invisible prompt-injection attack vector that passes through encrypted blocks.
- It highlights the fundamental limits of encryption-based approaches and proposes concrete mitigation measures for both vendors and users.
Paper links
External research summaries. These are not HDATF publications or measured product results.