Stealing Reasoning Traces from Proprietary LLM APIs

Published
Source
arXiv
Paper number
858
Field
AI / General
arXiv ID
2608.09867

Key points

  • It discovers a fundamental design flaw in encrypted reasoning blocks that makes them interchangeable across sessions, users, and models.
  • It demonstrates a cross-model attack that turns the reasoning of a strong model into something that can be decoded by a weaker model against OpenAI, Anthropic, and Google.
  • It recovers real sensitive information from public developer logs, including 62 API keys, 33 passwords, and 30 emails.
  • It also confirms an invisible prompt-injection attack vector that passes through encrypted blocks.
  • It highlights the fundamental limits of encryption-based approaches and proposes concrete mitigation measures for both vendors and users.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)