Tell me about yourself: LLMs are aware of their learned behaviors

Published
Source
arXiv
Paper number
023
Field
Alignment / Interpretability
arXiv ID
2501.11120

Key points

  • Understanding the internal states and emergent, implicit behaviors of large language models (LLMs) remains difficult, especially when such behaviors are not explicitly described in the training data.
  • Detecting hidden backdoor behaviors in LLMs that activate only under specific and often obscure trigger conditions is a major AI safety challenge.
  • Ensuring LLM transparency and alignment requires methods that proactively identify and understand undesirable tendencies, including those arising from training-data bias or malicious contamination.
  • This study fine-tuned LLMs, GPT-4o and Llama-3.1-70B, on diverse datasets that implicitly exhibit specific behaviors such as economic decision making, strategic goal pursuit in games, and vulnerable code generation.
  • To evaluate the models' ability to explain the policies they learned, we probed them with out-of-distribution questions in free-form, numeric, multiple-choice, and two-step reasoning formats.
  • For backdoor detection, this study introduced reverse learning, which explicitly learns the reverse mapping needed to elicit a trigger by augmenting the training data with swapped user and assistant messages.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)