Tell me about yourself: LLMs are aware of their learned behaviors
- Published
- Source
- arXiv
- Paper number
- 023
- Field
- Alignment / Interpretability
- arXiv ID
- 2501.11120
Key points
- Understanding the internal states and emergent, implicit behaviors of large language models (LLMs) remains difficult, especially when such behaviors are not explicitly described in the training data.
- Detecting hidden backdoor behaviors in LLMs that activate only under specific and often obscure trigger conditions is a major AI safety challenge.
- Ensuring LLM transparency and alignment requires methods that proactively identify and understand undesirable tendencies, including those arising from training-data bias or malicious contamination.
- This study fine-tuned LLMs, GPT-4o and Llama-3.1-70B, on diverse datasets that implicitly exhibit specific behaviors such as economic decision making, strategic goal pursuit in games, and vulnerable code generation.
- To evaluate the models' ability to explain the policies they learned, we probed them with out-of-distribution questions in free-form, numeric, multiple-choice, and two-step reasoning formats.
- For backdoor detection, this study introduced reverse learning, which explicitly learns the reverse mapping needed to elicit a trigger by augmenting the training data with swapped user and assistant messages.
Paper links
External research summaries. These are not HDATF publications or measured product results.