Minimally Invasive Steering of Language Models

Published
Source
arXiv
Paper number
1131
Field
Machine Learning
arXiv ID
2609.30218

Key points

  • It adjusts model dispositions using only vectors added to the final hidden states without retraining, while suppressing distribution shift with a Fisher quadratic penalty.
  • It exactly decomposes the sequence-level KL gradient into an analytic term and a score-function term, and proves the validity of the approximation to first order.
  • The computation reduces to matrix-vector products with the frozen model head, so position-specific steering vectors can be optimized without any training.
  • It achieved the highest mean reward in six of seven settings on preference and code-generation tasks across 1B-14B models, while keeping diversity and coherence at Best-of-N levels.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)