Minimally Invasive Steering of Language Models
- Published
- Source
- arXiv
- Paper number
- 1131
- Field
- Machine Learning
- arXiv ID
- 2609.30218
Key points
- It adjusts model dispositions using only vectors added to the final hidden states without retraining, while suppressing distribution shift with a Fisher quadratic penalty.
- It exactly decomposes the sequence-level KL gradient into an analytic term and a score-function term, and proves the validity of the approximation to first order.
- The computation reduces to matrix-vector products with the frozen model head, so position-specific steering vectors can be optimized without any training.
- It achieved the highest mean reward in six of seven settings on preference and code-generation tasks across 1B-14B models, while keeping diversity and coherence at Best-of-N levels.
Paper links
External research summaries. These are not HDATF publications or measured product results.