STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations

Published
Source
arXiv
Paper number
305
Field
Machine Learning
arXiv ID
2606.05165

Key points

  • The paper proposes a shift in perspective: instead of estimating parameter changes, it models the functional effect of training data in activation space.
  • The authors introduce STRIDE, which stands for Steering-based Training Data Influence Decomposition, a framework that formulates TDA as a sparse recovery problem in the spirit of compressed sensing.
  • STRIDE achieves SOTA on LLM pretraining attribution while being one order of magnitude, specifically 13x, faster than existing methods.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)