STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations
- Published
- Source
- arXiv
- Paper number
- 305
- Field
- Machine Learning
- arXiv ID
- 2606.05165
Key points
- The paper proposes a shift in perspective: instead of estimating parameter changes, it models the functional effect of training data in activation space.
- The authors introduce STRIDE, which stands for Steering-based Training Data Influence Decomposition, a framework that formulates TDA as a sparse recovery problem in the spirit of compressed sensing.
- STRIDE achieves SOTA on LLM pretraining attribution while being one order of magnitude, specifically 13x, faster than existing methods.
Paper links
External research summaries. These are not HDATF publications or measured product results.