OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
- Published
- Source
- arXiv
- Paper number
- 054
- Field
- Interpretability
- arXiv ID
- 2504.07096
Key points
- LLMs behave as black boxes, making it hard to understand why they produce specific outputs.
- Existing interpretability methods such as influence functions are computationally intractable at the trillions-of-token scale of modern LLM training datasets.
- The lack of tools to efficiently analyze memorization in LLMs, trace factual claims back to sources, or identify the origins of generated text raises concerns about transparency, copyright, and PII.
- OLMoTrace developed a system that uses an expanded infini-gram index to efficiently find maximal exact-match spans within a trillion-token-scale training dataset.
- The method has five stages: identify maximal matching spans, filter long and unique spans, retrieve containing documents, merge duplicates, and rerank documents by relevance using BM25.
- The production system uses high-IOPS SSDs on Google Cloud to guarantee the disk I/O throughput needed for real-time performance on large-scale data.
Paper links
External research summaries. These are not HDATF publications or measured product results.