OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens

Published
Source
arXiv
Paper number
054
Field
Interpretability
arXiv ID
2504.07096

Key points

  • LLMs behave as black boxes, making it hard to understand why they produce specific outputs.
  • Existing interpretability methods such as influence functions are computationally intractable at the trillions-of-token scale of modern LLM training datasets.
  • The lack of tools to efficiently analyze memorization in LLMs, trace factual claims back to sources, or identify the origins of generated text raises concerns about transparency, copyright, and PII.
  • OLMoTrace developed a system that uses an expanded infini-gram index to efficiently find maximal exact-match spans within a trillion-token-scale training dataset.
  • The method has five stages: identify maximal matching spans, filter long and unique spans, retrieve containing documents, merge duplicates, and rerank documents by relevance using BM25.
  • The production system uses high-IOPS SSDs on Google Cloud to guarantee the disk I/O throughput needed for real-time performance on large-scale data.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)