A Bitter Lesson for Data Filtering

Published
Source
arXiv
Paper number
192
Field
Machine Learning
arXiv ID
2605.19407

Key points

  • Use a fixed number of epochs and train for four epochs.
  • Simplify the pipeline. Instead of spending thousands of GPU hours and a lot of human engineering effort on complex classifiers and heuristic rules, simply increasing training time on raw parsed web data can yield better returns.
  • Even trash has value. Robustness to mixed and random data suggests that, if the model is large enough, almost any data containing some statistical signal related to the target distribution can be useful.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)