A Bitter Lesson for Data Filtering
- Published
- Source
- arXiv
- Paper number
- 192
- Field
- Machine Learning
- arXiv ID
- 2605.19407
Key points
- Use a fixed number of epochs and train for four epochs.
- Simplify the pipeline. Instead of spending thousands of GPU hours and a lot of human engineering effort on complex classifiers and heuristic rules, simply increasing training time on raw parsed web data can yield better returns.
- Even trash has value. Robustness to mixed and random data suggests that, if the model is large enough, almost any data containing some statistical signal related to the target distribution can be useful.
Paper links
External research summaries. These are not HDATF publications or measured product results.