Sharpening Tax in Post-Training
- Published
- Source
- arXiv
- Paper number
- 1146
- Field
- AI / General
- arXiv ID
- 2610.01509
Key points
- The study analyzed 42 conditions by comparing 14 pairs of models before and after post-training across three benchmarks.
- The study found that base models equipped with a simple tool-execution harness could solve more tasks than post-trained models when given enough retries.
- The training procedure was designed to encourage more diverse actions on difficult tasks based on each task’s success history.
- The method improved both single-attempt success rates and the range of tasks solved through repeated attempts in two game environments.
- The authors stated that the observational results alone could not establish causality because the training data of the existing public models were unknown.
Paper links
External research summaries. These are not HDATF publications or measured product results.