Sharpening Tax in Post-Training

Published
Source
arXiv
Paper number
1146
Field
AI / General
arXiv ID
2610.01509

Key points

  • The study analyzed 42 conditions by comparing 14 pairs of models before and after post-training across three benchmarks.
  • The study found that base models equipped with a simple tool-execution harness could solve more tasks than post-trained models when given enough retries.
  • The training procedure was designed to encourage more diverse actions on difficult tasks based on each task’s success history.
  • The method improved both single-attempt success rates and the range of tasks solved through repeated attempts in two game environments.
  • The authors stated that the observational results alone could not establish causality because the training data of the existing public models were unknown.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)