AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

Published
Source
arXiv
Paper number
886
Field
Machine Learning
arXiv ID
2608.12307

Key points

  • It combines prompts, problem-specific handling rules, calculation code, and answer-format checks without changing the target model's weights.
  • The hidden test set contained 3,900 questions across four theory-of-mind tasks, and 195 separate validation items were provided for improving the harness.
  • GPT-5.4-mini's base average score was 0.488, while the overall average with the automatically generated harness was 0.763 and the best score was 0.912.
  • Deterministic code, problem-specific path selection, and strict output formatting had a larger effect than long reasoning or extra sampling.
  • The improvement was large on weak target models but some performance degradation also appeared on strong target models, and the experiments were limited to theory-of-mind tasks.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)