AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
- Published
- Source
- arXiv
- Paper number
- 886
- Field
- Machine Learning
- arXiv ID
- 2608.12307
Key points
- It combines prompts, problem-specific handling rules, calculation code, and answer-format checks without changing the target model's weights.
- The hidden test set contained 3,900 questions across four theory-of-mind tasks, and 195 separate validation items were provided for improving the harness.
- GPT-5.4-mini's base average score was 0.488, while the overall average with the automatically generated harness was 0.763 and the best score was 0.912.
- Deterministic code, problem-specific path selection, and strict output formatting had a larger effect than long reasoning or extra sampling.
- The improvement was large on weak target models but some performance degradation also appeared on strong target models, and the experiments were limited to theory-of-mind tasks.
Paper links
External research summaries. These are not HDATF publications or measured product results.