WHALE: A Simple Recipe for Joint Harness-Weight Optimization
- Published
- Source
- arXiv
- Paper number
- 1059
- Field
- Machine Learning
- arXiv ID
- 2609.00196
Key points
- It alternates between online rejection-sampling fine-tuning with the current harness fixed and Meta-Harness code search with the updated model fixed.
- Evaluations of Qwen3.5-2B and 4B on search-based question answering, mathematical reasoning, and chess puzzles achieved peak accuracies 7.67–24.38 percentage points higher than single-component baselines and 4.15–13.00 percentage points higher than FST, which changes only prompts.
- The harness was the main bottleneck for search tasks, but mathematical tasks required weight updates first before the same harness search became effective, showing that the bottleneck varies by task.
- Interleaving multiple small updates improved accuracy and rollout efficiency over a staged approach that trains all the weights and then searches for a harness once, providing grounds for designing agent-training pipelines as a joint optimization problem.
- Important limitations are that validation covered only Qwen3.5-2B and 4B and three tasks, and that the rollout-cost comparison excluded compute used by the model proposing harnesses.
Paper links
External research summaries. These are not HDATF publications or measured product results.