WHALE: A Simple Recipe for Joint Harness-Weight Optimization

Published
Source
arXiv
Paper number
1059
Field
Machine Learning
arXiv ID
2609.00196

Key points

  • It alternates between online rejection-sampling fine-tuning with the current harness fixed and Meta-Harness code search with the updated model fixed.
  • Evaluations of Qwen3.5-2B and 4B on search-based question answering, mathematical reasoning, and chess puzzles achieved peak accuracies 7.67–24.38 percentage points higher than single-component baselines and 4.15–13.00 percentage points higher than FST, which changes only prompts.
  • The harness was the main bottleneck for search tasks, but mathematical tasks required weight updates first before the same harness search became effective, showing that the bottleneck varies by task.
  • Interleaving multiple small updates improved accuracy and rollout efficiency over a staged approach that trains all the weights and then searches for a harness once, providing grounds for designing agent-training pipelines as a joint optimization problem.
  • Important limitations are that validation covered only Qwen3.5-2B and 4B and three tasks, and that the rollout-cost comparison excluded compute used by the model proposing harnesses.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)