JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

Published
Source
arXiv
Paper number
1023
Field
LLMs / NLP
arXiv ID
2608.25593

Key points

  • Viewing agent capability as the combined result of the model and harness, it proposed a JIT approach that generates a harness on the fly for each task instead of fixing one in advance.
  • It represents the harness as executable code divided into four modules, memory, planning, action, and capability, enabling generation, repair, and evolution.
  • It trained the generator by combining task-specific learning, failure-recovery learning, and Evo-GDPO, which accumulates better-performing harnesses.
  • DeepSeek-V4-Flash surpassed GPT-5.6 by 9.1 and 4.3 points on two benchmarks, while GLM-5.2 also improved by up to 20.2 points depending on the task.
  • Changing the entire execution framework on the fly leaves stability and verification issues in operational environments, and the authors themselves identified a stable core with verifiable partial modifications as future work.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)