LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

Published
Source
arXiv
Paper number
668
Field
Machine Learning
arXiv ID
2607.18110

Key points

  • A single score is an information bottleneck because it cannot distinguish answers that receive the same top score, especially at the upper end.
  • The solution is to use the same feedback model as a coach rather than a judge. The coach turns answer evaluation into generalized, reusable advice, or experiential knowledge.
  • That advice is given to the teacher model as context and internalized by the policy through on-policy context distillation, which minimizes token-level reverse KL divergence between the teacher and the policy output. At inference time, neither the coach nor the advice context is needed.
  • The motivation is illustrated by feedback bandwidth. A 1 to 10 integer score carries 3.3 bits per sample, while a 1024-token piece of advice over a 150,000-word vocabulary theoretically carries 17,600 bits.
  • Across two policy families, EL outperformed rubric-based reinforcement learning on held-out and novel benchmarks, whether the feedback model was the policy itself or a commercial API model.
  • While RL is prone to reward hacking because it learns to chase reward on the training set, EL is a distribution-matching method that follows the teacher distribution and therefore reduces that failure mode.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)