RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing

Published
Source
arXiv
Paper number
1180
Field
Agents
arXiv ID
2610.10507

Key points

  • They formalized the problem of deriving evidence through filtering, aggregation, and computation rather than mere retrieval as a sequential decision process.
  • RouterLM alternates between calling primitive operations and synthesizing custom code to build evidence, trained with SFT and GRPO.
  • It achieved a 75.6% mean success rate across six heterogeneous benchmarks, 15.9 percentage points above the strongest baseline.
  • The compact 9B trained model outperformed a training-free Gemini 3.5 Flash router by 5.0 percentage points.
  • On three unseen benchmarks it stayed 15.0 percentage points ahead on average, demonstrating zero-shot generalization.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)