ISO: An RLVR-Native Optimization Stack

Published
Source
arXiv
Paper number
693
Field
Machine Learning
arXiv ID
2607.19331

Key points

  • It discovers and validates a 'spectrum inheritance' phenomenon in which the weight spectrum changes very little after RLVR training.
  • ISO-Optimizer matches AdamW accuracy on Qwen3-8B using 2.7 times fewer steps, from 270 down to 100.
  • ISO-Merger merges expert models using checkpoints alone, without data, rollouts, or gradients.
  • Freezing the spectrum alone does not improve performance much, which means both frames need to remain trainable.
  • It shows consistent efficiency gains on reasoning and coding tasks at scales from 1.5B to 8B.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)