CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents

Published
Source
arXiv
Paper number
242
Field
AI / General
arXiv ID
2605.25624

Key points

  • The discriminator is isolated from the generator's implementation details and sees only the task t and the resulting environment state, and it then writes a reward.py script that should return 0.0 for the initial state and 1.0 for the golden state.
  • The Qwen3.5-35B model (A3B) improved success rate from 54.5 percent to 62.1 percent.
  • The Qwen3.5-397B model (A17B) reached a 72.6 percent success rate, which is an increase of more than 10 percentage points.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)