UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning

Published
Source
arXiv
Paper number
568
Field
LLMs / NLP
arXiv ID
2607.04425

Key points

  • MOPD, multi-teacher on-policy distillation, is introduced for the first time to continual cross-platform learning for GUI agents.
  • Uni-GUI is a dataset of about 10K high-quality cross-platform GUI interaction trajectories.
  • A platform router dynamically selects a teacher based on the environment, mixing behavior rules and preventing catastrophic forgetting.
  • It achieves relative gains of 38.2% on OSWorld and 12.0% on MobileWorld, or 12.7% and 55.8% depending on the baseline, showing that a single 8B model improves on both platforms at once.
  • GUI grounding ability is also preserved, with 43.14% on ScreenSpot-Pro and 90.88% on ScreenSpotV2, showing minimal loss relative to the base model.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)