AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems

Published
Source
arXiv
Paper number
516
Field
AI / General
arXiv ID
2606.26859

Key points

  • The four-stage closed loop is Brainstorm, which proposes evidence-based ideas; Developing, which generates repo-grounded code; Evaluation, which performs guardrail-vetoed A/B decisions; and Harness Evolution, which applies SGPO.
  • SGPO, or Semantic-Gradient-based Prompt Optimization, uses execution trajectories to update agent prompts and doubles weekly throughput.
  • Three AgentX workers turned 374 ideas into 10 launchable rollouts over 3 weeks, which is 8x the concurrency of human engineers and 3.7x the business value.
  • It uses online A/B feedback as a reward signal so that the work stays aligned with real business value, including a 0.561 percent increase in app dwell time and more than 100 million yuan in annual value.
  • It shows that the approach can extend beyond recommendation strategy to model-side research, such as paper replication, module pruning, and architecture combination.
  • It includes a negative-result assetization mechanism that turns failed experiments into structured knowledge assets.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)