AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems
- Published
- Source
- arXiv
- Paper number
- 516
- Field
- AI / General
- arXiv ID
- 2606.26859
Key points
- The four-stage closed loop is Brainstorm, which proposes evidence-based ideas; Developing, which generates repo-grounded code; Evaluation, which performs guardrail-vetoed A/B decisions; and Harness Evolution, which applies SGPO.
- SGPO, or Semantic-Gradient-based Prompt Optimization, uses execution trajectories to update agent prompts and doubles weekly throughput.
- Three AgentX workers turned 374 ideas into 10 launchable rollouts over 3 weeks, which is 8x the concurrency of human engineers and 3.7x the business value.
- It uses online A/B feedback as a reward signal so that the work stays aligned with real business value, including a 0.561 percent increase in app dwell time and more than 100 million yuan in annual value.
- It shows that the approach can extend beyond recommendation strategy to model-side research, such as paper replication, module pruning, and architecture combination.
- It includes a negative-result assetization mechanism that turns failed experiments into structured knowledge assets.
Paper links
External research summaries. These are not HDATF publications or measured product results.