ACE-RL: Adaptive Constraint-Enhanced Reward for Long-form Generation Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 081
- Field
- RL / Writing
- arXiv ID
- 2509.04903
Key points
- Existing supervised fine-tuning approaches for long-form generation rely heavily on high-quality datasets that are scarce and expensive.
- Current reinforcement learning, or RL, methods use coarse and subjective evaluation metrics, which fail to capture the subtle and instruction-specific requirements that matter for high-quality long-form writing.
- LLMs struggle to consistently satisfy fine-grained, multi-faceted, and instruction-specific requirements in complex long-form generation tasks.
- An automated pipeline using a strong LLM decomposes high-level user instructions into a verifiable checklist of constraints covering content completeness, structural logic, and stylistic formality.
- The new reward mechanism combines length rewards with fine-grained constraint rewards, and a smaller verifier LLM evaluates whether each constraint is satisfied with a three-level judgment.
- Training uses the Group Relative Policy Optimization, or GRPO, algorithm, which encourages the LLM to produce long-form responses that satisfy both length and instruction-specific requirements through this comprehensive reward.
Paper links
External research summaries. These are not HDATF publications or measured product results.