ACE-RL: Adaptive Constraint-Enhanced Reward for Long-form Generation Reinforcement Learning

Published
Source
arXiv
Paper number
081
Field
RL / Writing
arXiv ID
2509.04903

Key points

  • Existing supervised fine-tuning approaches for long-form generation rely heavily on high-quality datasets that are scarce and expensive.
  • Current reinforcement learning, or RL, methods use coarse and subjective evaluation metrics, which fail to capture the subtle and instruction-specific requirements that matter for high-quality long-form writing.
  • LLMs struggle to consistently satisfy fine-grained, multi-faceted, and instruction-specific requirements in complex long-form generation tasks.
  • An automated pipeline using a strong LLM decomposes high-level user instructions into a verifiable checklist of constraints covering content completeness, structural logic, and stylistic formality.
  • The new reward mechanism combines length rewards with fine-grained constraint rewards, and a smaller verifier LLM evaluates whether each constraint is satisfied with a three-level judgment.
  • Training uses the Group Relative Policy Optimization, or GRPO, algorithm, which encourages the LLM to produce long-form responses that satisfy both length and instruction-specific requirements through this comprehensive reward.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)