Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control?
- Published
- Source
- arXiv
- Paper number
- 799
- Field
- Robotics
- arXiv ID
- 2608.02547
Key points
- It tested the three existing hypotheses with controlled experiments and concluded that neither temporal consistency, reduced effective horizon, nor representation learning alone can fully explain the performance of action chunking.
- A lag policy that predicts only one step ahead from past observations replaces much of chunking's ability to use the past and reduce accumulated error. The validation error was lowest for lags between 5 and 15.
- The chunking policy learns relationships between actions at different time gaps from a single observation, so it behaves like an ensemble of several lagged policies. The authors call this an implicit ensemble.
- In every simulated and real-robot environment they tested, deploying the chunking policy as a random lag ensemble matched the performance of action chunking without using action chunks.
- An explicit ensemble built from multiple independently trained policies performed even better than chunking. They report about a 30% improvement on the Robomimic Transport task.
- At control frequencies between 15 and 20 Hz, the lag policy worked, but it failed when raised to 50 to 60 Hz. Treating partial chunks of length 5 as a single action made it work again, and the authors explain this by noting that human vision-based behavior updates at 2 to 10 Hz.
Paper links
External research summaries. These are not HDATF publications or measured product results.