Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control?

Published
Source
arXiv
Paper number
799
Field
Robotics
arXiv ID
2608.02547

Key points

  • It tested the three existing hypotheses with controlled experiments and concluded that neither temporal consistency, reduced effective horizon, nor representation learning alone can fully explain the performance of action chunking.
  • A lag policy that predicts only one step ahead from past observations replaces much of chunking's ability to use the past and reduce accumulated error. The validation error was lowest for lags between 5 and 15.
  • The chunking policy learns relationships between actions at different time gaps from a single observation, so it behaves like an ensemble of several lagged policies. The authors call this an implicit ensemble.
  • In every simulated and real-robot environment they tested, deploying the chunking policy as a random lag ensemble matched the performance of action chunking without using action chunks.
  • An explicit ensemble built from multiple independently trained policies performed even better than chunking. They report about a 30% improvement on the Robomimic Transport task.
  • At control frequencies between 15 and 20 Hz, the lag policy worked, but it failed when raised to 50 to 60 Hz. Treating partial chunks of length 5 as a single action made it work again, and the authors explain this by noting that human vision-based behavior updates at 2 to 10 Hz.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)