Neurosymbolic Embodied Agents

Published
Source
arXiv
Paper number
930
Field
Robotics
arXiv ID
2608.16794

Key points

  • The first stage checks object positions and relations through multi-view and in-simulator interactions, such as opening or picking up objects. It does not place unseen objects into the state using the model's own guess.
  • The second stage only generates action tokens that satisfy the current state constraints. Search time is up to 250 seconds per task, and the completed plan receives no feedback during execution.
  • We evaluate 200 VirtualHome tasks and 134 hidden splits from ALFWorld. The total action budget is 20 and 55, respectively, and the model is not shown scene graphs or ground-truth states.
  • For Qwen3.5-27B, ALFWorld success is 32.3% with constraints only, 29.2% with search only, and 95.5% when both are combined. Averaged across both environments at the same model scale, the token cost is up to 4x lower than long-reasoning methods.
  • Most remaining failures are concentrated in the first stage. The rate of getting the state wrong is about 2.3% on VirtualHome and about 4.5% on ALFWorld, and the typical visual error is missing small objects.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)