InstructSAM: Segment Any Instance with Any Instructions

Published
Source
arXiv
Paper number
249
Field
Computer Vision
arXiv ID
2605.26102

Key points

  • During decoding, the final score head determines whether the query corresponds to a valid target instance, and the segmentation head produces a binary mask.
  • For a single target, it identifies the one specific object mentioned in the instruction.
  • For multiple targets, it finds all instances that satisfy a complex condition, such as "all wooden chairs."

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)