MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows

Published
Source
arXiv
Paper number
1138
Field
AI / General
arXiv ID
2609.37053

Key points

  • MatToolBench provides 204 tasks using ten materials-science tools inside a Windows 11 virtual machine.
  • It evaluates GUI operation, OriginPro scripting, and code-based queries, distinguishing partial-credit scores from success that satisfies every criterion.
  • The highest average success rates among the evaluated models are 25% on GUI tasks and 45% on code tasks.
  • Performance drops when domain-specific workflow guidance is removed, highlighting the importance of tool knowledge and procedural support.
  • Evaluation is limited to English prompts and single sessions of at most 50 steps, excluding some specialist tools and longer workflows.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)