Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

Published
Source
arXiv
Paper number
702
Field
LLMs / NLP
arXiv ID
2607.20911

Key points

  • It reverse-engineers real commit, pull request, and security scenarios to create tasks that cannot be found through web search.
  • The benchmark contains 260 tasks across four domains: 80 code, 70 web, 50 office, and 60 security tasks.
  • Each task is rewritten as a everyday, conversational persona request to prevent contamination.
  • Claude Opus 4.8 ranks first with an overall score of 75.0 percent, and GLM-5.2 ranks second with 72.9 percent.
  • The authors release the full code, environment images, and evaluation tools so anyone can reproduce and audit the benchmark.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)