Last Translation Benchmark

Published
Source
arXiv
Paper number
1069
Field
LLMs / NLP
arXiv ID
2609.04173

Key points

  • Humans created and peer-reviewed 3456 challenging translation examples across 109 languages in text, image, audio, and video formats.
  • Each example includes verification rules specifying how it should be translated, and an LLM judge assigns a pass or fail for each rule. An example counts as successful only if every rule passes.
  • On a simple English-to-Czech sentence involving the gendered expression for a male nurse, Google Translate and Gemini 3.1 Pro scored 0%, and only GPT-5.6 Sol passed. This shows that even strong models make critical mistakes.
  • In an experiment transferring examples to other languages, automatic filtering retained 48%, while human review passed 64.1%, demonstrating the feasibility of semi-automatic expansion.
  • The data is released under CC BY 4.0, and this is a living benchmark that continues to accept contributions.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)