Last Translation Benchmark
- Published
- Source
- arXiv
- Paper number
- 1069
- Field
- LLMs / NLP
- arXiv ID
- 2609.04173
Key points
- Humans created and peer-reviewed 3456 challenging translation examples across 109 languages in text, image, audio, and video formats.
- Each example includes verification rules specifying how it should be translated, and an LLM judge assigns a pass or fail for each rule. An example counts as successful only if every rule passes.
- On a simple English-to-Czech sentence involving the gendered expression for a male nurse, Google Translate and Gemini 3.1 Pro scored 0%, and only GPT-5.6 Sol passed. This shows that even strong models make critical mistakes.
- In an experiment transferring examples to other languages, automatic filtering retained 48%, while human review passed 64.1%, demonstrating the feasibility of semi-automatic expansion.
- The data is released under CC BY 4.0, and this is a living benchmark that continues to accept contributions.
Paper links
External research summaries. These are not HDATF publications or measured product results.