The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
- Published
- Source
- arXiv
- Paper number
- 097
- Field
- Benchmarks / Factuality
- arXiv ID
- 2512.10791
Key points
- Large language models frequently generate factually incorrect information, which undermines reliable deployment in real-world applications.
- Existing benchmarks for LLM factuality often focus on narrow abilities, such as context-grounded or general knowledge recall, and therefore fail to provide a holistic evaluation.
- Real LLM applications require combinations of factual capabilities across different modalities and knowledge sources, which current evaluation frameworks do not adequately cover.
- The FACTS Leaderboard consists of four separate sub-leaderboards, FACTS Multimodal, FACTS Parametric, FACTS Search, and FACTS Grounding v2, each of which evaluates a specific dimension of LLM factuality.
- The benchmark evaluates model responses across thousands of diverse questions using sophisticated automatic judges that have been rigorously validated against human annotations.
- Kaggle conducts all evaluations fairly, with public and private data splits that preserve benchmark integrity and reduce overfitting.
Paper links
External research summaries. These are not HDATF publications or measured product results.