HDATF RESEARCH / 基准评测报告

如何验证AI 能否完成实际工作

公开成绩能说明什么,哪些问题必须通过实际运行验证。 从模型选择到调研、文档编写与成果检验。

核心结论

公开评测为模型选择和测试设计提供依据。能否投入工作,还要连接实际资料和工具,检验系统是否完成所要求的成果。

本报告从可用成果、可核查证据和完整执行三个角度解读公开成绩。HDATF尚未公布这些评测的自身实测成绩,最后一章将说明验证这一差距的计划。

73 评测与指标684 归档成绩

01 / 如何读懂成绩

先确定要交付的工作,再阅读分数。

基准评测用规定的任务和评分规则比较能力。推理题、文档比较和工具操作分别考察不同能力,读分数时应先明确测试对象。

报告任务需要文件可打开、论点有据可查、内容符合要求。数据分析还需要能用提供的数据重现结果。通用模型排名无法单独验证这些条件。

公开成绩用于缩小候选范围并确定要测试的失败类型。选择用于LabChin的方法时,需要结合自己的资料、权限、工具与成果检查。本报告依次讨论成果、证据和执行。[1]

02 / 成果质量

按实际任务要求评价成果。

GDPval比较专业工作成果。最初的Gold研究涵盖44个职业的220项任务,由专家在不知道作者的情况下比较。下图展示归档的三项结果。

胜出与持平比例反映这些成果比较,不能换算为企业可交给AI的工作比例。任务范围、专家判断和模型配置限定了结论。[2]

图 01

GDPval Gold

相对专家成果的胜出与持平比例 (%)

%越高越好
  1. Claude Opus 4.147.6%
  2. GPT-5 (high)38.8%
  3. o3 (high)34.1%
GDPval Gold / %
模型与运行配置成绩来源
Claude Opus 4.1Anthropic47.6%原始记录
GPT-5 (high)OpenAI38.8%原始记录
o3 (high)OpenAI34.1%原始记录

44 个职业的 220 项 Gold 任务。原始研究由专家盲评比较成果,不代表当前排名。

来源: OpenAI2025-09原始记录

HDATF需要确认文件生成后还有多少人工工作。应保留遗漏要求、无依据数值和人工修改。易读的布局是质量的一部分,但不能弥补错误证据或不可用文件。

03 / 调研与证据

同时检查寻找答案和提供依据的能力。

BrowseComp测试寻找难查事实的能力。Parallel从1,266题中随机抽取100题,报告正确率与请求成本。以下成绩来自服务提供方对部分题目的实验。[3][4]

图 02

BrowseComp / Parallel study

正确率 (%)

%越高越好
BrowseComp / Parallel study / %
模型与运行配置成绩来源
Parallel Ultra8xParallel58%原始记录
Parallel UltraParallel45%原始记录
GPT-5 (high)OpenAI38%原始记录

从 1,266 题随机抽取 100 题。Parallel 于 2025 年 8 月 11 至 29 日测量。每次费用:Ultra8x $2.40、Ultra $0.30、GPT-5 $0.488。配置与预算不同。

来源: Parallel2025-09-09原始记录

Ultra8x报告每次2.40美元、正确率58%,Ultra为0.30美元、45%。预算不同,这些数值呈现准确率与成本关系,但不能证明同等成本下的优势或基础模型的独立效果。

调研报告需要额外检查。DeepResearch Bench II和ResearchRubrics依据明确标准评估内容。LabChin计划保留支持主要论点的原文、矛盾资料及未回答的问题。

04 / 执行与记忆

同一模型,执行方式不同,结果也会改变。

执行框架向模型提供工具、上下文并控制任务。LangChain固定GPT-5.2-Codex,仅修改周边执行方式,89项Terminal-Bench 2.0任务的成绩从52.8%升至66.5%。[5][6]

图 03

Terminal-Bench 2.0

任务成功率 (%)

%越高越好
Terminal-Bench 2.0 / %
模型与运行配置成绩来源
Deep Agents / improvedLangChain66.5%原始记录
Deep Agents / baselineLangChain52.8%原始记录

89 项任务。LangChain 固定 GPT-5.2-Codex,修改提示、工具及执行控制,使用 Harbor / Daytona。公布提升为 13.7 个百分点。

来源: LangChain2026-02-17原始记录

实验同时改变指令、工具和执行控制,包括结束前检查与环境说明。13.7个百分点属于组合修改的结果,无法据此确定各功能贡献或预测HDATF的提升。

可以固定自己的任务与模型,对比有无成果检查的流程。保留文件、检查结果、重试和耗时,同时解释成功及需要人工干预的失败,才能判断执行差异。

记忆需要独立测试。回忆旧信息还不够,也要处理修正和访问范围。归档的LoCoMo结果仅限单跳事实回忆,不能视为已验证更新冲突或用户隔离,因此列入附录供参考。

05 / HDATF评测计划

下一步验证HDATF实际完成的工作。

公开研究有助于整理问题,产品成绩仍需直接测量。计划记录代表任务的执行条件,分别评估证据检索、分析和成果交付,避免某一项高分掩盖其他失败。

调研和数据工作应保留问题、资料版本及最终成果检查记录。报告需暴露无依据论点,分析需保存输入数据、运行记录与可重现成果。以下是待验证的标准,并非已完成测试。[1][7][8][9]

这是评测计划,HDATF实测成绩尚未公布。

能否找到所需证据

BrowseComp评估查找网络中难以找到的事实的能力。LabChin还计划单独检查引用来源是否真正支持相应论点。

能否依据证据回答问题

以DeepResearch Bench II和ResearchRubrics作为调研报告的评估参考。不只看行文流畅,还计划检查内容覆盖、事实依据和结论的分析过程。

能否交付可用成果

GDPval和Agents’ Last Exam面向专业任务。计划检查请求是否完成、文件能否使用,以及还需要多少人工修正。

评测对象保留的证据支持的判断
任务任务要求、输入文件、完成标准所要求的工作是否完成
执行模型、工具、权限、执行配置版本能否以相同条件重新运行
结果成果文件、来源检查、人工修改成果是否正确且可用
效率总耗时、工具使用、重试、总成本整个流程的时间与成本是否可承担
APPENDIX

附录。完整评测资料

按评测比较公开成绩,点击模型名称查看原文。数据以标注日期为准,模型系列比较采用各系列成绩最好的配置。

73 / 73 评测与指标

工作与文档

AA-Briefcase15 项成绩

基础模型成绩Artificial Analysis

长期知识工作

Elo越高越好
  1. Claude Fable 5.11,661.91
  2. Claude Opus 51,647.02
  3. GPT-6 Astra1,561.95
  4. Muse Spark 1.31,558.85
  5. Grok 4.61,545.82
  6. Claude Fable 51,533.52
  7. GLM-5.31,515.45
  8. Kimi K31,496.83
  9. GPT-5.6 Sol1,474.49
  10. GLM 5.3 Flash1,459.13
  11. Qwen3.8 2.4T A95B1,442.81
  12. DeepSeek V4 Flash 07311,433.23
  13. Qwen3.8 27B1,401.81
  14. Qwen3.8 Max1,392.33
  15. Claude Sonnet 51,358.11
AA-Briefcase / Elo
模型与运行配置成绩来源
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic1,661.91原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic1,647.02原始记录
GPT-6 Astra (max)OpenAI1,561.95原始记录
Muse Spark 1.3 (max)Meta1,558.85原始记录
Grok 4.6 (xhigh)SpaceXAI1,545.82原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic1,533.52原始记录
GLM-5.3 (max)Z AI1,515.45原始记录
Kimi K3 (max)Kimi1,496.83原始记录
GPT-5.6 Sol (max)OpenAI1,474.49原始记录
GLM 5.3 Flash (GLM-5.3-Flash)Z AI1,459.13原始记录
Qwen3.8 2.4T A95BAlibaba1,442.81原始记录
DeepSeek V4 Flash 0731 (DeepSeek V4 Flash Vision (Reasoning, Max Effort))DeepSeek1,433.23原始记录
Qwen3.8 27B (xhigh)Alibaba1,401.81原始记录
Qwen3.8 MaxAlibaba1,392.33原始记录
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic1,358.11原始记录
AA-Briefcase / Rubric15 项成绩

基础模型成绩Artificial Analysis

工作要求满足率

%越高越好
AA-Briefcase / Rubric / %
模型与运行配置成绩来源
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic61.52%原始记录
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic57.98%原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic55.96%原始记录
Muse Spark 1.3 (max)Meta54.87%原始记录
Grok 4.6 (xhigh)SpaceXAI53.84%原始记录
GPT-6 Astra (xhigh)OpenAI52.32%原始记录
GLM-5.3 (max)Z AI51.35%原始记录
Kimi K3 (max)Kimi50.97%原始记录
Qwen3.8 2.4T A95BAlibaba50%原始记录
Qwen3.8 MaxAlibaba49.6%原始记录
Qwen3.8 27B (xhigh)Alibaba47.8%原始记录
GLM 5.3 Flash (GLM-5.3-Flash)Z AI46.36%原始记录
DeepSeek V4 Flash 0731 (DeepSeek V4 Flash Vision (Reasoning, Max Effort))DeepSeek46.15%原始记录
Muse Spark 1.2 (xhigh)Meta45.15%原始记录
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek43.52%原始记录
AA-Briefcase / Analytical quality15 项成绩

基础模型成绩Artificial Analysis

分析质量

Elo越高越好
  1. Claude Fable 5.11,959.49
  2. Claude Opus 51,909.74
  3. GLM-5.31,788.62
  4. GPT-6 Astra1,753.64
  5. Muse Spark 1.31,711.52
  6. Grok 4.61,682.58
  7. Claude Fable 51,676.55
  8. Kimi K31,668.89
  9. GLM 5.3 Flash1,650.95
  10. Qwen3.8 2.4T A95B1,617.45
  11. DeepSeek V4 Flash 07311,617.37
  12. Qwen3.8 27B1,590.24
  13. GPT-5.6 Sol1,537.92
  14. Qwen3.8 Max1,533.14
  15. Claude Sonnet 51,418.22
AA-Briefcase / Analytical quality / Elo
模型与运行配置成绩来源
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic1,959.49原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic1,909.74原始记录
GLM-5.3 (max)Z AI1,788.62原始记录
GPT-6 Astra (max)OpenAI1,753.64原始记录
Muse Spark 1.3 (max)Meta1,711.52原始记录
Grok 4.6 (xhigh)SpaceXAI1,682.58原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic1,676.55原始记录
Kimi K3 (max)Kimi1,668.89原始记录
GLM 5.3 Flash (GLM-5.3-Flash)Z AI1,650.95原始记录
Qwen3.8 2.4T A95BAlibaba1,617.45原始记录
DeepSeek V4 Flash 0731 (DeepSeek V4 Flash Vision (Reasoning, Max Effort))DeepSeek1,617.37原始记录
Qwen3.8 27B (xhigh)Alibaba1,590.24原始记录
GPT-5.6 Sol (max)OpenAI1,537.92原始记录
Qwen3.8 MaxAlibaba1,533.14原始记录
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic1,418.22原始记录
AA-Briefcase / Presentation15 项成绩

基础模型成绩Artificial Analysis

成果呈现质量

Elo越高越好
  1. GPT-5.6 Sol1,633.93
  2. Claude Opus 51,544.01
  3. GPT-6 Astra1,535.29
  4. GPT-5.6 Terra1,519.13
  5. Grok 4.61,514.68
  6. Muse Spark 1.31,506.83
  7. Claude Fable 5.11,472.7
  8. GPT-5.6 Luna1,471.03
  9. Claude Fable 51,456.83
  10. Kimi K31,435.29
  11. Claude Sonnet 51,409.5
  12. GLM 5.3 Flash1,409.12
  13. DeepSeek V4 Flash 07311,355.99
  14. GLM-5.31,354.5
  15. Muse Spark 1.21,334.35
AA-Briefcase / Presentation / Elo
模型与运行配置成绩来源
GPT-5.6 Sol (max)OpenAI1,633.93原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic1,544.01原始记录
GPT-6 Astra (max)OpenAI1,535.29原始记录
GPT-5.6 Terra (max)OpenAI1,519.13原始记录
Grok 4.6 (xhigh)SpaceXAI1,514.68原始记录
Muse Spark 1.3 (max)Meta1,506.83原始记录
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic1,472.7原始记录
GPT-5.6 Luna (max)OpenAI1,471.03原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic1,456.83原始记录
Kimi K3 (max)Kimi1,435.29原始记录
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic1,409.5原始记录
GLM 5.3 Flash (GLM-5.3-Flash)Z AI1,409.12原始记录
DeepSeek V4 Flash 0731 (DeepSeek V4 Flash Vision (Reasoning, Max Effort))DeepSeek1,355.99原始记录
GLM-5.3 (max)Z AI1,354.5原始记录
Muse Spark 1.2 (xhigh)Meta1,334.35原始记录
GDPval-AA v215 项成绩

基础模型成绩Artificial Analysis

实际职业任务

Elo越高越好
GDPval-AA v2 / Elo
模型与运行配置成绩来源
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic1,765.9原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic1,738.08原始记录
Muse Spark 1.3 (max)Meta1,719.65原始记录
GLM-5.3 (max)Z AI1,678.47原始记录
GLM 5.3 Flash (GLM-5.3-Flash)Z AI1,672.94原始记录
Grok 4.6 (xhigh)SpaceXAI1,664.82原始记录
Qwen3.8-Flash-NextAlibaba1,649.93原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic1,635.39原始记录
Qwen3.8 MaxAlibaba1,633.53原始记录
Qwen3.8 2.4T A95BAlibaba1,630.39原始记录
GPT-5.6 Sol (max)OpenAI1,625.68原始记录
Kimi K3 (max)Kimi1,585.72原始记录
GPT-6 Astra (max)OpenAI1,582.24原始记录
DeepSeek V4 Flash 0731 (DeepSeek V4 Flash Vision (Reasoning, Max Effort))DeepSeek1,579.3原始记录
Muse Spark 1.2 (xhigh)Meta1,525.03原始记录
AA-AnalystAgent / pass^515 项成绩

基础模型成绩Artificial Analysis

表格与文档分析的重复成功

%越高越好
AA-AnalystAgent / pass^5 / %
模型与运行配置成绩来源
Gemini 3.7 Flash (high)Google60%原始记录
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic57.5%原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic53.75%原始记录
GPT-6 Astra (max)OpenAI51.25%原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic48.75%原始记录
GPT-5.6 Sol (max)OpenAI47.5%原始记录
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic46.25%原始记录
Gemini 3.1 Pro PreviewGoogle41.25%原始记录
Grok 4.6 (high)SpaceXAI41.25%原始记录
Kimi K3 (max)Kimi38.75%原始记录
Grok 4.5 (high)SpaceXAI35%原始记录
Inkling SmallThinking Machines27.5%原始记录
Inkling (xhigh)Thinking Machines23.75%原始记录
MiMo-V2.5-ProXiaomi20%原始记录
DeepSeek V4 Pro 0424 (DeepSeek V4 Pro (Reasoning, Max Effort))DeepSeek18.75%原始记录
GDP.pdf / All-pass15 项成绩

基础模型成绩Artificial Analysis

专业文档推理

%越高越好
GDP.pdf / All-pass / %
模型与运行配置成绩来源
GPT-6 Astra (max)OpenAI33.2%原始记录
GPT-5.6 Sol (max)OpenAI28.2%原始记录
Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Anthropic28%原始记录
Muse Spark 1.3 (max)Meta25.6%原始记录
GPT-5.6 Terra (max)OpenAI25.6%原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic24%原始记录
GPT-5.6 Luna (xhigh)OpenAI23.8%原始记录
Gemini 3.7 Flash (high)Google23.6%原始记录
Gemini 3.8 Flash (medium)Google22.8%原始记录
Qwen3.8 MaxAlibaba21.8%原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic21.6%原始记录
GPT-5.5 Instant (June 2026)OpenAI20.2%原始记录
Kimi K3 (max)Kimi19.6%原始记录
Grok 4.5 (high)SpaceXAI18.8%原始记录
Grok 4.6 (high)SpaceXAI18.8%原始记录
AutomationBench-AA15 项成绩

基础模型成绩Artificial Analysis

业务应用目标完成

%越高越好
AutomationBench-AA / %
模型与运行配置成绩来源
Gemini 3.7 Flash (high)Google62.75%原始记录
Kimi K3 (max)Kimi52.71%原始记录
Grok 4.5 (high)SpaceXAI51.44%原始记录
GPT-5.6 Sol (max)OpenAI51.19%原始记录
Gemini 3.6 Flash (high)Google51.08%原始记录
Gemini 3.8 Flash (high)Google50.72%原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic48.58%原始记录
GPT-5.6 Terra (max)OpenAI45.61%原始记录
GPT-5.6 Luna (max)OpenAI42.24%原始记录
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic39.19%原始记录
Gemini 3.1 Pro PreviewGoogle37.54%原始记录
Gemini 3.5 Flash-LiteGoogle32.66%原始记录
GLM-5.2 (max)Z AI27.84%原始记录
Kimi K2.7 CodeKimi22.53%原始记录
Qwen3.7 PlusAlibaba20.43%原始记录
EnterpriseOps-Gym-AA15 项成绩

基础模型成绩Artificial Analysis

企业运营任务

%越高越好
EnterpriseOps-Gym-AA / %
模型与运行配置成绩来源
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic51.12%原始记录
Gemini 3.7 Flash (medium)Google50.4%原始记录
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek49.6%原始记录
Grok 4.6 (high)SpaceXAI48.34%原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic47.48%原始记录
Qwen3.8 2.4T A95BAlibaba47.36%原始记录
Muse Spark 1.2 (xhigh)Meta47.27%原始记录
DeepSeek V4 Flash 0731 (DeepSeek V4 Flash Vision (Reasoning, Max Effort))DeepSeek47.09%原始记录
Kimi K3 (max)Kimi45.33%原始记录
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic44.67%原始记录
Qwen3.8 27B (xhigh)Alibaba44.23%原始记录
GPT-5.6 Sol (max)OpenAI42.91%原始记录
GLM-5.2 (max)Z AI42.73%原始记录
Inkling SmallThinking Machines42.7%原始记录
Gemini 3.5 Flash-LiteGoogle42.35%原始记录
APEX-Agents-AA14 项成绩

基础模型成绩Artificial Analysis

跨应用长期任务

%越高越好
APEX-Agents-AA / %
模型与运行配置成绩来源
Kimi K3 (max)Kimi41.3%原始记录
GPT-5.6 Terra (max)OpenAI38.94%原始记录
GPT-5.6 Luna (max)OpenAI35.84%原始记录
GLM-5.2 (max)Z AI33.7%原始记录
Gemini 3.1 Pro PreviewGoogle32.01%原始记录
Apodex 1.1Apodex31.19%原始记录
DeepSeek V4 Pro 0424 (DeepSeek V4 Pro (Reasoning, Max Effort))DeepSeek24.26%原始记录
Qwen3.7 PlusAlibaba22.42%原始记录
Qwen3.5 397B A17B (Reasoning)Alibaba15.34%原始记录
Step 3.7 FlashStepFun14.82%原始记录
gpt-oss-120b (high)OpenAI3.1%原始记录
MiMo-V2.5-ProXiaomi2.43%原始记录
Nemotron 3 Super 120B A12B (Reasoning)NVIDIA1.84%原始记录
gpt-oss-20b (high)OpenAI0.74%原始记录

系统研究

GDPval Gold3 项成绩

专家比较研究OpenAI

AI 能否交付专业人员认可的成果? 相对专家成果的胜出与持平比例 (%)

%越高越好
  1. Claude Opus 4.147.6%
  2. GPT-5 (high)38.8%
  3. o3 (high)34.1%
GDPval Gold / %
模型与运行配置成绩来源
Claude Opus 4.1Anthropic47.6%原始记录
GPT-5 (high)OpenAI38.8%原始记录
o3 (high)OpenAI34.1%原始记录
BrowseComp / Parallel study3 项成绩

检索智能体研究Parallel

智能体能否找到难以检索的信息并连接证据? 正确率 (%)

%越高越好
BrowseComp / Parallel study / %
模型与运行配置成绩来源
Parallel Ultra8xParallel58%原始记录
Parallel UltraParallel45%原始记录
GPT-5 (high)OpenAI38%原始记录
Terminal-Bench 2.02 项成绩

执行体系研究LangChain

改变执行机制能让结果发生多大变化? 任务成功率 (%)

%越高越好
Terminal-Bench 2.0 / %
模型与运行配置成绩来源
Deep Agents / improvedLangChain66.5%原始记录
Deep Agents / baselineLangChain52.8%原始记录
LoCoMo / single-hop3 项成绩

记忆系统研究Mem0 research team

跨越不同对话后,是否保留有用信息? 单跳事实回忆,AI 评分(0~100)

score越高越好
  1. Mem067.13
  2. LangMem62.23
  3. Zep61.7
LoCoMo / single-hop / score
模型与运行配置成绩来源
Mem0Mem067.13原始记录
LangMemLangChain62.23原始记录
ZepZep61.7原始记录

智能体与工具

τ³-Banking15 项成绩

基础模型成绩Artificial Analysis

银行政策检索与工具执行

%越高越好
τ³-Banking / %
模型与运行配置成绩来源
Muse Spark 1.3 (max)Meta52.37%原始记录
Qwen3.8 MaxAlibaba51.34%原始记录
Grok 4.6 (high)SpaceXAI50.72%原始记录
GLM-5.3 (max)Z AI50.31%原始记录
Qwen3.8 2.4T A95BAlibaba49.07%原始记录
Qwen3.8 27B (xhigh)Alibaba48.04%原始记录
GLM 5.3 Flash (GLM-5.3-Flash)Z AI47.22%原始记录
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic47.22%原始记录
Kimi K3 (max)Kimi45.98%原始记录
Gemini 3.8 Flash (medium)Google45.77%原始记录
Qwen3.8-Flash-NextAlibaba45.36%原始记录
Claude Opus 5 (Adaptive Reasoning, High Effort)Anthropic44.74%原始记录
GPT-5.6 Sol (max)OpenAI44.33%原始记录
GPT-6 Astra (xhigh)OpenAI43.09%原始记录
Grok 4.5 (high)SpaceXAI42.06%原始记录
τ²-Bench15 项成绩

基础模型成绩Artificial Analysis

对话工具使用

%越高越好
τ²-Bench / %
模型与运行配置成绩来源
JT-35B-FlashChina Mobile99.12%原始记录
GLM-5.2 (max)Z AI99.12%原始记录
Step 3.7 FlashStepFun98.54%原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic98.54%原始记录
DeepSeek V4 Pro 0424 (DeepSeek V4 Pro (Reasoning, Max Effort))DeepSeek96.2%原始记录
Qwen3.5 397B A17B (Reasoning)Alibaba95.61%原始记录
Gemini 3.5 Flash (medium)Google95.61%原始记录
Gemini 3.1 Pro PreviewGoogle95.61%原始记录
Qwen3.6 35B A3B (Reasoning)Alibaba95.32%原始记录
DeepSeek V4 Flash 0420 (DeepSeek V4 Flash (Non-reasoning))DeepSeek94.44%原始记录
MiMo-V2.5-ProXiaomi94.15%原始记录
Qwen3.6 27B (Reasoning)Alibaba94.15%原始记录
Mistral Medium 3.5Mistral94.15%原始记录
Qwen3.5 122B A10B (Reasoning)Alibaba93.57%原始记录
MiMo-V2-Flash (Feb 2026)Xiaomi93.27%原始记录
ITBench-AA15 项成绩

基础模型成绩Artificial Analysis

运维故障根因分析

%越高越好
ITBench-AA / %
模型与运行配置成绩来源
GPT-5.6 Sol (max)OpenAI56.21%原始记录
GPT-5.6 Terra (max)OpenAI51.04%原始记录
Kimi K3 (max)Kimi47.69%原始记录
GLM-5.2 (max)Z AI42.66%原始记录
GPT-5.6 Luna (max)OpenAI40.32%原始记录
DeepSeek V4 Pro 0424 (DeepSeek V4 Pro (Reasoning, Max Effort))DeepSeek38.32%原始记录
MiMo-V2.5-ProXiaomi38.23%原始记录
Gemma 4 31B (Reasoning)Google37.29%原始记录
Qwen3.5 397B A17B (Reasoning)Alibaba34.09%原始记录
Gemini 3.1 Pro PreviewGoogle30.33%原始记录
Step 3.7 FlashStepFun30.27%原始记录
Claude 4.5 Haiku (Reasoning)Anthropic27.31%原始记录
Gemma 4 26B A4B (Reasoning)Google23.63%原始记录
gpt-oss-120b (high)OpenAI5.65%原始记录
Nemotron 3 Super 120B A12B (Reasoning)NVIDIA1.13%原始记录
MCP-Atlas public set / Z.ai report7 项成绩

供应商发布的模型成绩Z.ai (model-card report)

使用连接工具完成任务。

%越高越好
MCP-Atlas public set / Z.ai report / %
模型与运行配置成绩来源
GPT-5.2 (xhigh)OpenAI68%原始记录
GLM-5 release reportZ.ai67.8%原始记录
Gemini 3 Pro (GLM-5 release comparison)Google66.6%原始记录
Claude Opus 4.5 (GLM-5 release comparison)Anthropic65.2%原始记录
Kimi K2.5 (GLM-5 release comparison)Moonshot AI63.8%原始记录
DeepSeek V3.2 (GLM-5 release comparison)DeepSeek62.2%原始记录
GLM-4.7 (GLM-5 release comparison)Z.ai52%原始记录

代码与执行

Terminal-Bench v2.115 项成绩

基础模型成绩Artificial Analysis

终端任务完成

%越高越好
Terminal-Bench v2.1 / %
模型与运行配置成绩来源
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic91.39%原始记录
GPT-6 Astra (high)OpenAI89.89%原始记录
GPT-5.6 Sol (xhigh)OpenAI89.51%原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic89.14%原始记录
Grok 4.6 (high)SpaceXAI88.39%原始记录
GPT-5.6 Terra (max)OpenAI88.01%原始记录
Gemini 3.8 Flash (high)Google87.64%原始记录
Qwen3.8-Flash-NextAlibaba86.14%原始记录
Muse Spark 1.3 (max)Meta85.77%原始记录
Gemini 3.7 Flash (high)Google85.77%原始记录
Kimi K3 (max)Kimi85.02%原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic84.64%原始记录
GLM 5.3 Flash (GLM-5.3-Flash)Z AI84.27%原始记录
GLM-5.3 (max)Z AI83.9%原始记录
Qwen3.8 2.4T A95BAlibaba82.02%原始记录
SciCode15 项成绩

基础模型成绩Artificial Analysis

科学编程

%越高越好
SciCode / %
模型与运行配置成绩来源
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic63.08%原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic61%原始记录
Gemini 3.7 Flash (medium)Google59.84%原始记录
Muse Spark 1.3 (xhigh)Meta59.72%原始记录
Kimi K3 (max)Kimi59.49%原始记录
GLM-5.3 (max)Z AI59.03%原始记录
Gemini 3.1 Pro PreviewGoogle58.68%原始记录
GPT-5.6 Sol (high)OpenAI57.75%原始记录
Muse Spark 1.2 (xhigh)Meta57.41%原始记录
Gemini 3.8 Flash (high)Google56.6%原始记录
GPT-6 Astra (max)OpenAI56.48%原始记录
Grok 4.6 (high)SpaceXAI56.48%原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic56.37%原始记录
Grok 4.5 (high)SpaceXAI54.98%原始记录
GPT-5.6 Terra (max)OpenAI54.98%原始记录
LiveCodeBench15 项成绩

基础模型成绩Artificial Analysis

编程问题求解

%越高越好
LiveCodeBench / %
模型与运行配置成绩来源
gpt-oss-120b (high)OpenAI87.83%原始记录
ERNIE 5.0 Thinking PreviewBaidu81.16%原始记录
o3OpenAI80.85%原始记录
Apriel-v1.6-15B-ThinkerServiceNow80.74%原始记录
Qwen3 Next 80B A3B (Reasoning)Alibaba78.41%原始记录
INTELLECT-3Prime Intellect77.67%原始记录
gpt-oss-20b (high)OpenAI77.67%原始记录
K-EXAONE (Reasoning)LG AI Research76.83%原始记录
Doubao Seed CodeByteDance Seed76.61%原始记录
Magistral Medium 1.2Mistral75.03%原始记录
EXAONE 4.0 32B (Reasoning)LG AI Research74.71%原始记录
NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)NVIDIA74.07%原始记录
Llama Nemotron Super 49B v1.5 (Reasoning)NVIDIA73.65%原始记录
Nova 2.0 Pro Preview (medium)Amazon73.02%原始记录
Falcon-H1R-7BTII UAE72.38%原始记录
SciCode / MiniMax report7 项成绩

供应商发布的模型成绩MiniMax (internal evaluation)

科学编程任务。

score越高越好
SciCode / MiniMax report / score
模型与运行配置成绩来源
Gemini 3 Pro (MiniMax internal evaluation)Google56原始记录
Claude Opus 4.6 (MiniMax internal evaluation)Anthropic52原始记录
GPT-5.2 (Thinking; MiniMax internal evaluation)OpenAI52原始记录
Claude Opus 4.5 (MiniMax internal evaluation)Anthropic50原始记录
Claude Sonnet 4.5 (MiniMax internal evaluation)Anthropic45原始记录
MiniMax-M2.5 (MiniMax internal evaluation)MiniMax44.4原始记录
MiniMax-M2.1 (MiniMax internal evaluation)MiniMax41原始记录
Terminal Bench 2.1 / DeepSeek GA report8 项成绩

供应商发布的模型成绩DeepSeek (GA model-card report)

终端中的编程与工具使用任务。

score越高越好
Terminal Bench 2.1 / DeepSeek GA report / score
模型与运行配置成绩来源
Kimi K3 (Code-agent evaluation; harness and reasoning effort not disclosed for this model)Moonshot AI88.3原始记录
Fable-5 (w/ fallback) (Code-agent evaluation; harness and reasoning effort not disclosed for this model; with fallback, details unspecified)Anthropic88原始记录
DeepSeek-V4-Pro-0813 (DeepSeek Harness minimal; max reasoning effort; temperature=1.0; top_p=0.95)DeepSeek87.9原始记录
Opus-4.8 (Code-agent evaluation; harness and reasoning effort not disclosed for this model)Anthropic85原始记录
DeepSeek-V4-Flash-0731 (Code-agent evaluation; harness and reasoning effort not disclosed for this model)DeepSeek82.7原始记录
GLM-5.2 (Code-agent evaluation; harness and reasoning effort not disclosed for this model)Z.ai81原始记录
DeepSeek-V4-Pro (Preview) (Code-agent evaluation; harness and reasoning effort not disclosed for this model)DeepSeek72.1原始记录
DeepSeek-V4-Flash (Preview) (Code-agent evaluation; harness and reasoning effort not disclosed for this model)DeepSeek61.8原始记录
SWE-bench Pro refined set / Qwen3.8-Max report5 项成绩

供应商发布的模型成绩Qwen (model-card report)

Qwen修订的SWE-bench Pro评测集上的代码修复任务。

score越高越好
SWE-bench Pro refined set / Qwen3.8-Max report / score
模型与运行配置成绩来源
Fable 5 (Qwen-refined set; Claude Code; temperature=1.0; top_p=0.95; 256K context; may involve fallback; reasoning effort not disclosed)Anthropic80原始记录
Opus 4.8 (Qwen-refined set; Claude Code; temperature=1.0; top_p=0.95; 256K context; reasoning effort not disclosed)Anthropic69.2原始记录
Qwen3.8-Max service model; Qwen-refined set; Claude Code; temperature=1.0; top_p=0.95; 256K context; reasoning effort not disclosedQwen67.7原始记录
GPT 5.6 Sol (max) (Qwen-refined set; Claude Code; max (table header); temperature=1.0; top_p=0.95; 256K context)OpenAI64.6原始记录
Qwen3.7-Max (Qwen-refined set; Claude Code; temperature=1.0; top_p=0.95; 256K context; reasoning effort not disclosed)Qwen60.6原始记录
SWE-bench Verified / Google G31 (2026-02-19)5 项成绩

供应商发布的模型成绩Google DeepMind (Gemini 3.1 Pro release comparison)

模型及其编程代理环境解决代码仓库问题的比例。

%越高越好
SWE-bench Verified / Google G31 (2026-02-19) / %
模型与运行配置成绩来源
Claude Opus 4.6 (Thinking (Max); Single attempt; provider-specific scaffolding)Anthropic80.8%原始记录
Gemini 3.1 Pro (Thinking (High); Single attempt; provider-specific scaffolding)Google80.6%原始记录
GPT-5.2 (Thinking (xhigh); Single attempt; provider-specific scaffolding)OpenAI80%原始记录
Claude Sonnet 4.6 (Thinking (Max); Single attempt; provider-specific scaffolding)Anthropic79.6%原始记录
Gemini 3 Pro (Thinking (High); Single attempt; provider-specific scaffolding)Google76.2%原始记录
SWE-bench Pro (Public) / Google G31 (2026-02-19)4 项成绩

供应商发布的模型成绩Google DeepMind (Gemini 3.1 Pro release comparison)

SWE-bench Pro 公开版软件工程任务。

%越高越好
SWE-bench Pro (Public) / Google G31 (2026-02-19) / %
模型与运行配置成绩来源
GPT-5.3-Codex (Thinking (xhigh); Public split; single attempt; provider-specific scaffolding)OpenAI56.8%原始记录
GPT-5.2 (Thinking (xhigh); Public split; single attempt; provider-specific scaffolding)OpenAI55.6%原始记录
Gemini 3.1 Pro (Thinking (High); Public split; single attempt; provider-specific scaffolding)Google54.2%原始记录
Gemini 3 Pro (Thinking (High); Public split; single attempt; provider-specific scaffolding)Google43.3%原始记录

模型推理

Intelligence Index v4.215 项成绩

基础模型成绩Artificial Analysis

10项评测综合指数

score越高越好
Intelligence Index v4.2 / score
模型与运行配置成绩来源
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic56.76原始记录
GPT-6 Astra (max)OpenAI54.66原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic54.05原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic53.19原始记录
Muse Spark 1.3 (max)Meta52.95原始记录
GPT-5.6 Sol (max)OpenAI51.26原始记录
Grok 4.6 (high)SpaceXAI50.58原始记录
Kimi K3 (max)Kimi50.23原始记录
GLM-5.3 (max)Z AI48.58原始记录
Gemini 3.8 Flash (high)Google47.07原始记录
Qwen3.8 MaxAlibaba46.91原始记录
Muse Spark 1.2 (xhigh)Meta46.84原始记录
GPT-5.6 Terra (max)OpenAI46.77原始记录
Qwen3.8 2.4T A95BAlibaba46.74原始记录
GLM 5.3 Flash (GLM-5.3-Flash)Z AI46.22原始记录
Humanity's Last Exam15 项成绩

基础模型成绩Artificial Analysis

专家知识与复杂推理

%越高越好
Humanity's Last Exam / %
模型与运行配置成绩来源
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic59.13%原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic55.47%原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic54.87%原始记录
GPT-6 Astra (max)OpenAI54.68%原始记录
GPT-5.6 Sol (max)OpenAI49.49%原始记录
Muse Spark 1.3 (max)Meta49.07%原始记录
Gemini 3.7 Flash (high)Google47.87%原始记录
Gemini 3.8 Flash (high)Google47.82%原始记录
Gemini 3.1 Pro PreviewGoogle47.03%原始记录
Kimi K3 (max)Kimi46.9%原始记录
Muse Spark 1.2 (xhigh)Meta45.46%原始记录
Grok 4.6 (xhigh)SpaceXAI44.07%原始记录
Qwen3.8 MaxAlibaba43.05%原始记录
GPT-5.6 Terra (max)OpenAI42.91%原始记录
Grok 4.5 (high)SpaceXAI42.68%原始记录
GPQA Diamond15 项成绩

基础模型成绩Artificial Analysis

研究生级科学推理

%越高越好
GPQA Diamond / %
模型与运行配置成绩来源
GPT-6 Astra (xhigh)OpenAI96.26%原始记录
Gemini 3.8 Flash (high)Google95.25%原始记录
Grok 4.6 (high)SpaceXAI94.95%原始记录
Gemini 3.7 Flash (high)Google94.55%原始记录
Gemini 3.1 Pro PreviewGoogle94.14%原始记录
Muse Spark 1.3 (xhigh)Meta94.14%原始记录
GPT-5.6 Sol (max)OpenAI94.14%原始记录
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic93.74%原始记录
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic93.74%原始记录
Qwen3.8 2.4T A95BAlibaba93.54%原始记录
Kimi K3 (max)Kimi93.54%原始记录
Grok 4.5 (high)SpaceXAI93.13%原始记录
MiniMax-M3MiniMax92.93%原始记录
Gemini 3.6 Flash (high)Google92.83%原始记录
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek92.83%原始记录
CritPt15 项成绩

基础模型成绩Artificial Analysis

研究级物理问题

%越高越好
CritPt / %
模型与运行配置成绩来源
GPT-5.6 Sol (max)OpenAI32.29%原始记录
GPT-6 Astra (max)OpenAI31.71%原始记录
Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Anthropic31.14%原始记录
GPT-5.5 Pro (xhigh)OpenAI30.57%原始记录
GPT-5.6 Terra (max)OpenAI30%原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic29.14%原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic28.57%原始记录
Muse Spark 1.3 (xhigh)Meta26%原始记录
Gemini 3 Deep ThinkGoogle25.71%原始记录
Kimi K3 (max)Kimi23.43%原始记录
GLM-5.2 (max)Z AI20.86%原始记录
GPT-5.6 Luna (max)OpenAI20.57%原始记录
Qwen3.8 MaxAlibaba20%原始记录
Qwen3.8 2.4T A95BAlibaba20%原始记录
Grok 4.6 (xhigh)SpaceXAI19.71%原始记录
AIME 202515 项成绩

基础模型成绩Artificial Analysis

竞赛数学

%越高越好
AIME 2025 / %
模型与运行配置成绩来源
Nova 2.0 Lite (high)Amazon94.33%原始记录
gpt-oss-120b (high)OpenAI93.44%原始记录
NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)NVIDIA91%原始记录
K-EXAONE (Reasoning)LG AI Research90.33%原始记录
Nova 2.0 Omni (medium)Amazon89.67%原始记录
gpt-oss-20b (high)OpenAI89.33%原始记录
Nova 2.0 Pro Preview (medium)Amazon89%原始记录
o3OpenAI88.33%原始记录
INTELLECT-3Prime Intellect88%原始记录
Apriel-v1.6-15B-ThinkerServiceNow88%原始记录
ERNIE 5.0 Thinking PreviewBaidu85%原始记录
Qwen3 Next 80B A3B (Reasoning)Alibaba84.33%原始记录
Ring-flash-2.0InclusionAI83.67%原始记录
Claude 4.5 Haiku (Reasoning)Anthropic83.67%原始记录
Magistral Medium 1.2Mistral82%原始记录
IFBench15 项成绩

基础模型成绩Artificial Analysis

复杂指令遵循

%越高越好
IFBench / %
模型与运行配置成绩来源
Grok 4.3 (medium)SpaceXAI83.33%原始记录
MiniMax-M3MiniMax82.86%原始记录
Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA81.36%原始记录
Nemotron Cascade 2 30B A3BNVIDIA80.41%原始记录
MiMo-V2.5-ProXiaomi79.86%原始记录
Nova 2.0 Pro Preview (low)Amazon79.59%原始记录
Qwen3.5 397B A17B (Reasoning)Alibaba78.78%原始记录
Qwen3.7 PlusAlibaba77.96%原始记录
Gemini 3.1 Pro PreviewGoogle77.14%原始记录
DeepSeek V4 Pro 0424 (DeepSeek V4 Pro (Reasoning, Max Effort))DeepSeek76.46%原始记录
Qwen3.5 122B A10B (Reasoning)Alibaba75.71%原始记录
Gemma 4 31B (Reasoning)Google75.58%原始记录
GPT-5.3 Codex (xhigh)OpenAI75.37%原始记录
Gemini 3.5 Flash (medium)Google74.56%原始记录
Command A+Cohere73.95%原始记录
LongBench v2 / Moonshot report5 项成绩

供应商发布的模型成绩Moonshot AI (model-card report)

长输入阅读与推理。

score越高越好
LongBench v2 / Moonshot report / score
模型与运行配置成绩来源
Gemini 3 Pro (High Thinking Level; publisher rerun (*))Google68.2原始记录
Claude Opus 4.5 (Extended Thinking; publisher rerun (*))Anthropic64.4原始记录
Kimi K2.5 (Thinking)Moonshot AI61原始记录
DeepSeek V3.2 (Thinking; publisher rerun (*))DeepSeek59.8原始记录
GPT-5.2 (xhigh; publisher rerun (*))OpenAI54.5原始记录
AIME 2026 I / Z.ai report6 项成绩

供应商发布的模型成绩Z.ai (model-card report)

2026 年 AIME 第一场,不是全年合并题集。

score越高越好
AIME 2026 I / Z.ai report / score
模型与运行配置成绩来源
Claude Opus 4.5 (GLM-5 release comparison)Anthropic93.3原始记录
GLM-4.7 (GLM-5 release comparison)Z.ai92.9原始记录
GLM-5 release reportZ.ai92.7原始记录
DeepSeek V3.2 (GLM-5 release comparison)DeepSeek92.7原始记录
Kimi K2.5 (GLM-5 release comparison)Moonshot AI92.5原始记录
Gemini 3 Pro (GLM-5 release comparison)Google90.6原始记录
HMMT November 2025 / Z.ai report7 项成绩

供应商发布的模型成绩Z.ai (model-card report)

2025 年 11 月数学竞赛题。

score越高越好
HMMT November 2025 / Z.ai report / score
模型与运行配置成绩来源
GPT-5.2 (xhigh)OpenAI97.1原始记录
GLM-5 release reportZ.ai96.9原始记录
GLM-4.7 (GLM-5 release comparison)Z.ai93.5原始记录
Gemini 3 Pro (GLM-5 release comparison)Google93原始记录
Claude Opus 4.5 (GLM-5 release comparison)Anthropic91.7原始记录
Kimi K2.5 (GLM-5 release comparison)Moonshot AI91.1原始记录
DeepSeek V3.2 (GLM-5 release comparison)DeepSeek90.2原始记录
AA-LCR / MiniMax report7 项成绩

供应商发布的模型成绩MiniMax (internal evaluation)

基于长输入的推理。

score越高越好
AA-LCR / MiniMax report / score
模型与运行配置成绩来源
Claude Opus 4.5 (MiniMax internal evaluation)Anthropic74原始记录
GPT-5.2 (Thinking; MiniMax internal evaluation)OpenAI73原始记录
Claude Opus 4.6 (MiniMax internal evaluation)Anthropic71原始记录
Gemini 3 Pro (MiniMax internal evaluation)Google71原始记录
MiniMax-M2.5 (MiniMax internal evaluation)MiniMax69.5原始记录
Claude Sonnet 4.5 (MiniMax internal evaluation)Anthropic66原始记录
MiniMax-M2.1 (MiniMax internal evaluation)MiniMax62原始记录
HLE no tools / DeepSeek GA report8 项成绩

供应商发布的模型成绩DeepSeek (GA model-card report)

不使用工具的跨学科推理评测。

score越高越好
HLE no tools / DeepSeek GA report / score
模型与运行配置成绩来源
Fable-5 (w/ fallback) (No tools; HLE reasoning effort and token budget not disclosed; with fallback, details unspecified)Anthropic53.3原始记录
Opus-4.8 (No tools; HLE reasoning effort and token budget not disclosed)Anthropic49.8原始记录
Kimi K3 (No tools; HLE reasoning effort and token budget not disclosed)Moonshot AI43.5原始记录
DeepSeek-V4-Pro-0813 (No tools; HLE reasoning effort and token budget not disclosed)DeepSeek42.7原始记录
GLM-5.2 (No tools; HLE reasoning effort and token budget not disclosed)Z.ai40.5原始记录
DeepSeek-V4-Flash-0731 (No tools; HLE reasoning effort and token budget not disclosed)DeepSeek37.8原始记录
DeepSeek-V4-Pro (Preview) (No tools; HLE reasoning effort and token budget not disclosed)DeepSeek37.7原始记录
DeepSeek-V4-Flash (Preview) (No tools; HLE reasoning effort and token budget not disclosed)DeepSeek34.8原始记录
HLE with tools / DeepSeek GA report8 项成绩

供应商发布的模型成绩DeepSeek (GA model-card report)

使用工具的跨学科推理评测。

score越高越好
HLE with tools / DeepSeek GA report / score
模型与运行配置成绩来源
Fable-5 (w/ fallback) (With tools; specific tools, HLE reasoning effort and token budget not disclosed; with fallback, details unspecified)Anthropic63原始记录
DeepSeek-V4-Pro-0813 (With tools; specific tools, HLE reasoning effort and token budget not disclosed)DeepSeek60原始记录
Opus-4.8 (With tools; specific tools, HLE reasoning effort and token budget not disclosed)Anthropic57.9原始记录
Kimi K3 (With tools; specific tools, HLE reasoning effort and token budget not disclosed)Moonshot AI56原始记录
GLM-5.2 (With tools; specific tools, HLE reasoning effort and token budget not disclosed)Z.ai54.7原始记录
DeepSeek-V4-Flash-0731 (With tools; specific tools, HLE reasoning effort and token budget not disclosed)DeepSeek51.5原始记录
DeepSeek-V4-Pro (Preview) (With tools; specific tools, HLE reasoning effort and token budget not disclosed)DeepSeek48.2原始记录
DeepSeek-V4-Flash (Preview) (With tools; specific tools, HLE reasoning effort and token budget not disclosed)DeepSeek45.1原始记录
HMMT February 2026 / DeepSeek Preview report6 项成绩

供应商发布的模型成绩DeepSeek (Preview technical report)

以单个答案成功率Pass@1衡量的数学竞赛题。

%越高越好
HMMT February 2026 / DeepSeek Preview report / %
模型与运行配置成绩来源
GPT-5.4 xHigh (xHigh; DeepSeek Preview report comparison)OpenAI97.7%原始记录
Opus-4.6 Max (Max; DeepSeek Preview report comparison)Anthropic96.2%原始记录
DS-V4-Pro Max (Preview; Max; temperature=1.0; 384K-token context; distinct rigorous-proof math prompt)DeepSeek95.2%原始记录
Gemini-3.1-Pro High (High; DeepSeek Preview report comparison)Google94.7%原始记录
K2.6 Thinking (Thinking; DeepSeek Preview report comparison)Moonshot AI92.7%原始记录
GLM-5.1 Thinking (Thinking; DeepSeek Preview report comparison)Z.ai89.4%原始记录
LongBench v2 / Qwen3.8-Max report4 项成绩

供应商发布的模型成绩Qwen (model-card report)

长文档与长输入的阅读和推理。

score越高越好
LongBench v2 / Qwen3.8-Max report / score
模型与运行配置成绩来源
Opus 4.8 (LongBench v2; tools, reasoning effort and token budget not disclosed)Anthropic69.1原始记录
GPT 5.6 Sol (max) (max (table header); LongBench v2 tools and token budget not disclosed)OpenAI67.1原始记录
Qwen3.8-Max service model; LongBench v2 tools, reasoning effort and token budget not disclosedQwen66.3原始记录
Qwen3.7-Max (LongBench v2; tools, reasoning effort and token budget not disclosed)Qwen65.3原始记录
GPQA Diamond / Google G31 (2026-02-19)5 项成绩

供应商发布的模型成绩Google DeepMind (Gemini 3.1 Pro release comparison)

不使用工具的科学知识与推理评测。

%越高越好
GPQA Diamond / Google G31 (2026-02-19) / %
模型与运行配置成绩来源
Gemini 3.1 Pro (Thinking (High); No tools)Google94.3%原始记录
GPT-5.2 (Thinking (xhigh); No tools)OpenAI92.4%原始记录
Gemini 3 Pro (Thinking (High); No tools)Google91.9%原始记录
Claude Opus 4.6 (Thinking (Max); No tools)Anthropic91.3%原始记录
Claude Sonnet 4.6 (Thinking (Max); No tools)Anthropic89.9%原始记录
HLE, no tools / Google G31 (2026-02-19)5 项成绩

供应商发布的模型成绩Google DeepMind (Gemini 3.1 Pro release comparison)

不使用工具完成 HLE 全部文本与多模态题目。

%越高越好
HLE, no tools / Google G31 (2026-02-19) / %
模型与运行配置成绩来源
Gemini 3.1 Pro (Thinking (High); No tools)Google44.4%原始记录
Claude Opus 4.6 (Thinking (Max); No tools)Anthropic40%原始记录
Gemini 3 Pro (Thinking (High); No tools)Google37.5%原始记录
GPT-5.2 (Thinking (xhigh); No tools)OpenAI34.5%原始记录
Claude Sonnet 4.6 (Thinking (Max); No tools)Anthropic33.2%原始记录
HLE, search and code / Google G31 (2026-02-19)5 项成绩

供应商发布的模型成绩Google DeepMind (Gemini 3.1 Pro release comparison)

使用搜索与代码执行的 HLE 推理评测。

%越高越好
HLE, search and code / Google G31 (2026-02-19) / %
模型与运行配置成绩来源
Claude Opus 4.6 (Thinking (Max); Search and code; provider-specific tools)Anthropic53.1%原始记录
Gemini 3.1 Pro (Thinking (High); Search and code; provider-specific tools)Google51.4%原始记录
Claude Sonnet 4.6 (Thinking (Max); Search and code; provider-specific tools)Anthropic49%原始记录
Gemini 3 Pro (Thinking (High); Search and code; provider-specific tools)Google45.8%原始记录
GPT-5.2 (Thinking (xhigh); Search and code; provider-specific tools)OpenAI45.5%原始记录
ARC-AGI-2 / Google G31 (2026-02-19)5 项成绩

供应商发布的模型成绩Google DeepMind (Gemini 3.1 Pro release comparison)

从网格示例中推断陌生规律。

%越高越好
ARC-AGI-2 / Google G31 (2026-02-19) / %
模型与运行配置成绩来源
Gemini 3.1 Pro (Thinking (High); ARC Prize Verified; semi-private set)Google77.1%原始记录
Claude Opus 4.6 (Thinking (Max); ARC Prize Verified; semi-private set)Anthropic68.8%原始记录
Claude Sonnet 4.6 (Thinking (Max); ARC Prize Verified; semi-private set)Anthropic58.3%原始记录
GPT-5.2 (Thinking (xhigh); ARC Prize Verified; semi-private set)OpenAI52.9%原始记录
Gemini 3 Pro (Thinking (High); ARC Prize Verified; semi-private set)Google31.1%原始记录
AIME 2025, no tools / Google G3F (2025-12-17)7 项成绩

供应商发布的模型成绩Google DeepMind (Gemini 3 Flash release comparison)

解答 2025 年 AIME 数学竞赛题。

%越高越好
AIME 2025, no tools / Google G3F (2025-12-17) / %
模型与运行配置成绩来源
GPT-5.2 (Extra high; No tools)OpenAI100%原始记录
Gemini 3 Flash (Thinking; exact level unspecified; No tools)Google95.2%原始记录
Gemini 3 Pro (Thinking; exact level unspecified; No tools)Google95%原始记录
Grok 4.1 Fast (Reasoning; No tools)xAI91.9%原始记录
Gemini 2.5 Pro (Thinking; exact level unspecified; No tools)Google88%原始记录
Claude Sonnet 4.5 (Thinking; high preferred, otherwise best available reported setting; No tools)Anthropic87%原始记录
Gemini 2.5 Flash (Thinking; exact level unspecified; No tools)Google72%原始记录
AIME 2025, code execution / Google G3F (2025-12-17)4 项成绩

供应商发布的模型成绩Google DeepMind (Gemini 3 Flash release comparison)

解答 2025 年 AIME 数学竞赛题。

%越高越好
AIME 2025, code execution / Google G3F (2025-12-17) / %
模型与运行配置成绩来源
Gemini 3 Pro (Thinking; exact level unspecified; Code execution)Google100%原始记录
Claude Sonnet 4.5 (Thinking; high preferred, otherwise best available reported setting; Code execution)Anthropic100%原始记录
Gemini 3 Flash (Thinking; exact level unspecified; Code execution)Google99.7%原始记录
Gemini 2.5 Flash (Thinking; exact level unspecified; Code execution)Google75.7%原始记录
HLE, no tools / Anthropic A51 (2026-09-01)3 项成绩

供应商发布的模型成绩Anthropic (Fable 5.1 release comparison)

不使用工具完成原版 HLE 的 2,500 道多模态题目。

%越高越好
HLE, no tools / Anthropic A51 (2026-09-01) / %
模型与运行配置成绩来源
Claude Fable 5.1 (Adaptive thinking (auto); max effort; default sampling; five-trial mean; total token cap 1M; no compaction; production safeguards enabled; No tools)Anthropic60.9%原始记录
Claude Fable 5 (Detailed reasoning setting and budget unverified; production safeguards enabled; No tools)Anthropic57.8%原始记录
Claude Opus 5 (Detailed reasoning setting and budget unverified; No tools)Anthropic56.6%原始记录
HLE, tools / Anthropic A51 (2026-09-01)3 项成绩

供应商发布的模型成绩Anthropic (Fable 5.1 release comparison)

使用搜索、网页检索及代码工具的原版 HLE 推理评测。

%越高越好
HLE, tools / Anthropic A51 (2026-09-01) / %
模型与运行配置成绩来源
Claude Fable 5.1 (Adaptive thinking (auto); max effort; default sampling; five-trial mean; total token cap 1M; no compaction; production safeguards enabled; Search; restricted fetch; programmatic tool calling; code execution)Anthropic65%原始记录
Claude Fable 5 (Detailed reasoning setting and budget unverified; production safeguards enabled; Search; restricted fetch; programmatic tool calling; code execution)Anthropic63.8%原始记录
Claude Opus 5 (Detailed reasoning setting and budget unverified; Search; restricted fetch; programmatic tool calling; code execution)Anthropic63.6%原始记录
ARC-AGI-2 / Anthropic A51 (2026-09-01)4 项成绩

供应商发布的模型成绩Anthropic (Fable 5.1 release comparison)

九月发布比较中的陌生网格规律推理评测。

%越高越好
ARC-AGI-2 / Anthropic A51 (2026-09-01) / %
模型与运行配置成绩来源
GPT-5.6 Sol (Exact reasoning setting and token budget unverified)OpenAI92.5%原始记录
Claude Opus 5 (Exact reasoning setting and token budget unverified)Anthropic90.42%原始记录
Claude Fable 5.1 / Claude Mythos 5.1 (Fable 5.1: max effort; semi-private set; ARC Prize Verified)Anthropic90%原始记录
Claude Fable 5 / Claude Mythos 5 (Exact reasoning setting and token budget unverified)Anthropic89.2%原始记录
HLE-Verified / Google G38 (2026-09-02)6 项成绩

供应商发布的模型成绩Google DeepMind (Gemini 3.8 Flash release comparison)

采用经验证及修订的 1,811 道题进行专家推理评测。

%越高越好
HLE-Verified / Google G38 (2026-09-02) / %
模型与运行配置成绩来源
Gemini 3.8 Flash (Default sampling; exact thinking level unspecified; HLE-Verified; 1811 items; tool use unspecified)Google54.9%原始记录
GPT-5.6 Sol (Exact reasoning level unspecified; HLE-Verified; 1811 items; tool use unspecified)OpenAI54.5%原始记录
Claude Opus 5 (Exact thinking level unspecified; HLE-Verified; 1811 items; tool use unspecified)Anthropic54.4%原始记录
Gemini 3.7 Flash (Exact thinking level unspecified; HLE-Verified; 1811 items; tool use unspecified)Google53.6%原始记录
GPT-5.6 Terra (Maximum reasoning preferred; otherwise best available reported setting; HLE-Verified; 1811 items; tool use unspecified)OpenAI51.1%原始记录
Claude Sonnet 5 (Maximum thinking preferred; otherwise best available reported setting; HLE-Verified; 1811 items; tool use unspecified)Anthropic31%原始记录

文档与上下文理解

AA-LCR v1.115 项成绩

基础模型成绩Artificial Analysis

长上下文推理

%越高越好
AA-LCR v1.1 / %
模型与运行配置成绩来源
Kimi K3 (max)Kimi88.67%原始记录
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic85.33%原始记录
Muse Spark 1.3 (max)Meta84.33%原始记录
Gemini 3.8 Flash (medium)Google84%原始记录
GPT-5.6 Sol (max)OpenAI84%原始记录
GPT-5.6 Luna (max)OpenAI83.67%原始记录
Muse Glimmer (high)Meta83.33%原始记录
GPT-5.3 Codex (xhigh)OpenAI83.33%原始记录
MiniMax-M3MiniMax83%原始记录
GPT-5.6 Terra (max)OpenAI83%原始记录
Gemini 3.7 Flash (medium)Google83%原始记录
Agnes 2.5 Pro BetaSapiens AI83%原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic82.33%原始记录
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic82%原始记录
Qwen3.8 27B (xhigh)Alibaba82%原始记录
MMMU-Pro15 项成绩

基础模型成绩Artificial Analysis

视觉推理

%越高越好
MMMU-Pro / %
模型与运行配置成绩来源
GPT-6 Astra (max)OpenAI86.88%原始记录
Gemini 3.8 Flash (high)Google85.61%原始记录
Gemini 3.7 Flash (high)Google85.49%原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic84.74%原始记录
Gemini 3.5 Flash (medium)Google83.87%原始记录
GPT-5.6 Sol (max)OpenAI83.41%原始记录
Gemini 3.6 Flash (high)Google83.24%原始记录
Gemini 3.1 Pro PreviewGoogle82.43%原始记录
Qwen3.8 MaxAlibaba82.31%原始记录
Muse Spark 1.3 (xhigh)Meta82.02%原始记录
GPT-5.6 Terra (max)OpenAI80.69%原始记录
Kimi K3 (max)Kimi80.52%原始记录
Qwen3.7 PlusAlibaba80.46%原始记录
Grok 4.5 (high)SpaceXAI80.4%原始记录
Qwen3.8-Flash-NextAlibaba79.77%原始记录
AA-Omniscience / Index15 项成绩

基础模型成绩Artificial Analysis

知识可靠性

score越高越好
AA-Omniscience / Index / score
模型与运行配置成绩来源
GPT-6 Astra (high)OpenAI43.73原始记录
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic43.45原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic43.3原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic37.07原始记录
Gemini 3.1 Pro PreviewGoogle31.88原始记录
Grok 4.6 (high)SpaceXAI30.48原始记录
Gemini 3.8 Flash (high)Google29.55原始记录
Muse Spark 1.2 (xhigh)Meta27.2原始记录
Gemini 3.7 Flash (high)Google26.48原始记录
Grok 4.5 (high)SpaceXAI25.32原始记录
Muse Spark 1.3 (max)Meta24.93原始记录
Gemini 3.6 Flash (high)Google22.13原始记录
GPT-5.6 Sol (max)OpenAI21.97原始记录
Gemini 3.5 Flash (medium)Google20.82原始记录
Kimi K3 (max)Kimi19.7原始记录
AA-Omniscience / Accuracy15 项成绩

基础模型成绩Artificial Analysis

知识问题正确率

%越高越好
AA-Omniscience / Accuracy / %
模型与运行配置成绩来源
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic67.23%原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic65.35%原始记录
GPT-6 Astra (max)OpenAI62.6%原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic60.87%原始记录
GPT-5.6 Sol (max)OpenAI59.4%原始记录
Gemini 3.7 Flash (high)Google55.32%原始记录
Gemini 3.1 Pro PreviewGoogle54.85%原始记录
Gemini 3.8 Flash (high)Google54.6%原始记录
GPT-5.3 Codex (xhigh)OpenAI52.88%原始记录
Grok 4.5 (high)SpaceXAI51.55%原始记录
Gemini 3.5 Flash (medium)Google51.05%原始记录
Gemini 3.6 Flash (high)Google49.97%原始记录
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek49.1%原始记录
Grok 4.6 (high)SpaceXAI48.23%原始记录
Kimi K3 (max)Kimi47.58%原始记录
AA-Omniscience / Hallucination15 项成绩

基础模型成绩Artificial Analysis

幻觉率,越低越好

%越低越好
AA-Omniscience / Hallucination / %
模型与运行配置成绩来源
MiniCPM5-1B (Non-reasoning)OpenBMB0.9%原始记录
G9v3-3BAI9Stars11.66%原始记录
G9v3-39A5BAI9Stars13.04%原始记录
Command A+Cohere14.16%原始记录
LFM2.5-2.6BLiquid AI16%原始记录
Grok 4.3 (medium)SpaceXAI16.94%原始记录
Qwen3.8 27B (Non-reasoning)Alibaba18.09%原始记录
MiniMax-M3MiniMax18.43%原始记录
Quasar 438B (max, based on GLM-5.2)Multiverse Computing21.44%原始记录
K-EXAONE 2.0 0803 (K-EXAONE 2.0)LG AI Research22.6%原始记录
Grok 4.6 (medium)SpaceXAI24%原始记录
Solar Pro 4Upstage24.4%原始记录
MiMo-V2.5-ProXiaomi24.7%原始记录
Solar Open2 250BUpstage25.38%原始记录
Granite 4.2 30BIBM25.58%原始记录
MathVision / Moonshot report5 项成绩

供应商发布的模型成绩Moonshot AI (model-card report)

包含图像的数学题。

score越高越好
MathVision / Moonshot report / score
模型与运行配置成绩来源
Gemini 3 Pro (High Thinking Level; publisher rerun (*))Google86.1原始记录
Kimi K2.5 (Thinking)Moonshot AI84.2原始记录
GPT-5.2 (xhigh)OpenAI83原始记录
Claude Opus 4.5 (Extended Thinking; publisher rerun (*))Anthropic77.1原始记录
Qwen3-VL-235B-A22B (Thinking)Qwen74.6原始记录
OCRBench / Moonshot report5 项成绩

供应商发布的模型成绩Moonshot AI (model-card report)

图像中的文字识别。

score越高越好
OCRBench / Moonshot report / score
模型与运行配置成绩来源
Kimi K2.5 (Thinking)Moonshot AI92.3原始记录
Gemini 3 Pro (High Thinking Level; publisher rerun (*))Google90.3原始记录
Qwen3-VL-235B-A22B (Thinking)Qwen87.5原始记录
Claude Opus 4.5 (Extended Thinking; publisher rerun (*))Anthropic86.5原始记录
GPT-5.2 (xhigh; publisher rerun (*))OpenAI80.7原始记录
OmniDocBench 1.5 / Moonshot report5 项成绩

供应商发布的模型成绩Moonshot AI (model-card report)

文档阅读。公布分数为归一化编辑距离的补数乘以 100。

score越高越好
OmniDocBench 1.5 / Moonshot report / score
模型与运行配置成绩来源
Kimi K2.5 (Thinking)Moonshot AI88.8原始记录
Gemini 3 Pro (High Thinking Level)Google88.5原始记录
Claude Opus 4.5 (Extended Thinking; publisher rerun (*))Anthropic87.7原始记录
GPT-5.2 (xhigh)OpenAI85.7原始记录
Qwen3-VL-235B-A22B (Thinking; publisher rerun (*))Qwen82原始记录
VideoMMMU / Moonshot report5 项成绩

供应商发布的模型成绩Moonshot AI (model-card report)

跨学科视频理解。

score越高越好
VideoMMMU / Moonshot report / score
模型与运行配置成绩来源
Gemini 3 Pro (High Thinking Level)Google87.6原始记录
Kimi K2.5 (Thinking)Moonshot AI86.6原始记录
GPT-5.2 (xhigh)OpenAI85.9原始记录
Claude Opus 4.5 (Extended Thinking; publisher rerun (*))Anthropic84.4原始记录
Qwen3-VL-235B-A22B (Thinking)Qwen80原始记录
MotionBench / Moonshot report4 项成绩

供应商发布的模型成绩Moonshot AI (model-card report)

理解视频中的运动。

score越高越好
MotionBench / Moonshot report / score
模型与运行配置成绩来源
Kimi K2.5 (Thinking)Moonshot AI70.4原始记录
Gemini 3 Pro (High Thinking Level)Google70.3原始记录
GPT-5.2 (xhigh)OpenAI64.8原始记录
Claude Opus 4.5 (Extended Thinking)Anthropic60.3原始记录
IFBench / MiniMax report7 项成绩

供应商发布的模型成绩MiniMax (internal evaluation)

遵循详细指令。

score越高越好
IFBench / MiniMax report / score
模型与运行配置成绩来源
GPT-5.2 (Thinking; MiniMax internal evaluation)OpenAI75原始记录
MiniMax-M2.5 (MiniMax internal evaluation)MiniMax70原始记录
MiniMax-M2.1 (MiniMax internal evaluation)MiniMax70原始记录
Gemini 3 Pro (MiniMax internal evaluation)Google70原始记录
Claude Opus 4.5 (MiniMax internal evaluation)Anthropic58原始记录
Claude Sonnet 4.5 (MiniMax internal evaluation)Anthropic57原始记录
Claude Opus 4.6 (MiniMax internal evaluation)Anthropic53原始记录
MMMLU / DeepSeek Base report3 项成绩

供应商发布的模型成绩DeepSeek (Base technical report)

提供5个示例,以答案完全匹配率衡量的多语言知识。

%越高越好
MMMLU / DeepSeek Base report / %
模型与运行配置成绩来源
DeepSeek-V4-Pro-Base (Base; MMMLU EM; 5-shot; shared internal evaluation setup; sampling and tools not disclosed)DeepSeek90.3%原始记录
DeepSeek-V4-Flash-Base (Base; MMMLU EM; 5-shot; shared internal evaluation setup; sampling and tools not disclosed)DeepSeek88.8%原始记录
DeepSeek-V3.2-Base (Base; MMMLU EM; 5-shot; shared internal evaluation setup; sampling and tools not disclosed)DeepSeek87.9%原始记录
OmniDocBench 1.5 / Qwen3.8-27B report5 项成绩

供应商发布的模型成绩Qwen (model-card report)

使用提供方公布分数比较文档图像读取能力。

score越高越好
OmniDocBench 1.5 / Qwen3.8-27B report / score
模型与运行配置成绩来源
Qwen3.7-Plus (OmniDocBench 1.5; tools, reasoning effort and image settings not disclosed)Qwen91.4原始记录
Qwen3.8-27B (OmniDocBench 1.5; tools, reasoning effort and image settings not disclosed)Qwen91.1原始记录
Qwen3.6-27B (OmniDocBench 1.5; tools, reasoning effort and image settings not disclosed)Qwen89.4原始记录
Opus4.6 Max (Max (table header); OmniDocBench 1.5 tools and image settings not disclosed)Anthropic86.6原始记录
Muse Glimmer-30B (OmniDocBench 1.5; tools, reasoning effort and image settings not disclosed)Meta75.8原始记录
MMMU-Pro / Google G31 (2026-02-19)5 项成绩

供应商发布的模型成绩Google DeepMind (Gemini 3.1 Pro release comparison)

不使用工具理解多模态信息并进行推理。

%越高越好
MMMU-Pro / Google G31 (2026-02-19) / %
模型与运行配置成绩来源
Gemini 3 Pro (Thinking (High); No tools; Standard (10 options) and Vision average)Google81%原始记录
Gemini 3.1 Pro (Thinking (High); No tools; Standard (10 options) and Vision average)Google80.5%原始记录
GPT-5.2 (Thinking (xhigh); No tools; Standard (10 options) and Vision average)OpenAI79.5%原始记录
Claude Sonnet 4.6 (Thinking (Max); No tools; Standard (10 options) and Vision average)Anthropic74.5%原始记录
Claude Opus 4.6 (Thinking (Max); No tools; Standard (10 options) and Vision average)Anthropic73.9%原始记录
MRCR v2, 128k / Google G31 (2026-02-19)5 项成绩

供应商发布的模型成绩Google DeepMind (Gemini 3.1 Pro release comparison)

长上下文中检索 8 项指定信息,采用截至 128k 令牌的累计平均。

%越高越好
MRCR v2, 128k / Google G31 (2026-02-19) / %
模型与运行配置成绩来源
Gemini 3.1 Pro (Thinking (High); MRCR v2; 8 needles; 128k cumulative average)Google84.9%原始记录
Claude Sonnet 4.6 (Thinking (Max); MRCR v2; 8 needles; 128k cumulative average)Anthropic84.9%原始记录
Claude Opus 4.6 (Thinking (Max); MRCR v2; 8 needles; 128k cumulative average)Anthropic84%原始记录
GPT-5.2 (Thinking (xhigh); MRCR v2; 8 needles; 128k cumulative average)OpenAI83.8%原始记录
Gemini 3 Pro (Thinking (High); MRCR v2; 8 needles; 128k cumulative average)Google77%原始记录
MRCR v2, 1M / Google G31 (2026-02-19)2 项成绩

供应商发布的模型成绩Google DeepMind (Gemini 3.1 Pro release comparison)

在 1M 令牌长度处检索 8 项指定信息的评测。

%越高越好
  1. Gemini 3.1 Pro26.3%
  2. Gemini 3 Pro26.3%
MRCR v2, 1M / Google G31 (2026-02-19) / %
模型与运行配置成绩来源
Gemini 3.1 Pro (Thinking (High); MRCR v2; 8 needles; 1M pointwise)Google26.3%原始记录
Gemini 3 Pro (Thinking (High); MRCR v2; 8 needles; 1M pointwise)Google26.3%原始记录
CharXiv Reasoning / Google G38 (2026-09-02)6 项成绩

供应商发布的模型成绩Google DeepMind (Gemini 3.8 Flash release comparison)

不使用工具综合复杂图表中的信息。

%越高越好
CharXiv Reasoning / Google G38 (2026-09-02) / %
模型与运行配置成绩来源
Gemini 3.8 Flash (Default sampling; exact thinking level unspecified; No tools)Google86.2%原始记录
GPT-5.6 Terra (Maximum reasoning preferred; otherwise best available reported setting; No tools)OpenAI85.9%原始记录
GPT-5.6 Sol (Exact reasoning level unspecified; No tools)OpenAI85.8%原始记录
Gemini 3.7 Flash (Exact thinking level unspecified; No tools)Google84.5%原始记录
Claude Opus 5 (Exact thinking level unspecified; No tools)Anthropic83.7%原始记录
Claude Sonnet 5 (Maximum thinking preferred; otherwise best available reported setting; No tools)Anthropic70.1%原始记录

检索模型

MTEB / OpenAI embeddings report3 项成绩

检索模型成绩OpenAI (developer documentation)

文本嵌入模型评测,与生成式模型推理分开。

%越高越好
MTEB / OpenAI embeddings report / %
模型与运行配置成绩来源
text-embedding-3-large (Published MTEB result; evaluation dimensions unspecified)OpenAI64.6%原始记录
text-embedding-3-small (Published MTEB result; evaluation dimensions unspecified)OpenAI62.3%原始记录
text-embedding-ada-002 (Published MTEB result; evaluation dimensions unspecified)OpenAI61%原始记录

时间与成本

Intelligence Index / Cost per task15 项成绩

基础模型成绩Artificial Analysis

评测任务加权平均成本

USD越低越好
Intelligence Index / Cost per task / USD
模型与运行配置成绩来源
GLM 5.3 Flash (GLM-5.3-Flash)Z AI0.18原始记录
Muse Spark 1.2 (xhigh)Meta0.55原始记录
Gemini 3.8 Flash (high)Google0.74原始记录
GPT-5.6 Terra (max)OpenAI0.81原始记录
Muse Spark 1.3 (max)Meta0.96原始记录
Qwen3.8 2.4T A95BAlibaba1.1原始记录
Qwen3.8 MaxAlibaba1.19原始记录
GPT-5.6 Sol (max)OpenAI1.25原始记录
Grok 4.6 (high)SpaceXAI1.25原始记录
GLM-5.3 (max)Z AI1.26原始记录
Kimi K3 (max)Kimi1.58原始记录
GPT-6 Astra (max)OpenAI2.57原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic4.21原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic5.62原始记录
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic6.12原始记录
Output speed15 项成绩

基础模型成绩Artificial Analysis

输出生成速度

tokens/s越高越好
Output speed / tokens/s
模型与运行配置成绩来源
Gemini 3.8 Flash (high)Google280.76原始记录
Muse Spark 1.2 (xhigh)Meta269.24原始记录
Muse Spark 1.3 (max)Meta233.15原始记录
GPT-5.6 Terra (max)OpenAI113.93原始记录
GLM-5.3 (max)Z AI83.31原始记录
GPT-5.6 Sol (max)OpenAI75.78原始记录
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic68.22原始记录
GPT-6 Astra (max)OpenAI64.26原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic62.97原始记录
Grok 4.6 (high)SpaceXAI59.89原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic55.95原始记录
GLM 5.3 Flash (GLM-5.3-Flash)Z AI50.63原始记录
Kimi K3 (max)Kimi41.61原始记录
Qwen3.8 2.4T A95BAlibaba40.23原始记录
Qwen3.8 MaxAlibaba40.15原始记录
API input price15 项成绩

基础模型成绩Artificial Analysis

每百万输入令牌价格

USD / 1M越低越好
API input price / USD / 1M
模型与运行配置成绩来源
GLM 5.3 Flash (GLM-5.3-Flash)Z AI0.15原始记录
Gemini 3.8 Flash (high)Google0.75原始记录
Muse Spark 1.3 (max)Meta1.25原始记录
Muse Spark 1.2 (xhigh)Meta1.25原始记录
GLM-5.3 (max)Z AI1.4原始记录
Grok 4.6 (high)SpaceXAI2原始记录
Qwen3.8 MaxAlibaba2原始记录
GPT-5.6 Terra (max)OpenAI2原始记录
Qwen3.8 2.4T A95BAlibaba2原始记录
Kimi K3 (max)Kimi3原始记录
GPT-5.6 Sol (max)OpenAI4原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic5原始记录
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic10原始记录
GPT-6 Astra (max)OpenAI10原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic10原始记录
API output price15 项成绩

基础模型成绩Artificial Analysis

每百万输出令牌价格

USD / 1M越低越好
API output price / USD / 1M
模型与运行配置成绩来源
GLM 5.3 Flash (GLM-5.3-Flash)Z AI0.5原始记录
Gemini 3.8 Flash (high)Google3.75原始记录
Muse Spark 1.3 (max)Meta4.25原始记录
Muse Spark 1.2 (xhigh)Meta4.25原始记录
GLM-5.3 (max)Z AI4.4原始记录
Grok 4.6 (high)SpaceXAI6原始记录
Qwen3.8 MaxAlibaba6原始记录
Qwen3.8 2.4T A95BAlibaba6原始记录
GPT-5.6 Terra (max)OpenAI12原始记录
Kimi K3 (max)Kimi15原始记录
GPT-5.6 Sol (max)OpenAI20原始记录
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic25原始记录
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic50原始记录
GPT-6 Astra (max)OpenAI50原始记录
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic50原始记录

正文来源与延伸阅读

  1. Anthropic. Demystifying evals for AI agents
  2. OpenAI. GDPval
  3. OpenAI. BrowseComp
  4. Parallel. Deep Research price-performance
  5. LangChain. Improving Deep Agents with harness engineering
  6. LangChain. Deep Agents source code
  7. DeepResearch Bench II. Evaluation code
  8. ResearchRubrics. Evaluation code
  9. Agents’ Last Exam. Tasks and evaluation