HDATF RESEARCH / 벤치마크 보고서

업무를 끝내는 AI, 어떻게 검증할 것인가

공개 성적에서 읽을 수 있는 것과 직접 실행해 확인해야 할 것. 모델 선택부터 조사, 문서 작성, 결과 검증까지 살펴봅니다.

먼저 읽을 결론

공개 벤치마크는 모델을 고르고 시험을 설계하는 근거입니다. 업무에 쓸 준비가 됐는지는 실제 자료와 도구를 연결한 뒤, 요청한 결과물을 끝까지 만드는지 확인해야 알 수 있습니다.

이 보고서는 쓸 수 있는 결과물, 확인 가능한 근거, 끝까지 이어지는 실행을 기준으로 공개 성적을 읽습니다. HDATF가 이 평가들에서 직접 측정해 공개한 성적은 아직 없습니다. 마지막 장에서는 그 차이를 확인하기 위한 평가 계획을 설명합니다.

73 평가와 지표684 보관된 성적

01 / 성적을 읽는 기준

성적을 읽기 전에, 맡길 일을 정해야 합니다.

벤치마크는 정해진 과제와 채점 기준으로 능력을 비교하는 시험입니다. 추론 문제를 푸는 시험, 문서를 비교하는 시험, 도구로 작업을 수행하는 시험은 확인하는 능력이 다릅니다. 점수를 읽을 때도 무엇을 시험했는지부터 봐야 합니다.

보고서 업무라면 파일이 열리고, 주장의 근거를 찾을 수 있고, 요청한 내용이 들어 있어야 합니다. 데이터 분석이라면 제공한 자료로 결과를 다시 만들 수 있어야 합니다. 일반 모델 순위만으로는 이 조건을 확인할 수 없습니다.

공개 성적은 후보를 좁히고 어떤 실패를 시험할지 정하는 데 씁니다. LabChin에 적용할 방법을 고를 때는 우리 자료와 권한, 도구, 결과 검사 기준을 연결한 과제가 필요합니다. 이 보고서는 결과물, 근거, 실행 방식의 순서로 그 판단 기준을 살펴봅니다.[1]

02 / 결과물의 품질

결과물은 요청한 업무를 기준으로 읽어야 합니다.

GDPval은 전문 업무의 산출물을 비교합니다. 최초 Gold 연구에서는 44개 직종의 과제 220개를 다뤘습니다. 전문가는 작성자를 모르는 상태에서 결과물을 비교했습니다. 아래 그림은 그 연구에서 보관한 세 결과입니다.

승리와 동률 비율은 해당 산출물 비교의 결과입니다. 회사 업무를 AI에 맡길 수 있는 비율로 바꿔 읽어서는 안 됩니다. 과제의 범위와 전문가 판단, 사용한 모델 구성이 해석의 범위를 정합니다.[2]

그림 01

GDPval Gold

전문가 산출물 대비 승리와 동률 비율 (%)

%높을수록 좋음
  1. Claude Opus 4.147.6%
  2. GPT-5 (high)38.8%
  3. o3 (high)34.1%
GDPval Gold / %
모델과 실행 구성성적출처
Claude Opus 4.1Anthropic47.6%원본 기록
GPT-5 (high)OpenAI38.8%원본 기록
o3 (high)OpenAI34.1%원본 기록

44개 직종의 Gold 과제 220개. 전문가가 작성자를 가린 채 산출물을 비교한 최초 연구 결과입니다. 현재 모델 순위가 아닙니다.

출처: OpenAI2025-09원본 기록

HDATF가 확인할 것은 파일을 만든 뒤 사람이 얼마나 더 일해야 하는가입니다. 빠진 요구사항, 근거 없는 수치, 사람이 고친 내용을 함께 남겨야 합니다. 읽기 좋은 배치는 품질의 한 부분입니다. 잘못된 근거나 쓸 수 없는 파일을 대신해 주지는 못합니다.

03 / 조사와 근거

답을 찾는 능력과 근거로 설명하는 능력을 함께 봅니다.

BrowseComp는 찾기 어려운 사실을 알아내는 능력을 시험합니다. Parallel의 공개 비교는 전체 1,266문항 중 무작위 100문항을 사용했습니다. 정답률과 요청 비용을 함께 발표했습니다. 아래 결과는 제공사가 일부 문항으로 수행한 실험입니다.[3][4]

그림 02

BrowseComp / Parallel study

정답률 (%)

%높을수록 좋음
BrowseComp / Parallel study / %
모델과 실행 구성성적출처
Parallel Ultra8xParallel58%원본 기록
Parallel UltraParallel45%원본 기록
GPT-5 (high)OpenAI38%원본 기록

전체 1,266문항 중 무작위 100문항. Parallel이 2025년 8월 11일부터 29일까지 측정했습니다. 요청당 비용은 Ultra8x 2.40달러, Ultra 0.30달러, GPT-5 0.488달러입니다. 구성과 예산이 다릅니다.

출처: Parallel2025-09-09원본 기록

Ultra8x는 요청당 2.40달러에서 정답률 58%, Ultra는 0.30달러에서 45%를 보고했습니다. 투입한 예산이 다릅니다. 정확도와 비용을 함께 볼 수 있지만, 같은 비용에서의 우열이나 기반 모델 하나의 효과를 증명하지는 않습니다.

조사 보고서는 더 확인할 것이 있습니다. DeepResearch Bench II와 ResearchRubrics는 명시한 기준으로 보고서 내용을 평가합니다. LabChin에서는 주요 주장에 쓰인 원문 구절, 상충하는 자료, 답하지 못한 질문을 함께 남기는 방식으로 검증하려고 합니다.

04 / 실행과 메모리

같은 모델도 실행 방식에 따라 결과가 달라집니다.

하네스는 모델에 도구와 참고 자료를 주고 작업 순서를 제어하는 실행 소프트웨어입니다. LangChain은 GPT-5.2-Codex 모델을 유지한 채 이 실행 방식을 바꿨습니다. 89개 과제의 Terminal-Bench 2.0 성적은 52.8%에서 66.5%로 달라졌습니다.[5][6]

그림 03

Terminal-Bench 2.0

과제 성공률 (%)

%높을수록 좋음
Terminal-Bench 2.0 / %
모델과 실행 구성성적출처
Deep Agents / improvedLangChain66.5%원본 기록
Deep Agents / baselineLangChain52.8%원본 기록

89개 과제. GPT-5.2-Codex 모델을 유지하고 프롬프트, 도구, 실행 제어를 바꾼 LangChain의 실험입니다. Harbor와 Daytona에서 실행했습니다. 발표된 차이는 13.7%p입니다.

출처: LangChain2026-02-17원본 기록

실험에서는 지시문, 도구, 실행 제어를 함께 바꿨습니다. 종료 전 검사와 실행 환경 안내도 포함됐습니다. 13.7%p 차이는 이 변경을 묶어 시험한 결과입니다. 개별 기능의 기여도나 HDATF에서 얻을 개선 폭은 이 수치로 정할 수 없습니다.

우리 업무에서는 과제와 모델을 고정하고 결과 검사 유무를 비교할 수 있습니다. 만든 파일과 검사 결과, 재시도, 소요 시간을 함께 남기는 방식입니다. 성공한 과제뿐 아니라 사람이 개입해야 했던 실패까지 설명해야 실행 방식의 차이를 판단할 수 있습니다.

메모리는 별도로 시험해야 합니다. 과거 정보를 떠올리더라도 수정된 내용과 접근 범위를 제대로 처리해야 업무에서 쓸 수 있습니다. 보관한 LoCoMo 성적은 단일 사실 회상에 한정됩니다. 변경 충돌과 사용자 간 분리를 검증한 결과로 읽지 않고, 부록의 참고 자료로 남겼습니다.

05 / HDATF 평가 계획

이제 확인할 것은 HDATF가 수행한 실제 업무입니다.

공개 연구로 필요한 질문은 정리할 수 있습니다. 우리 제품의 성적은 직접 측정해야 합니다. 대표 과제로 실행 조건을 기록한 비교를 진행하고, 근거 탐색과 분석, 결과물 완수를 각각 확인할 계획입니다. 한 부분의 높은 점수가 다른 부분의 실패를 가리지 않게 하려는 것입니다.

조사와 데이터 작업에서는 질문, 사용한 자료의 판본, 최종 결과 검사 기록을 남깁니다. 보고서는 근거 없는 주장을 찾을 수 있어야 합니다. 분석은 입력 데이터와 실행 기록, 재현할 결과물을 보존해야 합니다. 아래 기준은 앞으로 확인할 내용이며 완료한 시험은 아닙니다.[1][7][8][9]

평가 계획입니다. HDATF 실측 성적은 아직 공개하지 않았습니다.

필요한 근거를 찾는가

BrowseComp는 웹에서 찾기 어려운 사실을 찾아내는지 평가합니다. LabChin에서는 찾은 출처가 실제로 그 주장을 뒷받침하는지도 별도로 확인하려고 합니다.

근거로 질문에 답하는가

DeepResearch Bench II와 ResearchRubrics를 조사 보고서의 평가 기준으로 삼습니다. 문장이 자연스러운지에 그치지 않고, 필요한 내용과 사실 근거, 결론에 이르는 분석을 확인할 계획입니다.

실제로 쓸 결과물을 만드는가

GDPval과 Agents’ Last Exam은 전문 업무를 다룹니다. 요청한 일이 끝났는지, 파일을 실제로 사용할 수 있는지, 사람이 얼마나 고쳐야 하는지 확인할 계획입니다.

평가 대상남길 근거확인할 판단
과제요청 내용, 입력 파일, 완료 기준요청한 일이 끝났는가
실행모델, 도구, 권한, 실행 구성 판본같은 조건으로 다시 실행할 수 있는가
결과결과 파일, 출처 검사, 사람의 수정 내용결과물을 실제로 쓸 수 있는가
효율전체 시간, 도구 사용, 재시도, 총비용업무 전체의 시간과 비용을 감당할 수 있는가
APPENDIX

부록. 전체 평가 자료

평가별 공개 성적을 비교합니다. 모델 이름을 누르면 원문을 볼 수 있습니다. 표기된 날짜 기준이며, 모델군 비교에는 각 모델군의 최고 성적 구성을 표시했습니다.

73 / 73 평가와 지표

업무와 문서

AA-Briefcase15 개 성적

기반 모델 성적Artificial Analysis

장기 사무 업무 수행

Elo높을수록 좋음
  1. Claude Fable 5.11,661.91
  2. Claude Opus 51,647.02
  3. GPT-6 Astra1,561.95
  4. Muse Spark 1.31,558.85
  5. Grok 4.61,545.82
  6. Claude Fable 51,533.52
  7. GLM-5.31,515.45
  8. Kimi K31,496.83
  9. GPT-5.6 Sol1,474.49
  10. GLM 5.3 Flash1,459.13
  11. Qwen3.8 2.4T A95B1,442.81
  12. DeepSeek V4 Flash 07311,433.23
  13. Qwen3.8 27B1,401.81
  14. Qwen3.8 Max1,392.33
  15. Claude Sonnet 51,358.11
AA-Briefcase / Elo
모델과 실행 구성성적출처
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic1,661.91원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic1,647.02원본 기록
GPT-6 Astra (max)OpenAI1,561.95원본 기록
Muse Spark 1.3 (max)Meta1,558.85원본 기록
Grok 4.6 (xhigh)SpaceXAI1,545.82원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic1,533.52원본 기록
GLM-5.3 (max)Z AI1,515.45원본 기록
Kimi K3 (max)Kimi1,496.83원본 기록
GPT-5.6 Sol (max)OpenAI1,474.49원본 기록
GLM 5.3 Flash (GLM-5.3-Flash)Z AI1,459.13원본 기록
Qwen3.8 2.4T A95BAlibaba1,442.81원본 기록
DeepSeek V4 Flash 0731 (DeepSeek V4 Flash Vision (Reasoning, Max Effort))DeepSeek1,433.23원본 기록
Qwen3.8 27B (xhigh)Alibaba1,401.81원본 기록
Qwen3.8 MaxAlibaba1,392.33원본 기록
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic1,358.11원본 기록
AA-Briefcase / Rubric15 개 성적

기반 모델 성적Artificial Analysis

업무 요구사항 충족

%높을수록 좋음
AA-Briefcase / Rubric / %
모델과 실행 구성성적출처
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic61.52%원본 기록
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic57.98%원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic55.96%원본 기록
Muse Spark 1.3 (max)Meta54.87%원본 기록
Grok 4.6 (xhigh)SpaceXAI53.84%원본 기록
GPT-6 Astra (xhigh)OpenAI52.32%원본 기록
GLM-5.3 (max)Z AI51.35%원본 기록
Kimi K3 (max)Kimi50.97%원본 기록
Qwen3.8 2.4T A95BAlibaba50%원본 기록
Qwen3.8 MaxAlibaba49.6%원본 기록
Qwen3.8 27B (xhigh)Alibaba47.8%원본 기록
GLM 5.3 Flash (GLM-5.3-Flash)Z AI46.36%원본 기록
DeepSeek V4 Flash 0731 (DeepSeek V4 Flash Vision (Reasoning, Max Effort))DeepSeek46.15%원본 기록
Muse Spark 1.2 (xhigh)Meta45.15%원본 기록
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek43.52%원본 기록
AA-Briefcase / Analytical quality15 개 성적

기반 모델 성적Artificial Analysis

분석 내용의 품질

Elo높을수록 좋음
  1. Claude Fable 5.11,959.49
  2. Claude Opus 51,909.74
  3. GLM-5.31,788.62
  4. GPT-6 Astra1,753.64
  5. Muse Spark 1.31,711.52
  6. Grok 4.61,682.58
  7. Claude Fable 51,676.55
  8. Kimi K31,668.89
  9. GLM 5.3 Flash1,650.95
  10. Qwen3.8 2.4T A95B1,617.45
  11. DeepSeek V4 Flash 07311,617.37
  12. Qwen3.8 27B1,590.24
  13. GPT-5.6 Sol1,537.92
  14. Qwen3.8 Max1,533.14
  15. Claude Sonnet 51,418.22
AA-Briefcase / Analytical quality / Elo
모델과 실행 구성성적출처
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic1,959.49원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic1,909.74원본 기록
GLM-5.3 (max)Z AI1,788.62원본 기록
GPT-6 Astra (max)OpenAI1,753.64원본 기록
Muse Spark 1.3 (max)Meta1,711.52원본 기록
Grok 4.6 (xhigh)SpaceXAI1,682.58원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic1,676.55원본 기록
Kimi K3 (max)Kimi1,668.89원본 기록
GLM 5.3 Flash (GLM-5.3-Flash)Z AI1,650.95원본 기록
Qwen3.8 2.4T A95BAlibaba1,617.45원본 기록
DeepSeek V4 Flash 0731 (DeepSeek V4 Flash Vision (Reasoning, Max Effort))DeepSeek1,617.37원본 기록
Qwen3.8 27B (xhigh)Alibaba1,590.24원본 기록
GPT-5.6 Sol (max)OpenAI1,537.92원본 기록
Qwen3.8 MaxAlibaba1,533.14원본 기록
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic1,418.22원본 기록
AA-Briefcase / Presentation15 개 성적

기반 모델 성적Artificial Analysis

산출물 표현 품질

Elo높을수록 좋음
  1. GPT-5.6 Sol1,633.93
  2. Claude Opus 51,544.01
  3. GPT-6 Astra1,535.29
  4. GPT-5.6 Terra1,519.13
  5. Grok 4.61,514.68
  6. Muse Spark 1.31,506.83
  7. Claude Fable 5.11,472.7
  8. GPT-5.6 Luna1,471.03
  9. Claude Fable 51,456.83
  10. Kimi K31,435.29
  11. Claude Sonnet 51,409.5
  12. GLM 5.3 Flash1,409.12
  13. DeepSeek V4 Flash 07311,355.99
  14. GLM-5.31,354.5
  15. Muse Spark 1.21,334.35
AA-Briefcase / Presentation / Elo
모델과 실행 구성성적출처
GPT-5.6 Sol (max)OpenAI1,633.93원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic1,544.01원본 기록
GPT-6 Astra (max)OpenAI1,535.29원본 기록
GPT-5.6 Terra (max)OpenAI1,519.13원본 기록
Grok 4.6 (xhigh)SpaceXAI1,514.68원본 기록
Muse Spark 1.3 (max)Meta1,506.83원본 기록
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic1,472.7원본 기록
GPT-5.6 Luna (max)OpenAI1,471.03원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic1,456.83원본 기록
Kimi K3 (max)Kimi1,435.29원본 기록
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic1,409.5원본 기록
GLM 5.3 Flash (GLM-5.3-Flash)Z AI1,409.12원본 기록
DeepSeek V4 Flash 0731 (DeepSeek V4 Flash Vision (Reasoning, Max Effort))DeepSeek1,355.99원본 기록
GLM-5.3 (max)Z AI1,354.5원본 기록
Muse Spark 1.2 (xhigh)Meta1,334.35원본 기록
GDPval-AA v215 개 성적

기반 모델 성적Artificial Analysis

실제 직무 과제 수행

Elo높을수록 좋음
GDPval-AA v2 / Elo
모델과 실행 구성성적출처
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic1,765.9원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic1,738.08원본 기록
Muse Spark 1.3 (max)Meta1,719.65원본 기록
GLM-5.3 (max)Z AI1,678.47원본 기록
GLM 5.3 Flash (GLM-5.3-Flash)Z AI1,672.94원본 기록
Grok 4.6 (xhigh)SpaceXAI1,664.82원본 기록
Qwen3.8-Flash-NextAlibaba1,649.93원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic1,635.39원본 기록
Qwen3.8 MaxAlibaba1,633.53원본 기록
Qwen3.8 2.4T A95BAlibaba1,630.39원본 기록
GPT-5.6 Sol (max)OpenAI1,625.68원본 기록
Kimi K3 (max)Kimi1,585.72원본 기록
GPT-6 Astra (max)OpenAI1,582.24원본 기록
DeepSeek V4 Flash 0731 (DeepSeek V4 Flash Vision (Reasoning, Max Effort))DeepSeek1,579.3원본 기록
Muse Spark 1.2 (xhigh)Meta1,525.03원본 기록
AA-AnalystAgent / pass^515 개 성적

기반 모델 성적Artificial Analysis

스프레드시트와 문서 분석의 반복 성공

%높을수록 좋음
AA-AnalystAgent / pass^5 / %
모델과 실행 구성성적출처
Gemini 3.7 Flash (high)Google60%원본 기록
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic57.5%원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic53.75%원본 기록
GPT-6 Astra (max)OpenAI51.25%원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic48.75%원본 기록
GPT-5.6 Sol (max)OpenAI47.5%원본 기록
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic46.25%원본 기록
Gemini 3.1 Pro PreviewGoogle41.25%원본 기록
Grok 4.6 (high)SpaceXAI41.25%원본 기록
Kimi K3 (max)Kimi38.75%원본 기록
Grok 4.5 (high)SpaceXAI35%원본 기록
Inkling SmallThinking Machines27.5%원본 기록
Inkling (xhigh)Thinking Machines23.75%원본 기록
MiMo-V2.5-ProXiaomi20%원본 기록
DeepSeek V4 Pro 0424 (DeepSeek V4 Pro (Reasoning, Max Effort))DeepSeek18.75%원본 기록
GDP.pdf / All-pass15 개 성적

기반 모델 성적Artificial Analysis

업무 문서를 읽고 추론

%높을수록 좋음
GDP.pdf / All-pass / %
모델과 실행 구성성적출처
GPT-6 Astra (max)OpenAI33.2%원본 기록
GPT-5.6 Sol (max)OpenAI28.2%원본 기록
Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Anthropic28%원본 기록
Muse Spark 1.3 (max)Meta25.6%원본 기록
GPT-5.6 Terra (max)OpenAI25.6%원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic24%원본 기록
GPT-5.6 Luna (xhigh)OpenAI23.8%원본 기록
Gemini 3.7 Flash (high)Google23.6%원본 기록
Gemini 3.8 Flash (medium)Google22.8%원본 기록
Qwen3.8 MaxAlibaba21.8%원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic21.6%원본 기록
GPT-5.5 Instant (June 2026)OpenAI20.2%원본 기록
Kimi K3 (max)Kimi19.6%원본 기록
Grok 4.5 (high)SpaceXAI18.8%원본 기록
Grok 4.6 (high)SpaceXAI18.8%원본 기록
AutomationBench-AA15 개 성적

기반 모델 성적Artificial Analysis

업무 앱에서 목표 달성

%높을수록 좋음
AutomationBench-AA / %
모델과 실행 구성성적출처
Gemini 3.7 Flash (high)Google62.75%원본 기록
Kimi K3 (max)Kimi52.71%원본 기록
Grok 4.5 (high)SpaceXAI51.44%원본 기록
GPT-5.6 Sol (max)OpenAI51.19%원본 기록
Gemini 3.6 Flash (high)Google51.08%원본 기록
Gemini 3.8 Flash (high)Google50.72%원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic48.58%원본 기록
GPT-5.6 Terra (max)OpenAI45.61%원본 기록
GPT-5.6 Luna (max)OpenAI42.24%원본 기록
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic39.19%원본 기록
Gemini 3.1 Pro PreviewGoogle37.54%원본 기록
Gemini 3.5 Flash-LiteGoogle32.66%원본 기록
GLM-5.2 (max)Z AI27.84%원본 기록
Kimi K2.7 CodeKimi22.53%원본 기록
Qwen3.7 PlusAlibaba20.43%원본 기록
EnterpriseOps-Gym-AA15 개 성적

기반 모델 성적Artificial Analysis

기업 운영 업무 수행

%높을수록 좋음
EnterpriseOps-Gym-AA / %
모델과 실행 구성성적출처
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic51.12%원본 기록
Gemini 3.7 Flash (medium)Google50.4%원본 기록
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek49.6%원본 기록
Grok 4.6 (high)SpaceXAI48.34%원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic47.48%원본 기록
Qwen3.8 2.4T A95BAlibaba47.36%원본 기록
Muse Spark 1.2 (xhigh)Meta47.27%원본 기록
DeepSeek V4 Flash 0731 (DeepSeek V4 Flash Vision (Reasoning, Max Effort))DeepSeek47.09%원본 기록
Kimi K3 (max)Kimi45.33%원본 기록
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic44.67%원본 기록
Qwen3.8 27B (xhigh)Alibaba44.23%원본 기록
GPT-5.6 Sol (max)OpenAI42.91%원본 기록
GLM-5.2 (max)Z AI42.73%원본 기록
Inkling SmallThinking Machines42.7%원본 기록
Gemini 3.5 Flash-LiteGoogle42.35%원본 기록
APEX-Agents-AA14 개 성적

기반 모델 성적Artificial Analysis

여러 앱을 쓰는 장기 업무

%높을수록 좋음
APEX-Agents-AA / %
모델과 실행 구성성적출처
Kimi K3 (max)Kimi41.3%원본 기록
GPT-5.6 Terra (max)OpenAI38.94%원본 기록
GPT-5.6 Luna (max)OpenAI35.84%원본 기록
GLM-5.2 (max)Z AI33.7%원본 기록
Gemini 3.1 Pro PreviewGoogle32.01%원본 기록
Apodex 1.1Apodex31.19%원본 기록
DeepSeek V4 Pro 0424 (DeepSeek V4 Pro (Reasoning, Max Effort))DeepSeek24.26%원본 기록
Qwen3.7 PlusAlibaba22.42%원본 기록
Qwen3.5 397B A17B (Reasoning)Alibaba15.34%원본 기록
Step 3.7 FlashStepFun14.82%원본 기록
gpt-oss-120b (high)OpenAI3.1%원본 기록
MiMo-V2.5-ProXiaomi2.43%원본 기록
Nemotron 3 Super 120B A12B (Reasoning)NVIDIA1.84%원본 기록
gpt-oss-20b (high)OpenAI0.74%원본 기록

시스템 연구

GDPval Gold3 개 성적

전문가 비교 연구OpenAI

전문가가 받아들일 수 있는 업무 산출물을 만드는가? 전문가 산출물 대비 승리와 동률 비율 (%)

%높을수록 좋음
  1. Claude Opus 4.147.6%
  2. GPT-5 (high)38.8%
  3. o3 (high)34.1%
GDPval Gold / %
모델과 실행 구성성적출처
Claude Opus 4.1Anthropic47.6%원본 기록
GPT-5 (high)OpenAI38.8%원본 기록
o3 (high)OpenAI34.1%원본 기록
BrowseComp / Parallel study3 개 성적

조사 에이전트 연구Parallel

찾기 어려운 정보를 탐색하고 근거를 연결하는가? 정답률 (%)

%높을수록 좋음
BrowseComp / Parallel study / %
모델과 실행 구성성적출처
Parallel Ultra8xParallel58%원본 기록
Parallel UltraParallel45%원본 기록
GPT-5 (high)OpenAI38%원본 기록
Terminal-Bench 2.02 개 성적

실행 체계 연구LangChain

실행 방식을 바꾸면 작업 결과가 얼마나 달라지는가? 과제 성공률 (%)

%높을수록 좋음
Terminal-Bench 2.0 / %
모델과 실행 구성성적출처
Deep Agents / improvedLangChain66.5%원본 기록
Deep Agents / baselineLangChain52.8%원본 기록
LoCoMo / single-hop3 개 성적

기억 시스템 연구Mem0 research team

대화가 바뀌어도 필요한 정보를 기억하는가? 단일 사실 회상, AI 채점 점수 (0~100)

score높을수록 좋음
  1. Mem067.13
  2. LangMem62.23
  3. Zep61.7
LoCoMo / single-hop / score
모델과 실행 구성성적출처
Mem0Mem067.13원본 기록
LangMemLangChain62.23원본 기록
ZepZep61.7원본 기록

에이전트와 도구

τ³-Banking15 개 성적

기반 모델 성적Artificial Analysis

은행 규정 검색과 도구 실행

%높을수록 좋음
τ³-Banking / %
모델과 실행 구성성적출처
Muse Spark 1.3 (max)Meta52.37%원본 기록
Qwen3.8 MaxAlibaba51.34%원본 기록
Grok 4.6 (high)SpaceXAI50.72%원본 기록
GLM-5.3 (max)Z AI50.31%원본 기록
Qwen3.8 2.4T A95BAlibaba49.07%원본 기록
Qwen3.8 27B (xhigh)Alibaba48.04%원본 기록
GLM 5.3 Flash (GLM-5.3-Flash)Z AI47.22%원본 기록
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic47.22%원본 기록
Kimi K3 (max)Kimi45.98%원본 기록
Gemini 3.8 Flash (medium)Google45.77%원본 기록
Qwen3.8-Flash-NextAlibaba45.36%원본 기록
Claude Opus 5 (Adaptive Reasoning, High Effort)Anthropic44.74%원본 기록
GPT-5.6 Sol (max)OpenAI44.33%원본 기록
GPT-6 Astra (xhigh)OpenAI43.09%원본 기록
Grok 4.5 (high)SpaceXAI42.06%원본 기록
τ²-Bench15 개 성적

기반 모델 성적Artificial Analysis

대화 중 도구 사용

%높을수록 좋음
τ²-Bench / %
모델과 실행 구성성적출처
JT-35B-FlashChina Mobile99.12%원본 기록
GLM-5.2 (max)Z AI99.12%원본 기록
Step 3.7 FlashStepFun98.54%원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic98.54%원본 기록
DeepSeek V4 Pro 0424 (DeepSeek V4 Pro (Reasoning, Max Effort))DeepSeek96.2%원본 기록
Qwen3.5 397B A17B (Reasoning)Alibaba95.61%원본 기록
Gemini 3.5 Flash (medium)Google95.61%원본 기록
Gemini 3.1 Pro PreviewGoogle95.61%원본 기록
Qwen3.6 35B A3B (Reasoning)Alibaba95.32%원본 기록
DeepSeek V4 Flash 0420 (DeepSeek V4 Flash (Non-reasoning))DeepSeek94.44%원본 기록
MiMo-V2.5-ProXiaomi94.15%원본 기록
Qwen3.6 27B (Reasoning)Alibaba94.15%원본 기록
Mistral Medium 3.5Mistral94.15%원본 기록
Qwen3.5 122B A10B (Reasoning)Alibaba93.57%원본 기록
MiMo-V2-Flash (Feb 2026)Xiaomi93.27%원본 기록
ITBench-AA15 개 성적

기반 모델 성적Artificial Analysis

운영 장애 원인 분석

%높을수록 좋음
ITBench-AA / %
모델과 실행 구성성적출처
GPT-5.6 Sol (max)OpenAI56.21%원본 기록
GPT-5.6 Terra (max)OpenAI51.04%원본 기록
Kimi K3 (max)Kimi47.69%원본 기록
GLM-5.2 (max)Z AI42.66%원본 기록
GPT-5.6 Luna (max)OpenAI40.32%원본 기록
DeepSeek V4 Pro 0424 (DeepSeek V4 Pro (Reasoning, Max Effort))DeepSeek38.32%원본 기록
MiMo-V2.5-ProXiaomi38.23%원본 기록
Gemma 4 31B (Reasoning)Google37.29%원본 기록
Qwen3.5 397B A17B (Reasoning)Alibaba34.09%원본 기록
Gemini 3.1 Pro PreviewGoogle30.33%원본 기록
Step 3.7 FlashStepFun30.27%원본 기록
Claude 4.5 Haiku (Reasoning)Anthropic27.31%원본 기록
Gemma 4 26B A4B (Reasoning)Google23.63%원본 기록
gpt-oss-120b (high)OpenAI5.65%원본 기록
Nemotron 3 Super 120B A12B (Reasoning)NVIDIA1.13%원본 기록
MCP-Atlas public set / Z.ai report7 개 성적

제공사 발표 모델 성적Z.ai (model-card report)

연결된 도구로 과제 완료.

%높을수록 좋음
MCP-Atlas public set / Z.ai report / %
모델과 실행 구성성적출처
GPT-5.2 (xhigh)OpenAI68%원본 기록
GLM-5 release reportZ.ai67.8%원본 기록
Gemini 3 Pro (GLM-5 release comparison)Google66.6%원본 기록
Claude Opus 4.5 (GLM-5 release comparison)Anthropic65.2%원본 기록
Kimi K2.5 (GLM-5 release comparison)Moonshot AI63.8%원본 기록
DeepSeek V3.2 (GLM-5 release comparison)DeepSeek62.2%원본 기록
GLM-4.7 (GLM-5 release comparison)Z.ai52%원본 기록

코드와 실행

Terminal-Bench v2.115 개 성적

기반 모델 성적Artificial Analysis

터미널 작업 완료

%높을수록 좋음
Terminal-Bench v2.1 / %
모델과 실행 구성성적출처
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic91.39%원본 기록
GPT-6 Astra (high)OpenAI89.89%원본 기록
GPT-5.6 Sol (xhigh)OpenAI89.51%원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic89.14%원본 기록
Grok 4.6 (high)SpaceXAI88.39%원본 기록
GPT-5.6 Terra (max)OpenAI88.01%원본 기록
Gemini 3.8 Flash (high)Google87.64%원본 기록
Qwen3.8-Flash-NextAlibaba86.14%원본 기록
Muse Spark 1.3 (max)Meta85.77%원본 기록
Gemini 3.7 Flash (high)Google85.77%원본 기록
Kimi K3 (max)Kimi85.02%원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic84.64%원본 기록
GLM 5.3 Flash (GLM-5.3-Flash)Z AI84.27%원본 기록
GLM-5.3 (max)Z AI83.9%원본 기록
Qwen3.8 2.4T A95BAlibaba82.02%원본 기록
SciCode15 개 성적

기반 모델 성적Artificial Analysis

과학 문제를 코드로 해결

%높을수록 좋음
SciCode / %
모델과 실행 구성성적출처
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic63.08%원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic61%원본 기록
Gemini 3.7 Flash (medium)Google59.84%원본 기록
Muse Spark 1.3 (xhigh)Meta59.72%원본 기록
Kimi K3 (max)Kimi59.49%원본 기록
GLM-5.3 (max)Z AI59.03%원본 기록
Gemini 3.1 Pro PreviewGoogle58.68%원본 기록
GPT-5.6 Sol (high)OpenAI57.75%원본 기록
Muse Spark 1.2 (xhigh)Meta57.41%원본 기록
Gemini 3.8 Flash (high)Google56.6%원본 기록
GPT-6 Astra (max)OpenAI56.48%원본 기록
Grok 4.6 (high)SpaceXAI56.48%원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic56.37%원본 기록
Grok 4.5 (high)SpaceXAI54.98%원본 기록
GPT-5.6 Terra (max)OpenAI54.98%원본 기록
LiveCodeBench15 개 성적

기반 모델 성적Artificial Analysis

코딩 문제 풀이

%높을수록 좋음
LiveCodeBench / %
모델과 실행 구성성적출처
gpt-oss-120b (high)OpenAI87.83%원본 기록
ERNIE 5.0 Thinking PreviewBaidu81.16%원본 기록
o3OpenAI80.85%원본 기록
Apriel-v1.6-15B-ThinkerServiceNow80.74%원본 기록
Qwen3 Next 80B A3B (Reasoning)Alibaba78.41%원본 기록
INTELLECT-3Prime Intellect77.67%원본 기록
gpt-oss-20b (high)OpenAI77.67%원본 기록
K-EXAONE (Reasoning)LG AI Research76.83%원본 기록
Doubao Seed CodeByteDance Seed76.61%원본 기록
Magistral Medium 1.2Mistral75.03%원본 기록
EXAONE 4.0 32B (Reasoning)LG AI Research74.71%원본 기록
NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)NVIDIA74.07%원본 기록
Llama Nemotron Super 49B v1.5 (Reasoning)NVIDIA73.65%원본 기록
Nova 2.0 Pro Preview (medium)Amazon73.02%원본 기록
Falcon-H1R-7BTII UAE72.38%원본 기록
SciCode / MiniMax report7 개 성적

제공사 발표 모델 성적MiniMax (internal evaluation)

과학 분야 프로그래밍 과제.

score높을수록 좋음
SciCode / MiniMax report / score
모델과 실행 구성성적출처
Gemini 3 Pro (MiniMax internal evaluation)Google56원본 기록
Claude Opus 4.6 (MiniMax internal evaluation)Anthropic52원본 기록
GPT-5.2 (Thinking; MiniMax internal evaluation)OpenAI52원본 기록
Claude Opus 4.5 (MiniMax internal evaluation)Anthropic50원본 기록
Claude Sonnet 4.5 (MiniMax internal evaluation)Anthropic45원본 기록
MiniMax-M2.5 (MiniMax internal evaluation)MiniMax44.4원본 기록
MiniMax-M2.1 (MiniMax internal evaluation)MiniMax41원본 기록
Terminal Bench 2.1 / DeepSeek GA report8 개 성적

제공사 발표 모델 성적DeepSeek (GA model-card report)

터미널에서 수행하는 코딩과 도구 사용 과제.

score높을수록 좋음
Terminal Bench 2.1 / DeepSeek GA report / score
모델과 실행 구성성적출처
Kimi K3 (Code-agent evaluation; harness and reasoning effort not disclosed for this model)Moonshot AI88.3원본 기록
Fable-5 (w/ fallback) (Code-agent evaluation; harness and reasoning effort not disclosed for this model; with fallback, details unspecified)Anthropic88원본 기록
DeepSeek-V4-Pro-0813 (DeepSeek Harness minimal; max reasoning effort; temperature=1.0; top_p=0.95)DeepSeek87.9원본 기록
Opus-4.8 (Code-agent evaluation; harness and reasoning effort not disclosed for this model)Anthropic85원본 기록
DeepSeek-V4-Flash-0731 (Code-agent evaluation; harness and reasoning effort not disclosed for this model)DeepSeek82.7원본 기록
GLM-5.2 (Code-agent evaluation; harness and reasoning effort not disclosed for this model)Z.ai81원본 기록
DeepSeek-V4-Pro (Preview) (Code-agent evaluation; harness and reasoning effort not disclosed for this model)DeepSeek72.1원본 기록
DeepSeek-V4-Flash (Preview) (Code-agent evaluation; harness and reasoning effort not disclosed for this model)DeepSeek61.8원본 기록
SWE-bench Pro refined set / Qwen3.8-Max report5 개 성적

제공사 발표 모델 성적Qwen (model-card report)

Qwen이 수정한 SWE-bench Pro 평가셋의 코드 수정 과제.

score높을수록 좋음
SWE-bench Pro refined set / Qwen3.8-Max report / score
모델과 실행 구성성적출처
Fable 5 (Qwen-refined set; Claude Code; temperature=1.0; top_p=0.95; 256K context; may involve fallback; reasoning effort not disclosed)Anthropic80원본 기록
Opus 4.8 (Qwen-refined set; Claude Code; temperature=1.0; top_p=0.95; 256K context; reasoning effort not disclosed)Anthropic69.2원본 기록
Qwen3.8-Max service model; Qwen-refined set; Claude Code; temperature=1.0; top_p=0.95; 256K context; reasoning effort not disclosedQwen67.7원본 기록
GPT 5.6 Sol (max) (Qwen-refined set; Claude Code; max (table header); temperature=1.0; top_p=0.95; 256K context)OpenAI64.6원본 기록
Qwen3.7-Max (Qwen-refined set; Claude Code; temperature=1.0; top_p=0.95; 256K context; reasoning effort not disclosed)Qwen60.6원본 기록
SWE-bench Verified / Google G31 (2026-02-19)5 개 성적

제공사 발표 모델 성적Google DeepMind (Gemini 3.1 Pro release comparison)

모델과 코딩 실행 환경이 저장소 문제를 해결하는 비율입니다.

%높을수록 좋음
SWE-bench Verified / Google G31 (2026-02-19) / %
모델과 실행 구성성적출처
Claude Opus 4.6 (Thinking (Max); Single attempt; provider-specific scaffolding)Anthropic80.8%원본 기록
Gemini 3.1 Pro (Thinking (High); Single attempt; provider-specific scaffolding)Google80.6%원본 기록
GPT-5.2 (Thinking (xhigh); Single attempt; provider-specific scaffolding)OpenAI80%원본 기록
Claude Sonnet 4.6 (Thinking (Max); Single attempt; provider-specific scaffolding)Anthropic79.6%원본 기록
Gemini 3 Pro (Thinking (High); Single attempt; provider-specific scaffolding)Google76.2%원본 기록
SWE-bench Pro (Public) / Google G31 (2026-02-19)4 개 성적

제공사 발표 모델 성적Google DeepMind (Gemini 3.1 Pro release comparison)

SWE-bench Pro 공개판의 소프트웨어 개발 과제입니다.

%높을수록 좋음
SWE-bench Pro (Public) / Google G31 (2026-02-19) / %
모델과 실행 구성성적출처
GPT-5.3-Codex (Thinking (xhigh); Public split; single attempt; provider-specific scaffolding)OpenAI56.8%원본 기록
GPT-5.2 (Thinking (xhigh); Public split; single attempt; provider-specific scaffolding)OpenAI55.6%원본 기록
Gemini 3.1 Pro (Thinking (High); Public split; single attempt; provider-specific scaffolding)Google54.2%원본 기록
Gemini 3 Pro (Thinking (High); Public split; single attempt; provider-specific scaffolding)Google43.3%원본 기록

모델 추론

Intelligence Index v4.215 개 성적

기반 모델 성적Artificial Analysis

10개 평가 종합 지수

score높을수록 좋음
Intelligence Index v4.2 / score
모델과 실행 구성성적출처
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic56.76원본 기록
GPT-6 Astra (max)OpenAI54.66원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic54.05원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic53.19원본 기록
Muse Spark 1.3 (max)Meta52.95원본 기록
GPT-5.6 Sol (max)OpenAI51.26원본 기록
Grok 4.6 (high)SpaceXAI50.58원본 기록
Kimi K3 (max)Kimi50.23원본 기록
GLM-5.3 (max)Z AI48.58원본 기록
Gemini 3.8 Flash (high)Google47.07원본 기록
Qwen3.8 MaxAlibaba46.91원본 기록
Muse Spark 1.2 (xhigh)Meta46.84원본 기록
GPT-5.6 Terra (max)OpenAI46.77원본 기록
Qwen3.8 2.4T A95BAlibaba46.74원본 기록
GLM 5.3 Flash (GLM-5.3-Flash)Z AI46.22원본 기록
Humanity's Last Exam15 개 성적

기반 모델 성적Artificial Analysis

전문 지식과 고난도 추론

%높을수록 좋음
Humanity's Last Exam / %
모델과 실행 구성성적출처
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic59.13%원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic55.47%원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic54.87%원본 기록
GPT-6 Astra (max)OpenAI54.68%원본 기록
GPT-5.6 Sol (max)OpenAI49.49%원본 기록
Muse Spark 1.3 (max)Meta49.07%원본 기록
Gemini 3.7 Flash (high)Google47.87%원본 기록
Gemini 3.8 Flash (high)Google47.82%원본 기록
Gemini 3.1 Pro PreviewGoogle47.03%원본 기록
Kimi K3 (max)Kimi46.9%원본 기록
Muse Spark 1.2 (xhigh)Meta45.46%원본 기록
Grok 4.6 (xhigh)SpaceXAI44.07%원본 기록
Qwen3.8 MaxAlibaba43.05%원본 기록
GPT-5.6 Terra (max)OpenAI42.91%원본 기록
Grok 4.5 (high)SpaceXAI42.68%원본 기록
GPQA Diamond15 개 성적

기반 모델 성적Artificial Analysis

대학원 수준 과학 추론

%높을수록 좋음
GPQA Diamond / %
모델과 실행 구성성적출처
GPT-6 Astra (xhigh)OpenAI96.26%원본 기록
Gemini 3.8 Flash (high)Google95.25%원본 기록
Grok 4.6 (high)SpaceXAI94.95%원본 기록
Gemini 3.7 Flash (high)Google94.55%원본 기록
Gemini 3.1 Pro PreviewGoogle94.14%원본 기록
Muse Spark 1.3 (xhigh)Meta94.14%원본 기록
GPT-5.6 Sol (max)OpenAI94.14%원본 기록
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic93.74%원본 기록
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic93.74%원본 기록
Qwen3.8 2.4T A95BAlibaba93.54%원본 기록
Kimi K3 (max)Kimi93.54%원본 기록
Grok 4.5 (high)SpaceXAI93.13%원본 기록
MiniMax-M3MiniMax92.93%원본 기록
Gemini 3.6 Flash (high)Google92.83%원본 기록
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek92.83%원본 기록
CritPt15 개 성적

기반 모델 성적Artificial Analysis

연구 수준 물리 문제

%높을수록 좋음
CritPt / %
모델과 실행 구성성적출처
GPT-5.6 Sol (max)OpenAI32.29%원본 기록
GPT-6 Astra (max)OpenAI31.71%원본 기록
Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Anthropic31.14%원본 기록
GPT-5.5 Pro (xhigh)OpenAI30.57%원본 기록
GPT-5.6 Terra (max)OpenAI30%원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic29.14%원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic28.57%원본 기록
Muse Spark 1.3 (xhigh)Meta26%원본 기록
Gemini 3 Deep ThinkGoogle25.71%원본 기록
Kimi K3 (max)Kimi23.43%원본 기록
GLM-5.2 (max)Z AI20.86%원본 기록
GPT-5.6 Luna (max)OpenAI20.57%원본 기록
Qwen3.8 MaxAlibaba20%원본 기록
Qwen3.8 2.4T A95BAlibaba20%원본 기록
Grok 4.6 (xhigh)SpaceXAI19.71%원본 기록
AIME 202515 개 성적

기반 모델 성적Artificial Analysis

경시대회 수학 문제

%높을수록 좋음
AIME 2025 / %
모델과 실행 구성성적출처
Nova 2.0 Lite (high)Amazon94.33%원본 기록
gpt-oss-120b (high)OpenAI93.44%원본 기록
NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)NVIDIA91%원본 기록
K-EXAONE (Reasoning)LG AI Research90.33%원본 기록
Nova 2.0 Omni (medium)Amazon89.67%원본 기록
gpt-oss-20b (high)OpenAI89.33%원본 기록
Nova 2.0 Pro Preview (medium)Amazon89%원본 기록
o3OpenAI88.33%원본 기록
INTELLECT-3Prime Intellect88%원본 기록
Apriel-v1.6-15B-ThinkerServiceNow88%원본 기록
ERNIE 5.0 Thinking PreviewBaidu85%원본 기록
Qwen3 Next 80B A3B (Reasoning)Alibaba84.33%원본 기록
Ring-flash-2.0InclusionAI83.67%원본 기록
Claude 4.5 Haiku (Reasoning)Anthropic83.67%원본 기록
Magistral Medium 1.2Mistral82%원본 기록
IFBench15 개 성적

기반 모델 성적Artificial Analysis

복잡한 지시 이행

%높을수록 좋음
IFBench / %
모델과 실행 구성성적출처
Grok 4.3 (medium)SpaceXAI83.33%원본 기록
MiniMax-M3MiniMax82.86%원본 기록
Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA81.36%원본 기록
Nemotron Cascade 2 30B A3BNVIDIA80.41%원본 기록
MiMo-V2.5-ProXiaomi79.86%원본 기록
Nova 2.0 Pro Preview (low)Amazon79.59%원본 기록
Qwen3.5 397B A17B (Reasoning)Alibaba78.78%원본 기록
Qwen3.7 PlusAlibaba77.96%원본 기록
Gemini 3.1 Pro PreviewGoogle77.14%원본 기록
DeepSeek V4 Pro 0424 (DeepSeek V4 Pro (Reasoning, Max Effort))DeepSeek76.46%원본 기록
Qwen3.5 122B A10B (Reasoning)Alibaba75.71%원본 기록
Gemma 4 31B (Reasoning)Google75.58%원본 기록
GPT-5.3 Codex (xhigh)OpenAI75.37%원본 기록
Gemini 3.5 Flash (medium)Google74.56%원본 기록
Command A+Cohere73.95%원본 기록
LongBench v2 / Moonshot report5 개 성적

제공사 발표 모델 성적Moonshot AI (model-card report)

긴 입력을 읽고 추론하는 능력.

score높을수록 좋음
LongBench v2 / Moonshot report / score
모델과 실행 구성성적출처
Gemini 3 Pro (High Thinking Level; publisher rerun (*))Google68.2원본 기록
Claude Opus 4.5 (Extended Thinking; publisher rerun (*))Anthropic64.4원본 기록
Kimi K2.5 (Thinking)Moonshot AI61원본 기록
DeepSeek V3.2 (Thinking; publisher rerun (*))DeepSeek59.8원본 기록
GPT-5.2 (xhigh; publisher rerun (*))OpenAI54.5원본 기록
AIME 2026 I / Z.ai report6 개 성적

제공사 발표 모델 성적Z.ai (model-card report)

2026년 AIME 첫 번째 시험입니다. 2026년 전체 시험과 구분합니다.

score높을수록 좋음
AIME 2026 I / Z.ai report / score
모델과 실행 구성성적출처
Claude Opus 4.5 (GLM-5 release comparison)Anthropic93.3원본 기록
GLM-4.7 (GLM-5 release comparison)Z.ai92.9원본 기록
GLM-5 release reportZ.ai92.7원본 기록
DeepSeek V3.2 (GLM-5 release comparison)DeepSeek92.7원본 기록
Kimi K2.5 (GLM-5 release comparison)Moonshot AI92.5원본 기록
Gemini 3 Pro (GLM-5 release comparison)Google90.6원본 기록
HMMT November 2025 / Z.ai report7 개 성적

제공사 발표 모델 성적Z.ai (model-card report)

2025년 11월 수학 경시대회 문제.

score높을수록 좋음
HMMT November 2025 / Z.ai report / score
모델과 실행 구성성적출처
GPT-5.2 (xhigh)OpenAI97.1원본 기록
GLM-5 release reportZ.ai96.9원본 기록
GLM-4.7 (GLM-5 release comparison)Z.ai93.5원본 기록
Gemini 3 Pro (GLM-5 release comparison)Google93원본 기록
Claude Opus 4.5 (GLM-5 release comparison)Anthropic91.7원본 기록
Kimi K2.5 (GLM-5 release comparison)Moonshot AI91.1원본 기록
DeepSeek V3.2 (GLM-5 release comparison)DeepSeek90.2원본 기록
AA-LCR / MiniMax report7 개 성적

제공사 발표 모델 성적MiniMax (internal evaluation)

긴 입력을 바탕으로 한 추론.

score높을수록 좋음
AA-LCR / MiniMax report / score
모델과 실행 구성성적출처
Claude Opus 4.5 (MiniMax internal evaluation)Anthropic74원본 기록
GPT-5.2 (Thinking; MiniMax internal evaluation)OpenAI73원본 기록
Claude Opus 4.6 (MiniMax internal evaluation)Anthropic71원본 기록
Gemini 3 Pro (MiniMax internal evaluation)Google71원본 기록
MiniMax-M2.5 (MiniMax internal evaluation)MiniMax69.5원본 기록
Claude Sonnet 4.5 (MiniMax internal evaluation)Anthropic66원본 기록
MiniMax-M2.1 (MiniMax internal evaluation)MiniMax62원본 기록
HLE no tools / DeepSeek GA report8 개 성적

제공사 발표 모델 성적DeepSeek (GA model-card report)

도구를 쓰지 않는 여러 분야의 추론 평가.

score높을수록 좋음
HLE no tools / DeepSeek GA report / score
모델과 실행 구성성적출처
Fable-5 (w/ fallback) (No tools; HLE reasoning effort and token budget not disclosed; with fallback, details unspecified)Anthropic53.3원본 기록
Opus-4.8 (No tools; HLE reasoning effort and token budget not disclosed)Anthropic49.8원본 기록
Kimi K3 (No tools; HLE reasoning effort and token budget not disclosed)Moonshot AI43.5원본 기록
DeepSeek-V4-Pro-0813 (No tools; HLE reasoning effort and token budget not disclosed)DeepSeek42.7원본 기록
GLM-5.2 (No tools; HLE reasoning effort and token budget not disclosed)Z.ai40.5원본 기록
DeepSeek-V4-Flash-0731 (No tools; HLE reasoning effort and token budget not disclosed)DeepSeek37.8원본 기록
DeepSeek-V4-Pro (Preview) (No tools; HLE reasoning effort and token budget not disclosed)DeepSeek37.7원본 기록
DeepSeek-V4-Flash (Preview) (No tools; HLE reasoning effort and token budget not disclosed)DeepSeek34.8원본 기록
HLE with tools / DeepSeek GA report8 개 성적

제공사 발표 모델 성적DeepSeek (GA model-card report)

도구를 사용하는 여러 분야의 추론 평가.

score높을수록 좋음
HLE with tools / DeepSeek GA report / score
모델과 실행 구성성적출처
Fable-5 (w/ fallback) (With tools; specific tools, HLE reasoning effort and token budget not disclosed; with fallback, details unspecified)Anthropic63원본 기록
DeepSeek-V4-Pro-0813 (With tools; specific tools, HLE reasoning effort and token budget not disclosed)DeepSeek60원본 기록
Opus-4.8 (With tools; specific tools, HLE reasoning effort and token budget not disclosed)Anthropic57.9원본 기록
Kimi K3 (With tools; specific tools, HLE reasoning effort and token budget not disclosed)Moonshot AI56원본 기록
GLM-5.2 (With tools; specific tools, HLE reasoning effort and token budget not disclosed)Z.ai54.7원본 기록
DeepSeek-V4-Flash-0731 (With tools; specific tools, HLE reasoning effort and token budget not disclosed)DeepSeek51.5원본 기록
DeepSeek-V4-Pro (Preview) (With tools; specific tools, HLE reasoning effort and token budget not disclosed)DeepSeek48.2원본 기록
DeepSeek-V4-Flash (Preview) (With tools; specific tools, HLE reasoning effort and token budget not disclosed)DeepSeek45.1원본 기록
HMMT February 2026 / DeepSeek Preview report6 개 성적

제공사 발표 모델 성적DeepSeek (Preview technical report)

답 하나의 성공률인 Pass@1로 측정한 수학 경시대회 문제 풀이.

%높을수록 좋음
HMMT February 2026 / DeepSeek Preview report / %
모델과 실행 구성성적출처
GPT-5.4 xHigh (xHigh; DeepSeek Preview report comparison)OpenAI97.7%원본 기록
Opus-4.6 Max (Max; DeepSeek Preview report comparison)Anthropic96.2%원본 기록
DS-V4-Pro Max (Preview; Max; temperature=1.0; 384K-token context; distinct rigorous-proof math prompt)DeepSeek95.2%원본 기록
Gemini-3.1-Pro High (High; DeepSeek Preview report comparison)Google94.7%원본 기록
K2.6 Thinking (Thinking; DeepSeek Preview report comparison)Moonshot AI92.7%원본 기록
GLM-5.1 Thinking (Thinking; DeepSeek Preview report comparison)Z.ai89.4%원본 기록
LongBench v2 / Qwen3.8-Max report4 개 성적

제공사 발표 모델 성적Qwen (model-card report)

긴 문서와 입력을 읽고 추론하는 능력.

score높을수록 좋음
LongBench v2 / Qwen3.8-Max report / score
모델과 실행 구성성적출처
Opus 4.8 (LongBench v2; tools, reasoning effort and token budget not disclosed)Anthropic69.1원본 기록
GPT 5.6 Sol (max) (max (table header); LongBench v2 tools and token budget not disclosed)OpenAI67.1원본 기록
Qwen3.8-Max service model; LongBench v2 tools, reasoning effort and token budget not disclosedQwen66.3원본 기록
Qwen3.7-Max (LongBench v2; tools, reasoning effort and token budget not disclosed)Qwen65.3원본 기록
GPQA Diamond / Google G31 (2026-02-19)5 개 성적

제공사 발표 모델 성적Google DeepMind (Gemini 3.1 Pro release comparison)

도구 없이 과학 지식과 추론을 평가합니다.

%높을수록 좋음
GPQA Diamond / Google G31 (2026-02-19) / %
모델과 실행 구성성적출처
Gemini 3.1 Pro (Thinking (High); No tools)Google94.3%원본 기록
GPT-5.2 (Thinking (xhigh); No tools)OpenAI92.4%원본 기록
Gemini 3 Pro (Thinking (High); No tools)Google91.9%원본 기록
Claude Opus 4.6 (Thinking (Max); No tools)Anthropic91.3%원본 기록
Claude Sonnet 4.6 (Thinking (Max); No tools)Anthropic89.9%원본 기록
HLE, no tools / Google G31 (2026-02-19)5 개 성적

제공사 발표 모델 성적Google DeepMind (Gemini 3.1 Pro release comparison)

전체 텍스트와 멀티모달 HLE 문항을 도구 없이 풉니다.

%높을수록 좋음
HLE, no tools / Google G31 (2026-02-19) / %
모델과 실행 구성성적출처
Gemini 3.1 Pro (Thinking (High); No tools)Google44.4%원본 기록
Claude Opus 4.6 (Thinking (Max); No tools)Anthropic40%원본 기록
Gemini 3 Pro (Thinking (High); No tools)Google37.5%원본 기록
GPT-5.2 (Thinking (xhigh); No tools)OpenAI34.5%원본 기록
Claude Sonnet 4.6 (Thinking (Max); No tools)Anthropic33.2%원본 기록
HLE, search and code / Google G31 (2026-02-19)5 개 성적

제공사 발표 모델 성적Google DeepMind (Gemini 3.1 Pro release comparison)

검색과 코드 실행을 사용한 HLE 추론 평가입니다.

%높을수록 좋음
HLE, search and code / Google G31 (2026-02-19) / %
모델과 실행 구성성적출처
Claude Opus 4.6 (Thinking (Max); Search and code; provider-specific tools)Anthropic53.1%원본 기록
Gemini 3.1 Pro (Thinking (High); Search and code; provider-specific tools)Google51.4%원본 기록
Claude Sonnet 4.6 (Thinking (Max); Search and code; provider-specific tools)Anthropic49%원본 기록
Gemini 3 Pro (Thinking (High); Search and code; provider-specific tools)Google45.8%원본 기록
GPT-5.2 (Thinking (xhigh); Search and code; provider-specific tools)OpenAI45.5%원본 기록
ARC-AGI-2 / Google G31 (2026-02-19)5 개 성적

제공사 발표 모델 성적Google DeepMind (Gemini 3.1 Pro release comparison)

격자 예시에서 처음 보는 규칙을 추론합니다.

%높을수록 좋음
ARC-AGI-2 / Google G31 (2026-02-19) / %
모델과 실행 구성성적출처
Gemini 3.1 Pro (Thinking (High); ARC Prize Verified; semi-private set)Google77.1%원본 기록
Claude Opus 4.6 (Thinking (Max); ARC Prize Verified; semi-private set)Anthropic68.8%원본 기록
Claude Sonnet 4.6 (Thinking (Max); ARC Prize Verified; semi-private set)Anthropic58.3%원본 기록
GPT-5.2 (Thinking (xhigh); ARC Prize Verified; semi-private set)OpenAI52.9%원본 기록
Gemini 3 Pro (Thinking (High); ARC Prize Verified; semi-private set)Google31.1%원본 기록
AIME 2025, no tools / Google G3F (2025-12-17)7 개 성적

제공사 발표 모델 성적Google DeepMind (Gemini 3 Flash release comparison)

2025년 AIME 수학 경시대회 문제 풀이입니다.

%높을수록 좋음
AIME 2025, no tools / Google G3F (2025-12-17) / %
모델과 실행 구성성적출처
GPT-5.2 (Extra high; No tools)OpenAI100%원본 기록
Gemini 3 Flash (Thinking; exact level unspecified; No tools)Google95.2%원본 기록
Gemini 3 Pro (Thinking; exact level unspecified; No tools)Google95%원본 기록
Grok 4.1 Fast (Reasoning; No tools)xAI91.9%원본 기록
Gemini 2.5 Pro (Thinking; exact level unspecified; No tools)Google88%원본 기록
Claude Sonnet 4.5 (Thinking; high preferred, otherwise best available reported setting; No tools)Anthropic87%원본 기록
Gemini 2.5 Flash (Thinking; exact level unspecified; No tools)Google72%원본 기록
AIME 2025, code execution / Google G3F (2025-12-17)4 개 성적

제공사 발표 모델 성적Google DeepMind (Gemini 3 Flash release comparison)

2025년 AIME 수학 경시대회 문제 풀이입니다.

%높을수록 좋음
AIME 2025, code execution / Google G3F (2025-12-17) / %
모델과 실행 구성성적출처
Gemini 3 Pro (Thinking; exact level unspecified; Code execution)Google100%원본 기록
Claude Sonnet 4.5 (Thinking; high preferred, otherwise best available reported setting; Code execution)Anthropic100%원본 기록
Gemini 3 Flash (Thinking; exact level unspecified; Code execution)Google99.7%원본 기록
Gemini 2.5 Flash (Thinking; exact level unspecified; Code execution)Google75.7%원본 기록
HLE, no tools / Anthropic A51 (2026-09-01)3 개 성적

제공사 발표 모델 성적Anthropic (Fable 5.1 release comparison)

기존 HLE 멀티모달 2,500문항을 도구 없이 풉니다.

%높을수록 좋음
HLE, no tools / Anthropic A51 (2026-09-01) / %
모델과 실행 구성성적출처
Claude Fable 5.1 (Adaptive thinking (auto); max effort; default sampling; five-trial mean; total token cap 1M; no compaction; production safeguards enabled; No tools)Anthropic60.9%원본 기록
Claude Fable 5 (Detailed reasoning setting and budget unverified; production safeguards enabled; No tools)Anthropic57.8%원본 기록
Claude Opus 5 (Detailed reasoning setting and budget unverified; No tools)Anthropic56.6%원본 기록
HLE, tools / Anthropic A51 (2026-09-01)3 개 성적

제공사 발표 모델 성적Anthropic (Fable 5.1 release comparison)

검색, 웹 문서 가져오기, 코드 도구를 사용하는 기존 HLE 추론 평가입니다.

%높을수록 좋음
HLE, tools / Anthropic A51 (2026-09-01) / %
모델과 실행 구성성적출처
Claude Fable 5.1 (Adaptive thinking (auto); max effort; default sampling; five-trial mean; total token cap 1M; no compaction; production safeguards enabled; Search; restricted fetch; programmatic tool calling; code execution)Anthropic65%원본 기록
Claude Fable 5 (Detailed reasoning setting and budget unverified; production safeguards enabled; Search; restricted fetch; programmatic tool calling; code execution)Anthropic63.8%원본 기록
Claude Opus 5 (Detailed reasoning setting and budget unverified; Search; restricted fetch; programmatic tool calling; code execution)Anthropic63.6%원본 기록
ARC-AGI-2 / Anthropic A51 (2026-09-01)4 개 성적

제공사 발표 모델 성적Anthropic (Fable 5.1 release comparison)

9월 발표 비교표의 새로운 격자 규칙 추론 평가입니다.

%높을수록 좋음
ARC-AGI-2 / Anthropic A51 (2026-09-01) / %
모델과 실행 구성성적출처
GPT-5.6 Sol (Exact reasoning setting and token budget unverified)OpenAI92.5%원본 기록
Claude Opus 5 (Exact reasoning setting and token budget unverified)Anthropic90.42%원본 기록
Claude Fable 5.1 / Claude Mythos 5.1 (Fable 5.1: max effort; semi-private set; ARC Prize Verified)Anthropic90%원본 기록
Claude Fable 5 / Claude Mythos 5 (Exact reasoning setting and token budget unverified)Anthropic89.2%원본 기록
HLE-Verified / Google G38 (2026-09-02)6 개 성적

제공사 발표 모델 성적Google DeepMind (Gemini 3.8 Flash release comparison)

검증과 수정을 거친 1,811문항의 전문 지식 추론 평가입니다.

%높을수록 좋음
HLE-Verified / Google G38 (2026-09-02) / %
모델과 실행 구성성적출처
Gemini 3.8 Flash (Default sampling; exact thinking level unspecified; HLE-Verified; 1811 items; tool use unspecified)Google54.9%원본 기록
GPT-5.6 Sol (Exact reasoning level unspecified; HLE-Verified; 1811 items; tool use unspecified)OpenAI54.5%원본 기록
Claude Opus 5 (Exact thinking level unspecified; HLE-Verified; 1811 items; tool use unspecified)Anthropic54.4%원본 기록
Gemini 3.7 Flash (Exact thinking level unspecified; HLE-Verified; 1811 items; tool use unspecified)Google53.6%원본 기록
GPT-5.6 Terra (Maximum reasoning preferred; otherwise best available reported setting; HLE-Verified; 1811 items; tool use unspecified)OpenAI51.1%원본 기록
Claude Sonnet 5 (Maximum thinking preferred; otherwise best available reported setting; HLE-Verified; 1811 items; tool use unspecified)Anthropic31%원본 기록

문서와 맥락 이해

AA-LCR v1.115 개 성적

기반 모델 성적Artificial Analysis

긴 자료를 읽고 추론

%높을수록 좋음
AA-LCR v1.1 / %
모델과 실행 구성성적출처
Kimi K3 (max)Kimi88.67%원본 기록
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic85.33%원본 기록
Muse Spark 1.3 (max)Meta84.33%원본 기록
Gemini 3.8 Flash (medium)Google84%원본 기록
GPT-5.6 Sol (max)OpenAI84%원본 기록
GPT-5.6 Luna (max)OpenAI83.67%원본 기록
Muse Glimmer (high)Meta83.33%원본 기록
GPT-5.3 Codex (xhigh)OpenAI83.33%원본 기록
MiniMax-M3MiniMax83%원본 기록
GPT-5.6 Terra (max)OpenAI83%원본 기록
Gemini 3.7 Flash (medium)Google83%원본 기록
Agnes 2.5 Pro BetaSapiens AI83%원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic82.33%원본 기록
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic82%원본 기록
Qwen3.8 27B (xhigh)Alibaba82%원본 기록
MMMU-Pro15 개 성적

기반 모델 성적Artificial Analysis

이미지를 읽고 추론

%높을수록 좋음
MMMU-Pro / %
모델과 실행 구성성적출처
GPT-6 Astra (max)OpenAI86.88%원본 기록
Gemini 3.8 Flash (high)Google85.61%원본 기록
Gemini 3.7 Flash (high)Google85.49%원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic84.74%원본 기록
Gemini 3.5 Flash (medium)Google83.87%원본 기록
GPT-5.6 Sol (max)OpenAI83.41%원본 기록
Gemini 3.6 Flash (high)Google83.24%원본 기록
Gemini 3.1 Pro PreviewGoogle82.43%원본 기록
Qwen3.8 MaxAlibaba82.31%원본 기록
Muse Spark 1.3 (xhigh)Meta82.02%원본 기록
GPT-5.6 Terra (max)OpenAI80.69%원본 기록
Kimi K3 (max)Kimi80.52%원본 기록
Qwen3.7 PlusAlibaba80.46%원본 기록
Grok 4.5 (high)SpaceXAI80.4%원본 기록
Qwen3.8-Flash-NextAlibaba79.77%원본 기록
AA-Omniscience / Index15 개 성적

기반 모델 성적Artificial Analysis

지식 정확성과 잘못된 답변

score높을수록 좋음
AA-Omniscience / Index / score
모델과 실행 구성성적출처
GPT-6 Astra (high)OpenAI43.73원본 기록
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic43.45원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic43.3원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic37.07원본 기록
Gemini 3.1 Pro PreviewGoogle31.88원본 기록
Grok 4.6 (high)SpaceXAI30.48원본 기록
Gemini 3.8 Flash (high)Google29.55원본 기록
Muse Spark 1.2 (xhigh)Meta27.2원본 기록
Gemini 3.7 Flash (high)Google26.48원본 기록
Grok 4.5 (high)SpaceXAI25.32원본 기록
Muse Spark 1.3 (max)Meta24.93원본 기록
Gemini 3.6 Flash (high)Google22.13원본 기록
GPT-5.6 Sol (max)OpenAI21.97원본 기록
Gemini 3.5 Flash (medium)Google20.82원본 기록
Kimi K3 (max)Kimi19.7원본 기록
AA-Omniscience / Accuracy15 개 성적

기반 모델 성적Artificial Analysis

지식 질문 정답률

%높을수록 좋음
AA-Omniscience / Accuracy / %
모델과 실행 구성성적출처
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic67.23%원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic65.35%원본 기록
GPT-6 Astra (max)OpenAI62.6%원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic60.87%원본 기록
GPT-5.6 Sol (max)OpenAI59.4%원본 기록
Gemini 3.7 Flash (high)Google55.32%원본 기록
Gemini 3.1 Pro PreviewGoogle54.85%원본 기록
Gemini 3.8 Flash (high)Google54.6%원본 기록
GPT-5.3 Codex (xhigh)OpenAI52.88%원본 기록
Grok 4.5 (high)SpaceXAI51.55%원본 기록
Gemini 3.5 Flash (medium)Google51.05%원본 기록
Gemini 3.6 Flash (high)Google49.97%원본 기록
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek49.1%원본 기록
Grok 4.6 (high)SpaceXAI48.23%원본 기록
Kimi K3 (max)Kimi47.58%원본 기록
AA-Omniscience / Hallucination15 개 성적

기반 모델 성적Artificial Analysis

잘못된 답변 비율, 낮을수록 좋음

%낮을수록 좋음
AA-Omniscience / Hallucination / %
모델과 실행 구성성적출처
MiniCPM5-1B (Non-reasoning)OpenBMB0.9%원본 기록
G9v3-3BAI9Stars11.66%원본 기록
G9v3-39A5BAI9Stars13.04%원본 기록
Command A+Cohere14.16%원본 기록
LFM2.5-2.6BLiquid AI16%원본 기록
Grok 4.3 (medium)SpaceXAI16.94%원본 기록
Qwen3.8 27B (Non-reasoning)Alibaba18.09%원본 기록
MiniMax-M3MiniMax18.43%원본 기록
Quasar 438B (max, based on GLM-5.2)Multiverse Computing21.44%원본 기록
K-EXAONE 2.0 0803 (K-EXAONE 2.0)LG AI Research22.6%원본 기록
Grok 4.6 (medium)SpaceXAI24%원본 기록
Solar Pro 4Upstage24.4%원본 기록
MiMo-V2.5-ProXiaomi24.7%원본 기록
Solar Open2 250BUpstage25.38%원본 기록
Granite 4.2 30BIBM25.58%원본 기록
MathVision / Moonshot report5 개 성적

제공사 발표 모델 성적Moonshot AI (model-card report)

그림이 포함된 수학 문제 풀이.

score높을수록 좋음
MathVision / Moonshot report / score
모델과 실행 구성성적출처
Gemini 3 Pro (High Thinking Level; publisher rerun (*))Google86.1원본 기록
Kimi K2.5 (Thinking)Moonshot AI84.2원본 기록
GPT-5.2 (xhigh)OpenAI83원본 기록
Claude Opus 4.5 (Extended Thinking; publisher rerun (*))Anthropic77.1원본 기록
Qwen3-VL-235B-A22B (Thinking)Qwen74.6원본 기록
OCRBench / Moonshot report5 개 성적

제공사 발표 모델 성적Moonshot AI (model-card report)

이미지 안의 글자 인식.

score높을수록 좋음
OCRBench / Moonshot report / score
모델과 실행 구성성적출처
Kimi K2.5 (Thinking)Moonshot AI92.3원본 기록
Gemini 3 Pro (High Thinking Level; publisher rerun (*))Google90.3원본 기록
Qwen3-VL-235B-A22B (Thinking)Qwen87.5원본 기록
Claude Opus 4.5 (Extended Thinking; publisher rerun (*))Anthropic86.5원본 기록
GPT-5.2 (xhigh; publisher rerun (*))OpenAI80.7원본 기록
OmniDocBench 1.5 / Moonshot report5 개 성적

제공사 발표 모델 성적Moonshot AI (model-card report)

문서 읽기. 발표 점수는 정규화 편집거리의 보수를 100점으로 환산한 값입니다.

score높을수록 좋음
OmniDocBench 1.5 / Moonshot report / score
모델과 실행 구성성적출처
Kimi K2.5 (Thinking)Moonshot AI88.8원본 기록
Gemini 3 Pro (High Thinking Level)Google88.5원본 기록
Claude Opus 4.5 (Extended Thinking; publisher rerun (*))Anthropic87.7원본 기록
GPT-5.2 (xhigh)OpenAI85.7원본 기록
Qwen3-VL-235B-A22B (Thinking; publisher rerun (*))Qwen82원본 기록
VideoMMMU / Moonshot report5 개 성적

제공사 발표 모델 성적Moonshot AI (model-card report)

여러 분야의 영상 이해.

score높을수록 좋음
VideoMMMU / Moonshot report / score
모델과 실행 구성성적출처
Gemini 3 Pro (High Thinking Level)Google87.6원본 기록
Kimi K2.5 (Thinking)Moonshot AI86.6원본 기록
GPT-5.2 (xhigh)OpenAI85.9원본 기록
Claude Opus 4.5 (Extended Thinking; publisher rerun (*))Anthropic84.4원본 기록
Qwen3-VL-235B-A22B (Thinking)Qwen80원본 기록
MotionBench / Moonshot report4 개 성적

제공사 발표 모델 성적Moonshot AI (model-card report)

영상 속 움직임 이해.

score높을수록 좋음
MotionBench / Moonshot report / score
모델과 실행 구성성적출처
Kimi K2.5 (Thinking)Moonshot AI70.4원본 기록
Gemini 3 Pro (High Thinking Level)Google70.3원본 기록
GPT-5.2 (xhigh)OpenAI64.8원본 기록
Claude Opus 4.5 (Extended Thinking)Anthropic60.3원본 기록
IFBench / MiniMax report7 개 성적

제공사 발표 모델 성적MiniMax (internal evaluation)

세부 지시 준수.

score높을수록 좋음
IFBench / MiniMax report / score
모델과 실행 구성성적출처
GPT-5.2 (Thinking; MiniMax internal evaluation)OpenAI75원본 기록
MiniMax-M2.5 (MiniMax internal evaluation)MiniMax70원본 기록
MiniMax-M2.1 (MiniMax internal evaluation)MiniMax70원본 기록
Gemini 3 Pro (MiniMax internal evaluation)Google70원본 기록
Claude Opus 4.5 (MiniMax internal evaluation)Anthropic58원본 기록
Claude Sonnet 4.5 (MiniMax internal evaluation)Anthropic57원본 기록
Claude Opus 4.6 (MiniMax internal evaluation)Anthropic53원본 기록
MMMLU / DeepSeek Base report3 개 성적

제공사 발표 모델 성적DeepSeek (Base technical report)

예시 5개를 제공하고 정답 일치율로 측정한 다국어 지식.

%높을수록 좋음
MMMLU / DeepSeek Base report / %
모델과 실행 구성성적출처
DeepSeek-V4-Pro-Base (Base; MMMLU EM; 5-shot; shared internal evaluation setup; sampling and tools not disclosed)DeepSeek90.3%원본 기록
DeepSeek-V4-Flash-Base (Base; MMMLU EM; 5-shot; shared internal evaluation setup; sampling and tools not disclosed)DeepSeek88.8%원본 기록
DeepSeek-V3.2-Base (Base; MMMLU EM; 5-shot; shared internal evaluation setup; sampling and tools not disclosed)DeepSeek87.9%원본 기록
OmniDocBench 1.5 / Qwen3.8-27B report5 개 성적

제공사 발표 모델 성적Qwen (model-card report)

제공사가 발표한 점수로 비교하는 문서 이미지 읽기.

score높을수록 좋음
OmniDocBench 1.5 / Qwen3.8-27B report / score
모델과 실행 구성성적출처
Qwen3.7-Plus (OmniDocBench 1.5; tools, reasoning effort and image settings not disclosed)Qwen91.4원본 기록
Qwen3.8-27B (OmniDocBench 1.5; tools, reasoning effort and image settings not disclosed)Qwen91.1원본 기록
Qwen3.6-27B (OmniDocBench 1.5; tools, reasoning effort and image settings not disclosed)Qwen89.4원본 기록
Opus4.6 Max (Max (table header); OmniDocBench 1.5 tools and image settings not disclosed)Anthropic86.6원본 기록
Muse Glimmer-30B (OmniDocBench 1.5; tools, reasoning effort and image settings not disclosed)Meta75.8원본 기록
MMMU-Pro / Google G31 (2026-02-19)5 개 성적

제공사 발표 모델 성적Google DeepMind (Gemini 3.1 Pro release comparison)

도구 없이 이미지 등 여러 형태의 정보를 이해하고 추론합니다.

%높을수록 좋음
MMMU-Pro / Google G31 (2026-02-19) / %
모델과 실행 구성성적출처
Gemini 3 Pro (Thinking (High); No tools; Standard (10 options) and Vision average)Google81%원본 기록
Gemini 3.1 Pro (Thinking (High); No tools; Standard (10 options) and Vision average)Google80.5%원본 기록
GPT-5.2 (Thinking (xhigh); No tools; Standard (10 options) and Vision average)OpenAI79.5%원본 기록
Claude Sonnet 4.6 (Thinking (Max); No tools; Standard (10 options) and Vision average)Anthropic74.5%원본 기록
Claude Opus 4.6 (Thinking (Max); No tools; Standard (10 options) and Vision average)Anthropic73.9%원본 기록
MRCR v2, 128k / Google G31 (2026-02-19)5 개 성적

제공사 발표 모델 성적Google DeepMind (Gemini 3.1 Pro release comparison)

긴 문맥에서 지정 정보 8개를 찾는 평가의 128k까지 누적 평균입니다.

%높을수록 좋음
MRCR v2, 128k / Google G31 (2026-02-19) / %
모델과 실행 구성성적출처
Gemini 3.1 Pro (Thinking (High); MRCR v2; 8 needles; 128k cumulative average)Google84.9%원본 기록
Claude Sonnet 4.6 (Thinking (Max); MRCR v2; 8 needles; 128k cumulative average)Anthropic84.9%원본 기록
Claude Opus 4.6 (Thinking (Max); MRCR v2; 8 needles; 128k cumulative average)Anthropic84%원본 기록
GPT-5.2 (Thinking (xhigh); MRCR v2; 8 needles; 128k cumulative average)OpenAI83.8%원본 기록
Gemini 3 Pro (Thinking (High); MRCR v2; 8 needles; 128k cumulative average)Google77%원본 기록
MRCR v2, 1M / Google G31 (2026-02-19)2 개 성적

제공사 발표 모델 성적Google DeepMind (Gemini 3.1 Pro release comparison)

1M 토큰 길이에서 지정 정보 8개를 찾는 평가입니다.

%높을수록 좋음
  1. Gemini 3.1 Pro26.3%
  2. Gemini 3 Pro26.3%
MRCR v2, 1M / Google G31 (2026-02-19) / %
모델과 실행 구성성적출처
Gemini 3.1 Pro (Thinking (High); MRCR v2; 8 needles; 1M pointwise)Google26.3%원본 기록
Gemini 3 Pro (Thinking (High); MRCR v2; 8 needles; 1M pointwise)Google26.3%원본 기록
CharXiv Reasoning / Google G38 (2026-09-02)6 개 성적

제공사 발표 모델 성적Google DeepMind (Gemini 3.8 Flash release comparison)

도구 없이 복잡한 차트의 정보를 종합합니다.

%높을수록 좋음
CharXiv Reasoning / Google G38 (2026-09-02) / %
모델과 실행 구성성적출처
Gemini 3.8 Flash (Default sampling; exact thinking level unspecified; No tools)Google86.2%원본 기록
GPT-5.6 Terra (Maximum reasoning preferred; otherwise best available reported setting; No tools)OpenAI85.9%원본 기록
GPT-5.6 Sol (Exact reasoning level unspecified; No tools)OpenAI85.8%원본 기록
Gemini 3.7 Flash (Exact thinking level unspecified; No tools)Google84.5%원본 기록
Claude Opus 5 (Exact thinking level unspecified; No tools)Anthropic83.7%원본 기록
Claude Sonnet 5 (Maximum thinking preferred; otherwise best available reported setting; No tools)Anthropic70.1%원본 기록

검색 모델

MTEB / OpenAI embeddings report3 개 성적

검색 모델 성적OpenAI (developer documentation)

텍스트를 검색용 수치로 바꾸는 임베딩 모델 평가입니다. 생성형 모델 추론과 구분합니다.

%높을수록 좋음
MTEB / OpenAI embeddings report / %
모델과 실행 구성성적출처
text-embedding-3-large (Published MTEB result; evaluation dimensions unspecified)OpenAI64.6%원본 기록
text-embedding-3-small (Published MTEB result; evaluation dimensions unspecified)OpenAI62.3%원본 기록
text-embedding-ada-002 (Published MTEB result; evaluation dimensions unspecified)OpenAI61%원본 기록

시간과 비용

Intelligence Index / Cost per task15 개 성적

기반 모델 성적Artificial Analysis

평가 과제당 가중 평균 비용

USD낮을수록 좋음
Intelligence Index / Cost per task / USD
모델과 실행 구성성적출처
GLM 5.3 Flash (GLM-5.3-Flash)Z AI0.18원본 기록
Muse Spark 1.2 (xhigh)Meta0.55원본 기록
Gemini 3.8 Flash (high)Google0.74원본 기록
GPT-5.6 Terra (max)OpenAI0.81원본 기록
Muse Spark 1.3 (max)Meta0.96원본 기록
Qwen3.8 2.4T A95BAlibaba1.1원본 기록
Qwen3.8 MaxAlibaba1.19원본 기록
GPT-5.6 Sol (max)OpenAI1.25원본 기록
Grok 4.6 (high)SpaceXAI1.25원본 기록
GLM-5.3 (max)Z AI1.26원본 기록
Kimi K3 (max)Kimi1.58원본 기록
GPT-6 Astra (max)OpenAI2.57원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic4.21원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic5.62원본 기록
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic6.12원본 기록
Output speed15 개 성적

기반 모델 성적Artificial Analysis

출력 생성 속도

tokens/s높을수록 좋음
Output speed / tokens/s
모델과 실행 구성성적출처
Gemini 3.8 Flash (high)Google280.76원본 기록
Muse Spark 1.2 (xhigh)Meta269.24원본 기록
Muse Spark 1.3 (max)Meta233.15원본 기록
GPT-5.6 Terra (max)OpenAI113.93원본 기록
GLM-5.3 (max)Z AI83.31원본 기록
GPT-5.6 Sol (max)OpenAI75.78원본 기록
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic68.22원본 기록
GPT-6 Astra (max)OpenAI64.26원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic62.97원본 기록
Grok 4.6 (high)SpaceXAI59.89원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic55.95원본 기록
GLM 5.3 Flash (GLM-5.3-Flash)Z AI50.63원본 기록
Kimi K3 (max)Kimi41.61원본 기록
Qwen3.8 2.4T A95BAlibaba40.23원본 기록
Qwen3.8 MaxAlibaba40.15원본 기록
API input price15 개 성적

기반 모델 성적Artificial Analysis

입력 100만 토큰 가격

USD / 1M낮을수록 좋음
API input price / USD / 1M
모델과 실행 구성성적출처
GLM 5.3 Flash (GLM-5.3-Flash)Z AI0.15원본 기록
Gemini 3.8 Flash (high)Google0.75원본 기록
Muse Spark 1.3 (max)Meta1.25원본 기록
Muse Spark 1.2 (xhigh)Meta1.25원본 기록
GLM-5.3 (max)Z AI1.4원본 기록
Grok 4.6 (high)SpaceXAI2원본 기록
Qwen3.8 MaxAlibaba2원본 기록
GPT-5.6 Terra (max)OpenAI2원본 기록
Qwen3.8 2.4T A95BAlibaba2원본 기록
Kimi K3 (max)Kimi3원본 기록
GPT-5.6 Sol (max)OpenAI4원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic5원본 기록
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic10원본 기록
GPT-6 Astra (max)OpenAI10원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic10원본 기록
API output price15 개 성적

기반 모델 성적Artificial Analysis

출력 100만 토큰 가격

USD / 1M낮을수록 좋음
API output price / USD / 1M
모델과 실행 구성성적출처
GLM 5.3 Flash (GLM-5.3-Flash)Z AI0.5원본 기록
Gemini 3.8 Flash (high)Google3.75원본 기록
Muse Spark 1.3 (max)Meta4.25원본 기록
Muse Spark 1.2 (xhigh)Meta4.25원본 기록
GLM-5.3 (max)Z AI4.4원본 기록
Grok 4.6 (high)SpaceXAI6원본 기록
Qwen3.8 MaxAlibaba6원본 기록
Qwen3.8 2.4T A95BAlibaba6원본 기록
GPT-5.6 Terra (max)OpenAI12원본 기록
Kimi K3 (max)Kimi15원본 기록
GPT-5.6 Sol (max)OpenAI20원본 기록
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic25원본 기록
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic50원본 기록
GPT-6 Astra (max)OpenAI50원본 기록
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic50원본 기록

본문 출처와 더 읽을 자료

  1. Anthropic. Demystifying evals for AI agents
  2. OpenAI. GDPval
  3. OpenAI. BrowseComp
  4. Parallel. Deep Research price-performance
  5. LangChain. Improving Deep Agents with harness engineering
  6. LangChain. Deep Agents source code
  7. DeepResearch Bench II. Evaluation code
  8. ResearchRubrics. Evaluation code
  9. Agents’ Last Exam. Tasks and evaluation