System Card
← leaderboard
Anthropic
Anthropic
model systemcard
Claude Sonnet 5
#7
of 27 · general
adaptive-reasoning efforts collapsed
1M
context
Proprietary
non-commercial
2026-06-30
released
paper ↗
Standing
dot = this model · shaded = the field's middle 60% · track = full field
General
60.3
#7
/27
Math
61.7
#8
/27
Coding
62.8
#7
/27
Health
62.0
#8
/26
Engineering
62.5
#7
/17
Analyst
60.0
#5
/11
Legal
36.8
#19
/26
Instruct
36.1
#15
/27
Agents
82.8
#2
/27
Professional
58.6
#8
/27
Multimodal
41.8
#10
/19
Hallucination
23.0
#18
/27
Speed
56.9
#9
/22
Value
49.9
#12
/27
Measured
≈ imputed · amber dot = lab-self-reported among the sources
GPQA Diamond
90.5%
#15
HLE
49.3%
#6
MMLU-Pro
87.5%
#12
LiveBench
76.0%
#8
SWE-bench V
82.4%
#7
LiveCodeBench
82.4%
#17
SciCode
53.6%
#10
CritPt
16.9%
#10
Terminal-Bench
80.4%
#6
APEX Agents
32.5%
#12
GDPval
54.9%
#7
Arena
1460
#15
AA Index
55.3
#9
Vals Index
68.6
#6
Epoch Capabilities Index
155.5
#9
Operations
$2
/
$10
$ per 1M tokens, in / out
$4
blended
73
tokens / sec
$0.51
$ per solved task
1460
arena elo · 15,622 votes
Series
Anthropic's models on the board, by release
Claude Fable 5
2026-06-09 · 96
→
Claude Sonnet 5
2026-06-30 · 60
→
Claude Opus 5
2026-07-24 · 89
Neighbors
nearest capability fingerprints · deltas describe the neighbor
xAI
Grok 4.5
→ weaker agent
Zhipu (智谱)
GLM-5.2
→ weaker agent · cheaper
GPT-5.6 Terra
→ weaker agent
Moonshot (月之暗面)
Kimi K3
→ stronger professional · pricier
Alibaba
Qwen3.8 Max
→ weaker math · follows instructions better