System Card
← leaderboard
Alibaba
Alibaba
model systemcard
Qwen3.8 Max
#12
of 27 · general
3.5–3.7 Max previews collapsed
1.0T
params
36T
training tokens
Proprietary
non-commercial
2025-12-15
released
paper ↗
repo ↗
Standing
dot = this model · shaded = the field's middle 60% · track = full field
General
50.8
#12
/27
Math
32.6
#17
/27
Coding
57.0
#10
/27
Health
37.3
#16
/26
Engineering
68.8
#6
/17
Legal
43.1
#18
/26
Instruct
63.0
#7
/27
Agents
62.6
#9
/27
Professional
71.2
#6
/27
Multimodal
28.7
#14
/19
Hallucination
39.8
#10
/27
Speed
21.2
#20
/22
Value
51.2
#10
/27
Measured
≈ imputed · amber dot = lab-self-reported among the sources
GPQA Diamond
92.5%
#12
HLE
42.2%
#13
MMLU-Pro
89.5%
#5
AIME 2025
81.6%
#7
LiveBench
73.7%
#12
SWE-bench V
77.3%
#19
LiveCodeBench
87.1%
#10
SciCode
52.9%
#12
CritPt
13.4%
#14
Terminal-Bench
69.7%
#13
GDPval
61.9%
#3
Arena
1475
#10
AA Index
58.1
#5
Vals Index
57.5
#15
Epoch Capabilities Index
156.1
#8
Operations
$2
/
$6
$ per 1M tokens, in / out
$3
blended
51
tokens / sec
$0.18
$ per solved task
1475
arena elo · 3,697 votes
Neighbors
nearest capability fingerprints · deltas describe the neighbor
xAI
Grok 4.5
→ weaker professional · stronger math
DeepMind
Gemini 3.6 Flash
→ stronger multimodal
Moonshot (月之暗面)
Kimi K3
→ broader · pricier
Anthropic
Claude Sonnet 5
→ stronger math · looser instructions
GPT-5.6 Terra
→ stronger math · pricier