System Card
← leaderboard
xAI
xAI
model systemcard
Grok 4.5
#9
of 27 · general
500k
context
Proprietary
non-commercial
2026-07-16
released
2026-02-01
knowledge cutoff
Standing
dot = this model · shaded = the field's middle 60% · track = full field
General
57.7
#9
/27
Math
54.6
#10
/27
Coding
48.0
#12
/27
Health
51.6
#13
/26
Engineering
56.2
#8
/17
Analyst
40.0
#7
/11
Legal
44.7
#16
/26
Instruct
49.8
#12
/27
Agents
65.5
#6
/27
Professional
48.5
#13
/27
Multimodal
33.6
#13
/19
Hallucination
30.9
#13
/27
Speed
24.0
#19
/22
Value
59.5
#6
/27
Measured
≈ imputed · amber dot = lab-self-reported among the sources
GPQA Diamond
93.1%
#7
HLE
42.7%
#12
MMLU-Pro
89.2%
#7
LiveBench
75.8%
#9
SWE-bench V
86.6%
#6
LiveCodeBench
87.3%
#8
SciCode
54.0%
#8
CritPt
15.4%
#13
Terminal-Bench
74.7%
#11
APEX Agents
34.2%
#10
GDPval
51.3%
#11
Arena
1468
#12
AA Index
55.8
#8
Vals Index
65.3
#9
Epoch Capabilities Index
154.0
#14
Operations
$2
/
$6
$ per 1M tokens, in / out
$3
blended
56
tokens / sec
$0.12
$ per solved task
1468
arena elo · 9,998 votes
Neighbors
nearest capability fingerprints · deltas describe the neighbor
Anthropic
Claude Sonnet 5
→ stronger agent
Alibaba
Qwen3.8 Max
→ stronger professional · weaker math
GPT-5.6 Terra
→ looser instructions · pricier
Zhipu (智谱)
GLM-5.2
→ stronger math · cheaper
DeepMind
Gemini 3.6 Flash
→ stronger multimodal