System Card
← leaderboard
DeepMind
Google DeepMind
model systemcard
Gemini 3.5 Flash
#15
of 27 · general
high/medium effort collapsed
1M
context
Proprietary
non-commercial
2026-05-19
released
2026-01-31
knowledge cutoff
Standing
dot = this model · shaded = the field's middle 60% · track = full field
General
47.4
#15
/27
Math
50.6
#11
/27
Coding
33.9
#18
/27
Health
30.9
#19
/26
Cyber
46.2
#8
/14
Legal
6.2
#26
/26
Instruct
71.1
#3
/27
Agents
30.1
#18
/27
Professional
51.6
#12
/27
Multimodal
62.2
#5
/19
Hallucination
49.5
#6
/27
Speed
92.4
#2
/22
Value
44.8
#15
/27
Measured
≈ imputed · amber dot = lab-self-reported among the sources
GPQA Diamond
92.7%
#9
HLE
40.8%
#15
MMLU-Pro
89.5%
#4
LiveBench
74.8%
#11
SWE-bench V
79.1%
#10
LiveCodeBench
87.6%
#7
SciCode
53.1%
#11
CritPt
12.0%
#16
Terminal-Bench
75.2%
#10
Arena
1476
#9
AA Index
46.7
#15
Vals Index
62.8
#12
Epoch Capabilities Index
154.6
#12
Operations
$1.5
/
$9
$ per 1M tokens, in / out
$3.38
blended
174
tokens / sec
$0.25
$ per solved task
1476
arena elo · 10,011 votes
Series
DeepMind's models on the board, by release
Gemini 3.1 Pro
2026-02-19 · 55
→
Gemini 3.5 Flash
2026-05-19 · 47
→
Gemini 3.6 Flash
2026-07-21 · 49
Neighbors
nearest capability fingerprints · deltas describe the neighbor
DeepMind
Gemini 3.6 Flash
→ near-twin
DeepSeek V4 Pro
→ looser instructions · cheaper
GPT-5.5
→ stronger coding
DeepMind
Gemini 3.1 Pro
→ near-twin
Xiaomi
MiMo V2.5 Pro
→ cheaper