System Card
← leaderboard
Meta AI
model systemcard
Muse Spark
#5
of 27 · general
1.0 and 1.1 collapsed
1M
context
Proprietary
non-commercial
2026-07-09
released
paper ↗
Standing
dot = this model · shaded = the field's middle 60% · track = full field
General
73.2
#5
/27
Math
43.9
#13
/27
Coding
69.7
#5
/27
Health
87.5
#1
/26
Engineering
75.0
#5
/17
Legal
82.4
#1
/26
Instruct
59.8
#9
/27
Agents
73.0
#4
/27
Professional
83.0
#4
/27
Multimodal
58.4
#6
/19
Hallucination
51.1
#5
/27
Value
69.5
#2
/27
Measured
≈ imputed · amber dot = lab-self-reported among the sources
GPQA Diamond
90.1%
#19
HLE
45.5%
#9
MMLU-Pro
88.7%
#9
AIME 2025
96.9%
#3
LiveBench
75.3%
#10
SWE-bench V
79.7%
#8
LiveCodeBench
85.9%
#13
SciCode
57.3%
#4
CritPt
16.4%
#12
Terminal-Bench
69.3%
#14
APEX Agents
41.9%
#3
GDPval
56.4%
#6
Arena
1491
#3
AA Index
56.8
#6
Vals Index
68.4
#7
Epoch Capabilities Index
154.8
#11
Operations
$1.25
/
$4.25
$ per 1M tokens, in / out
$2
blended
$0.19
$ per solved task
1491
arena elo · 9,941 votes
Neighbors
nearest capability fingerprints · deltas describe the neighbor
Moonshot (月之暗面)
Kimi K3
→ pricier
GPT-5.6 Sol
→ stronger math · pricier
Anthropic
Claude Opus 5
→ stronger math · pricier
Zhipu (智谱)
GLM-5.2
→ stronger math
Anthropic
Claude Fable 5
→ stronger math · pricier