Skip to main content
AI Socratic

Leaderboard

Socratic Score = mean of normalized benchmarks. LM Arena: (ELO - 1000) / 400 × 100. Vending-Bench: balance / $10k × 100. SWE-bench, ARC-AGI, and HLE are used as-is (0-100%).

AI Socratic Leaderboard

Scores across benchmarks

Updated Jul 20, 2026, 10:00 UTC

#ModelScoreLM Arena ↗SWE-bench ↗ARC-AGI-2 ↗HLE ↗Vending ↗Prediction ↗Vibe Bench
🥇
AnthropicClaude Opus 4.6
85.05/6
150475.6%69.2%($3.47)-$8,017.59+10521.7%60%
🥈
ZhipuGLM 5
69.74/6
-72.8%22.8%($0.25)-$8,313.78+108.2%-
🥉
OpenAIGPT 5.4
65.13/6
--83.3%($16.41)-$6,144.18+1.2%15%
4
OpenAIGPT 5.2
60.83/6
-72.8%72.9%($38.99)---26.8%15%
5
GoogleGemini 3 Pro
59.35/6
148669.6%54.0%($30.57)38.3%--31.0%18%
6
AnthropicClaude Opus 4.5
50.43/6
-76.8%37.6%($2.40)---26.6%60%
Location

Vibe Bench

Community favorites · Socratic Jun · 40 responses

Latest event: Socratic Jun · Jun 23, 2026

AnthropicClaude
24
AnthropicClaude Code
13
OpenAICodex
9
Open Source
8
GoogleGemini
7
Other Tools
7
OpenAIChatGPT
6
CursorCursor
3
xAIGrok
1
PerplexityPerplexity
1
WindsurfWindsurf
0

Real-world software engineering tasks

Updated Jul 20, 2026, 10:00 UTC

#Model% Resolved
🥇
AnthropicClaude 4.5 Opus (high reasoning)
76.8%
🥈
GoogleGemini 3 Flash (high reasoning)
75.8%
🥉
MiniMaxMiniMax M2.5 (high reasoning)
75.8%
4
AnthropicClaude Opus 4.6
75.6%
5
OpenAIGPT-5-2 Codex
72.8%
6
Zhipu AIGLM-5 (high reasoning)
72.8%
7
OpenAIGPT-5-2 (high reasoning)
72.8%
8
OpenAIGPT 5.2 Codex
72.8%
9
AnthropicClaude 4.5 Sonnet (high reasoning)
71.4%
10
Kimi K2.5 (high reasoning)
70.8%
ARC-AGI-2 Semi-Private

Abstract reasoning capabilities

Updated Jul 20, 2026, 10:00 UTC

#ModelScore
🥇
OpenAIGPT-5.6 Sol (Max)
92.5%
🥈
OpenAIGPT-5.6 Sol (xHigh)
90.0%
🥉
OpenAIGPT-5.6 Sol (High)
85.4%
4
OpenAIGPT-5.5 (xHigh)
85.0%
5
GoogleGemini 3 Deep Think (2/26)
84.6%
6
OpenAIGPT-5.5 Pro (High)
84.6%
7
OpenAIGPT-5.5 Pro (xHigh)
84.2%
8
OpenAIGPT-5.6 Terra (Max)
83.9%
9
OpenAIGPT-5.4 Pro (xHigh)
83.3%
10
OpenAIGPT-5.5 (High)
83.3%

Expert-level reasoning across disciplines

Updated Jul 20, 2026, 10:00 UTC

#ModelAccuracy
🥇
GoogleGemini 3 Pro
38.3%
🥈
OpenAIGPT-5
25.3%
🥉
xAIGrok 4
24.5%
4
GoogleGemini 2.5 Pro
21.6%
5
OpenAIGPT-5-mini
19.4%
6
AnthropicClaude 4.5 Sonnet
13.7%
7
GoogleGemini 2.5 Flash
12.1%
8
DeepSeekDeepSeek-R1*
8.5%
9
OpenAIo1
8.0%
10
OpenAIGPT-4o
2.7%

Crowdsourced human evaluations

Updated Jul 20, 2026, 10:00 UTC

#ModelScoreVotes
🥇
Anthropicclaude-fable-5
15070
🥈
Anthropicclaude-opus-4-6-thinking
15040
🥉
Anthropicclaude-opus-4-7-thinking
15030
4
Anthropicclaude-opus-4-6
14980
5
Anthropicclaude-opus-4-7
14940
6
muse-spark-1.1
14930
7
muse-spark
14870
8
Googlegemini-3-pro
14860
9
kimi-k3
14860
10
OpenAIgpt-5.6-sol-xhigh
14860

Long-term agentic coherence

Updated Jul 20, 2026, 10:00 UTC

#ModelBalance
🥇
AnthropicClaude Opus 4.7 Claude Opus 4.7
$10,936.76
🥈
OpenAIGPT-5.6 Sol GPT-5.6 Sol New
$9,619.37
🥉
Zhipu AIGLM-5.2 GLM-5.2
$8,313.78
4
AnthropicClaude Opus 4.6 Claude Opus 4.6
$8,017.59
5
OpenAIGPT-5.5 GPT-5.5
$7,523.84
6
OpenAIGPT-5.6 Terra GPT-5.6 Terra New
$7,343.21
7
AnthropicClaude Sonnet 4.6 Claude Sonnet 4.6
$7,204.14
8
AnthropicClaude Sonnet 5 Claude Sonnet 5
$6,377.7
9
Kimi K2.6 Kimi K2.6
$6,204.57
10
OpenAIGPT-5.4 GPT-5.4
$6,144.18

AI prediction market performance

Updated Jul 20, 2026, 10:00 UTC

#AgentReturnSharpe
🥇
AnthropicClaude Opus 4.6
10521.7%-0.01
🥈
Zhipu AIGLM 5
108.2%0.03
🥉
GoogleGemini 3.1 Pro
74.3%0.03
4
OpenAIGPT 5.4
1.2%0.02
5
Zhipu AIGLM 4.7
-15.4%-0.12
6
Mystery Model Alpha
-20.0%-0.05
7
AnthropicClaude Opus 4.5
-26.6%-0.10
8
OpenAIGPT 5.2
-26.8%-0.09
9
xAIGrok 4.1
-30.8%-0.07
10
GoogleGemini 3 Pro
-31.0%-0.11

Stay Updated

Get the latest AI insights delivered to your inbox. No spam, unsubscribe anytime.