Skip to content

Leaderboard

SWE-Bench Proofficial site ↗

51 models scored · metric: resolve rate. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the agentic token profile — tool loops re-send a growing transcript and think between calls.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 12k in · 2k out · 3k reasoning · 70% cache hit · 32k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. SWE-Bench Pro is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · agentic profile
  1. 1.Laguna XS 2.140,476
  2. 2.MiniMax-M2.725,841
  3. 3.Qwen3.8 27B23,514
  4. 4.GPT-5.6 Luna23,413
  5. 5.Gemini 3 Flash Preview11,114
51 models
#
🥇Claude Fable 5anthropic
80.3
$3$18.50$0.1117221M
🥈Claude Opus 5anthropic
79.2
$5$25$0.1585001M
🥉Fugu Ultrasakana
73.7
$5$30$0.1814071.05M
4Claude Opus 4.8anthropic
69.2
$0.425$2.13$0.0164,1961.05M
5Qwen3.8 Max Previewalibaba
67.7
$2$6$0.0431,5781M
6Grok 4.5xai
64.7
$2$6$0.0431,4941M
7GPT-5.6 Solopenai
64.6
$2$10$0.0631,0191.05M
8GPT-5.6 Terraopenai
63.4
$1.50$2$0.0193,2751.05M
9Claude Sonnet 5anthropic
63.2
$1.44$7.20$0.0561,1311M
10GPT-5.6 Lunaopenai
62.7
$0.06$0.370.268¢23,4131.05M
11GLM-5.2zhipuai
62.1
$0.3$1.050.729¢8,5191.05M
12Qwen3.8 27Balibaba
61.7
$0.1$0.40.262¢23,5141M
13Muse Spark 1.1meta
61.5
$1.25$4.25$0.0292,1021.05M
14Qwen3.7 Maxalibaba
60.6
$0.825$2.48$0.0183,3261.06M
15LongCat-2.0meituan
59.5
$0.3$1.200.767¢7,7571.05M
16GPT-5.4openai
59.1
$0.75$6$0.0351,7041.05M
17MiniMax-M3minimax
59
$0.225$0.90.609¢9,6831.05M
18Fugusakana
59
1M
19Gemini 3.6 Flashgoogle
58.7
$0.375$1.88$0.0115,2991.05M
20MiMo-V2.5-Proxiaomi
57.2
$0.4$0.80.619¢9,2481.05M

Score reports: anthropic.com · anthropic.com · console.sakana.ai · anthropic.com · alibabacloud.com · x.ai — full source URLs are linked on each model page.

Cheapest scorer: GPT-5.6 Luna at $0.06 input /M.