Skip to content

Leaderboard

DeepSWE

13 models scored · metric: resolve rate. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the agentic token profile — tool loops re-send a growing transcript and think between calls.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 12k in · 2k out · 3k reasoning · 70% cache hit · 32k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. DeepSWE is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · agentic profile
  1. 1.DeepSeek V4 Flash 073191,000
  2. 2.GPT-5.6 Luna25,093
  3. 3.GLM-5.26,337
  4. 4.Gemini 3.6 Flash4,423
  5. 5.GPT-5.6 Terra3,595
13 models
#
🥇GPT-6 Astraopenai
74.1
$10$50$0.3122371.05M
🥈GPT-5.6 Solopenai
72.7
$2$10$0.0631,1471.05M
🥉GPT-5.6 Terraopenai
69.6
$1.50$2$0.0193,5951.05M
4Claude Opus 5anthropic
68.8
$5$25$0.1584341M
5Kimi K3moonshotai
67.5
$2$8$0.0531,2761.05M
6GPT-5.6 Lunaopenai
67.2
$0.06$0.370.268¢25,0931.05M
7Gemini 3.7 Flashgoogle
65.3
$0.75$3.75$0.0222,9481.05M
8Grok 4.5xai
62
$2$6$0.0431,4311M
9Qwen3.8 Max Previewalibaba
56.6
$2$6$0.0431,3191M
10DeepSeek V4 Flash 0731deepseek
54.4
$0.035$0.070.060¢91,0001.31M
11Muse Spark 1.1meta
53.3
$1.25$4.25$0.0291,8221.05M
12Gemini 3.6 Flashgoogle
49
$0.375$1.88$0.0114,4231.05M
13GLM-5.2zhipuai
46.2
$0.3$1.050.729¢6,3371.05M

Score reports: openai.com · openai.com · anthropic.com · kimi.com · deepmind.google · x.ai — full source URLs are linked on each model page.

Cheapest scorer: DeepSeek V4 Flash 0731 at $0.035 input /M.