Skip to content

Leaderboard

SWE-Atlas Codebase QnA

14 models scored · metric: score. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the long context token profile — a whole codebase or filing per request — the case tiered pricing exists for.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 250k in · 2k out · 30% cache hit · 256k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. SWE-Atlas Codebase QnA is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · long context profile
  1. 1.GPT-5.51,502
  2. 2.DeepSeek V4 Pro944
  3. 3.GLM-5.1822
  4. 4.Kimi K2.6768
  5. 5.GPT-5.4435
14 models
#
🥇Claude Opus 4.7anthropic
81
$4$20$0.89590.51.05M
🥈GPT-5.5openai
80.8
$0.188$1.13$0.0541,5021.05M
🥉GLM-5.1zhipuai
73.2
$0.45$2.15$0.089822205K
4GPT-5.4openai
72.9
$0.75$6$0.1684351.05M
5Claude Opus 4.6anthropic
71.9
$4$20$0.89580.31M
6Claude Sonnet 4.6anthropic
70.3
$0.9$5.55$0.1983551M
7DeepSeek V4 Prodeepseek
67.8
$0.35$0.8$0.0729441.05M
8Kimi K2.6moonshotai
59.8
$0.275$1.10$0.078768262K
9Gemini 3.1 Pro Previewgoogle
45.6
$1$6$0.1992291.05M
10GPT-5.3 Codexopenai
32.6
$1.60$13$0.35891.1400K
11GLM-5zhipuai
20.5
$0.35$1.40$0.086239205K
12Kimi K2.5moonshotai
13.1
$0.3$1.90$0.075175262K
13MiniMax-M2.5minimax
10.3
$0.22$0.88$0.062165229K
14Gemini 3 Flash Previewgoogle
8.2
$0.07$0.43$0.0204081.05M

Score reports: artificialanalysis.ai · labs.scale.com — full source URLs are linked on each model page.

Cheapest scorer: Gemini 3 Flash Preview at $0.07 input /M.