Skip to content

Leaderboard

SWE-Atlas Refactoring

11 models scored · metric: score. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the long context token profile — a whole codebase or filing per request — the case tiered pricing exists for.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 250k in · 2k out · 30% cache hit · 256k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. SWE-Atlas Refactoring is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · long context profile
  1. 1.GPT-5.5832
  2. 2.Gemini 3 Flash Preview497
  3. 3.MiniMax-M2.5314
  4. 4.GLM-5282
  5. 5.Kimi K2.5279
11 models
#
🥇Claude Opus 4.7anthropic
48.57
$4$20$0.89554.31.05M
🥈GPT-5.5openai
44.79
$0.188$1.13$0.0548321.05M
🥉GPT-5.4openai
44.29
$0.75$6$0.1682641.05M
4GPT-5.3 Codexopenai
42.38
$1.60$13$0.358118400K
5Claude Opus 4.6anthropic
35.58
$4$20$0.89539.81M
6Gemini 3.1 Pro Previewgoogle
33.81
$1$6$0.1991701.05M
7Claude Sonnet 4.6anthropic
32.21
$0.9$5.55$0.1981631M
8GLM-5zhipuai
24.24
$0.35$1.40$0.086282205K
9Kimi K2.5moonshotai
20.95
$0.3$1.90$0.075279262K
10MiniMax-M2.5minimax
19.52
$0.22$0.88$0.062314229K
11Gemini 3 Flash Previewgoogle
10
$0.07$0.43$0.0204971.05M

Score reports: labs.scale.com — full source URLs are linked on each model page.

Cheapest scorer: Gemini 3 Flash Preview at $0.07 input /M.