Skip to content

Leaderboard

SWE-Atlas Test Writing

10 models scored · metric: score. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the long context token profile — a whole codebase or filing per request — the case tiered pricing exists for.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 250k in · 2k out · 30% cache hit · 256k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. SWE-Atlas Test Writing is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · long context profile
  1. 1.Gemini 3 Flash Preview1,507
  2. 2.GPT-5.5791
  3. 3.Kimi K2.5343
  4. 4.GLM-5334
  5. 5.MiniMax-M2.5299
10 models
#
🥇GPT-5.4openai
44.36
$0.75$6$0.1682651.05M
🥈GPT-5.5openai
42.59
$0.188$1.13$0.0547911.05M
🥉GPT-5.3 Codexopenai
38.98
$1.60$13$0.358109400K
4Claude Opus 4.6anthropic
36.67
$4$20$0.89541.01M
5Claude Sonnet 4.6anthropic
31.76
$0.9$5.55$0.1981611M
6Gemini 3 Flash Previewgoogle
30.3
$0.07$0.43$0.0201,5071.05M
7Gemini 3.1 Pro Previewgoogle
29.84
$1$6$0.1991501.05M
8GLM-5zhipuai
28.74
$0.35$1.40$0.086334205K
9Kimi K2.5moonshotai
25.77
$0.3$1.90$0.075343262K
10MiniMax-M2.5minimax
18.6
$0.22$0.88$0.062299229K

Score reports: labs.scale.com — full source URLs are linked on each model page.

Cheapest scorer: Gemini 3 Flash Preview at $0.07 input /M.