Leaderboard
SWE-Atlas Codebase QnA
14 models scored · metric: score. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the long context token profile — a whole codebase or filing per request — the case tiered pricing exists for.
A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 250k in · 2k out · 30% cache hit · 256k context.
Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. SWE-Atlas Codebase QnA is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.
- 1.GPT-5.51,502
- 2.DeepSeek V4 Pro944
- 3.GLM-5.1822
- 4.Kimi K2.6768
- 5.GPT-5.4435
Score reports: artificialanalysis.ai · labs.scale.com — full source URLs are linked on each model page.
Cheapest scorer: Gemini 3 Flash Preview at $0.07 input /M.