Leaderboard
SWE-Bench Multilingualofficial site ↗
7 models scored · metric: resolve rate. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the agentic token profile — tool loops re-send a growing transcript and think between calls.
A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 12k in · 2k out · 3k reasoning · 70% cache hit · 32k context.
Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. SWE-Bench Multilingual is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.
- 1.Laguna XS 2.153,656
- 2.Nemotron 3 Ultra 550B A55B36,011
- 3.LongCat-2.010,078
- 4.Qwen3.7 Max4,298
- 5.Grok 4.51,801
Score reports: anthropic.com · qwen.ai · anthropic.com · x.ai · github.com · huggingface.co — full source URLs are linked on each model page.
Cheapest scorer: Laguna XS 2.1 at $0.06 input /M.