Skip to content

Leaderboard

SWE-Bench Multilingualofficial site ↗

7 models scored · metric: resolve rate. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the agentic token profile — tool loops re-send a growing transcript and think between calls.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 12k in · 2k out · 3k reasoning · 70% cache hit · 32k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. SWE-Bench Multilingual is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · agentic profile
  1. 1.Laguna XS 2.153,656
  2. 2.Nemotron 3 Ultra 550B A55B36,011
  3. 3.LongCat-2.010,078
  4. 4.Qwen3.7 Max4,298
  5. 5.Grok 4.51,801
7 models
#
🥇Claude Opus 5anthropic
89.5
$5$25$0.1585651M
🥈Qwen3.7 Maxalibaba
78.3
$0.825$2.48$0.0184,2981.06M
🥉Claude Sonnet 5anthropic
78.3
$1.44$7.20$0.0561,4011M
4Grok 4.5xai
78
$2$6$0.0431,8011M
5LongCat-2.0meituan
77.3
$0.3$1.200.767¢10,0781.05M
6Nemotron 3 Ultra 550B A55Bnvidia
67.7
$0.1$0.10.188¢36,0111.05M
7Laguna XS 2.1poolside
63.1
$0.06$0.120.118¢53,656262K

Score reports: anthropic.com · qwen.ai · anthropic.com · x.ai · github.com · huggingface.co — full source URLs are linked on each model page.

Cheapest scorer: Laguna XS 2.1 at $0.06 input /M.