Skip to content

Leaderboard

SciCode

30 models scored · metric: percent correct. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the heavy reasoning token profile — thinking tokens outnumber the answer several times over.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 2k in · 1.5k out · 8k reasoning · 0% cache hit · 12k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. SciCode is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · heavy reasoning profile
  1. 1.Gemini 2.5 Flash-Lite16,783
  2. 2.Step 3.5 Flash13,333
  3. 3.Step 3.5 Flash 260312,623
  4. 4.Llama-3.3-70B-Instruct11,379
  5. 5.Qwen3-Coder 30B-A3B Instruct11,142
30 models
#
🥇Fugusakana
60.1
1M
🥈Fugu Ultrasakana
58.7
$5$30$0.2951991.05M
🥉Qwen3.7 Maxalibaba
53.5
$0.825$2.48$0.0252,1261.06M
4Gemini 2.5 Progoogle
42.8
$0.87$7$0.0686271.07M
5GPT-5-Codexopenai
40.9
$1.07$8.50$0.083493400K
6Step 3.5 Flashstepfun
40.4
$0.09$0.30.303¢13,333262K
7Step 3.7 Flashstepfun
40
$0.185$1.11$0.0113,665262K
8Gemini 2.5 Flashgoogle
39.4
$0.09$0.710.693¢5,6901.07M
9Step 3.5 Flash 2603stepfun
38.5
$0.1$0.30.305¢12,623262K
10GLM-4.6zhipuai
38.4
$0.286$1.14$0.0113,362205K
11Qwen3 Maxalibaba
38.3
$0.36$1.43$0.0142,677262K
12Mistral Small 4mistral
38
$0.143$0.5680.568¢6,688262K
13Mistral Large 3mistral
36.2
$0.5$1.50$0.0152,374262K
14DeepSeek-R1deepseek
35.7
$0.4$1.70$0.0172,106164K
15GLM-4.5zhipuai
34.8
$0.286$1.14$0.0113,047131K
16GPT-4o (2024-11-20)openai
33.3
$2.50$10$0.100333128K
17Devstral 2mistral
33.1
$0.4$2$0.0201,672262K
18Mistral Medium 3mistral
33.1
$0.4$2$0.0201,672131K
19GPT-4o (2024-08-06)openai
33.1
$2.50$10$0.100331128K
20GPT-4 Turboopenai
31.9
$9$27$0.274116128K

Score reports: console.sakana.ai · qwen.ai · openrouter.ai · openrouter.ai · openrouter.ai · openrouter.ai — full source URLs are linked on each model page.

Cheapest scorer: Llama-3.3-70B-Instruct at $0.05 input /M.