Leaderboard
SciCode
30 models scored · metric: percent correct. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the heavy reasoning token profile — thinking tokens outnumber the answer several times over.
A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 2k in · 1.5k out · 8k reasoning · 0% cache hit · 12k context.
Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. SciCode is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.
- 1.Gemini 2.5 Flash-Lite16,783
- 2.Step 3.5 Flash13,333
- 3.Step 3.5 Flash 260312,623
- 4.Llama-3.3-70B-Instruct11,379
- 5.Qwen3-Coder 30B-A3B Instruct11,142
Score reports: console.sakana.ai · qwen.ai · openrouter.ai · openrouter.ai · openrouter.ai · openrouter.ai — full source URLs are linked on each model page.
Cheapest scorer: Llama-3.3-70B-Instruct at $0.05 input /M.