Skip to content

Leaderboard

FrontierCode

6 models scored · metric: pass rate. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the heavy reasoning token profile — thinking tokens outnumber the answer several times over.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 2k in · 1.5k out · 8k reasoning · 0% cache hit · 12k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. FrontierCode is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · heavy reasoning profile
  1. 1.Gemini 3.7 Flash1,174
  2. 2.Claude Opus 4.8637
  3. 3.Claude Sonnet 5544
  4. 4.Claude Opus 5216
  5. 5.Claude Fable 5161
6 models
#
🥇GPT-6 Astraopenai
64.5
$10$50$0.4951301.05M
🥈Claude Opus 5anthropic
53.4
$5$25$0.2482161M
🥉Gemini 3.7 Flashgoogle
43.6
$0.75$3.75$0.0371,1741.05M
4Claude Sonnet 5anthropic
38.8
$1.44$7.20$0.0715441M
5Claude Fable 5anthropic
29.3
$3$18.50$0.1821611M
6Claude Opus 4.8anthropic
13.4
$0.425$2.13$0.0216371.05M

Score reports: openai.com · anthropic.com · deepmind.google · anthropic.com · anthropic.com — full source URLs are linked on each model page.

Cheapest scorer: Claude Opus 4.8 at $0.425 input /M.