Skip to content

Leaderboard

FrontierMathofficial site ↗

8 models scored · metric: accuracy. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the heavy reasoning token profile — thinking tokens outnumber the answer several times over.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 2k in · 1.5k out · 8k reasoning · 0% cache hit · 12k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. FrontierMath is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · heavy reasoning profile
  1. 1.GPT-5.6 Luna21,623
  2. 2.GPT-5.54,673
  3. 3.GPT-5.6 Terra3,859
  4. 4.GPT-5.6 Sol899
  5. 5.GPT-5.4814
8 models
#
🥇GPT-6 Astraopenai
97.6
$10$50$0.4951971.05M
🥈GPT-5.6 Solopenai
89
$2$10$0.0998991.05M
🥉GPT-5.6 Terraopenai
84.9
$1.50$2$0.0223,8591.05M
4GPT-5.6 Lunaopenai
78.6
$0.06$0.370.363¢21,6231.05M
5GPT-5.5 Proopenai
52.4
$27.27$163.64$1.60932.61.05M
6GPT-5.5openai
51.7
$0.188$1.13$0.0114,6731.05M
7GPT-5.4 Proopenai
50
$27$160$1.57431.81.05M
8GPT-5.4openai
47.6
$0.75$6$0.0598141.05M

Score reports: openai.com · openai.com · openai.com — full source URLs are linked on each model page.

Cheapest scorer: GPT-5.6 Luna at $0.06 input /M.