Skip to content

Leaderboard

MMMU Proofficial site ↗

10 models scored · metric: accuracy. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the heavy reasoning token profile — thinking tokens outnumber the answer several times over.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 2k in · 1.5k out · 8k reasoning · 0% cache hit · 12k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. MMMU Pro is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · heavy reasoning profile
  1. 1.GPT-5.6 Luna21,568
  2. 2.Gemini 3.5 Flash7,630
  3. 3.GPT-5.57,340
  4. 4.GPT-5.4 nano6,429
  5. 5.GPT-5.6 Terra3,668
10 models
#
🥇Gemini 3.5 Flashgoogle
83.6
$0.186$1.11$0.0117,6301.05M
🥈GPT-5.6 Solopenai
83
$2$10$0.0998381.05M
🥉Qwen3.8 Max Previewalibaba
82.3
$2$6$0.0611,3491M
4GPT-5.4openai
81.2
$0.75$6$0.0591,3881.05M
5GPT-5.5openai
81.2
$0.188$1.13$0.0117,3401.05M
6GPT-5.6 Terraopenai
80.7
$1.50$2$0.0223,6681.05M
7Gemini 3.1 Pro Previewgoogle
80.5
$1$6$0.0591,3641.05M
8GPT-5.6 Lunaopenai
78.4
$0.06$0.370.363¢21,5681.05M
9GPT-5.4 miniopenai
78
$0.375$4$0.0392,013400K
10GPT-5.4 nanoopenai
69.5
$0.18$1.10$0.0116,4291.05M

Score reports: deepmind.google · openai.com · alibabacloud.com · openai.com · openai.com — full source URLs are linked on each model page.

Cheapest scorer: GPT-5.6 Luna at $0.06 input /M.