Skip to content

Leaderboard

ARC-AGI-2official site ↗

6 models scored · metric: accuracy. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the heavy reasoning token profile — thinking tokens outnumber the answer several times over.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 2k in · 1.5k out · 8k reasoning · 0% cache hit · 12k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. ARC-AGI-2 is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · heavy reasoning profile
  1. 1.GPT-5.57,684
  2. 2.Gemini 3.5 Flash6,581
  3. 3.Gemini 3.1 Pro Preview1,307
  4. 4.GPT-5.41,253
  5. 5.Claude Opus 5365
6 models
#
🥇Claude Opus 5anthropic
90.4
$5$25$0.2483651M
🥈GPT-5.5openai
85
$0.188$1.13$0.0117,6841.05M
🥉GPT-5.4 Proopenai
83.3
$27$160$1.57452.91.05M
4Gemini 3.1 Pro Previewgoogle
77.1
$1$6$0.0591,3071.05M
5GPT-5.4openai
73.3
$0.75$6$0.0591,2531.05M
6Gemini 3.5 Flashgoogle
72.1
$0.186$1.11$0.0116,5811.05M

Score reports: anthropic.com · openai.com · deepmind.google — full source URLs are linked on each model page.

Cheapest scorer: Gemini 3.5 Flash at $0.186 input /M.