Skip to content

Leaderboard

BrowseCompofficial site ↗

15 models scored · metric: accuracy. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the agentic token profile — tool loops re-send a growing transcript and think between calls.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 12k in · 2k out · 3k reasoning · 70% cache hit · 32k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. BrowseComp is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · agentic profile
  1. 1.GPT-5.6 Luna31,105
  2. 2.Nemotron 3 Ultra 550B A55B23,617
  3. 3.MiniMax-M313,708
  4. 4.Step 3.7 Flash11,050
  5. 5.LongCat-2.010,417
15 models
#
🥇GPT-6 Astraopenai
91.5
$10$50$0.3122931.05M
🥈Kimi K3moonshotai
91.2
$2$8$0.0531,7241.05M
🥉Claude Opus 5anthropic
90.8
$5$25$0.1585731M
4GPT-5.6 Solopenai
90.4
$2$10$0.0631,4261.05M
5GPT-5.5 Proopenai
90.1
$27.27$163.64$1.19575.41.05M
6GPT-5.4 Proopenai
89.3
$27$160$1.17376.21.05M
7GPT-5.6 Terraopenai
87.5
$1.50$2$0.0194,5201.05M
8Claude Sonnet 5anthropic
84.7
$1.44$7.20$0.0561,5161M
9GPT-5.5openai
84.4
$0.188$1.130.821¢10,2771.05M
10MiniMax-M3minimax
83.52
$0.225$0.90.609¢13,7081.05M
11GPT-5.6 Lunaopenai
83.3
$0.06$0.370.268¢31,1051.05M
12GPT-5.4openai
82.7
$0.75$6$0.0352,3851.05M
13LongCat-2.0meituan
79.9
$0.3$1.200.767¢10,4171.05M
14Step 3.7 Flashstepfun
75.8
$0.185$1.110.686¢11,050262K
15Nemotron 3 Ultra 550B A55Bnvidia
44.4
$0.1$0.10.188¢23,6171.05M

Score reports: openai.com · kimi.com · anthropic.com · openai.com · openai.com · anthropic.com — full source URLs are linked on each model page.

Cheapest scorer: GPT-5.6 Luna at $0.06 input /M.