Skip to content

Leaderboard

Toolathlon

8 models scored · metric: success rate. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the agentic token profile — tool loops re-send a growing transcript and think between calls.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 12k in · 2k out · 3k reasoning · 70% cache hit · 32k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. Toolathlon is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · agentic profile
  1. 1.GPT-5.6 Luna19,940
  2. 2.Step 3.7 Flash7,216
  3. 3.Gemini 3.5 Flash6,946
  4. 4.GPT-5.56,770
  5. 5.GPT-5.4 nano5,360
8 models
#
🥇GPT-5.6 Solopenai
58
$2$10$0.0639151.05M
🥈Gemini 3.5 Flashgoogle
56.5
$0.186$1.110.813¢6,9461.05M
🥉GPT-5.5openai
55.6
$0.188$1.130.821¢6,7701.05M
4GPT-5.6 Lunaopenai
53.4
$0.06$0.370.268¢19,9401.05M
5GPT-5.6 Terraopenai
53.1
$1.50$2$0.0192,7431.05M
6Step 3.7 Flashstepfun
49.5
$0.185$1.110.686¢7,216262K
7GPT-5.4 miniopenai
42.9
$0.375$4$0.0221,920400K
8GPT-5.4 nanoopenai
35.5
$0.18$1.100.662¢5,3601.05M

Score reports: openai.com · deepmind.google · openai.com · static.stepfun.com · openai.com — full source URLs are linked on each model page.

Cheapest scorer: GPT-5.6 Luna at $0.06 input /M.