Skip to content

Leaderboard

AutomationBench

7 models scored · metric: pass@1. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the agentic token profile — tool loops re-send a growing transcript and think between calls.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 12k in · 2k out · 3k reasoning · 70% cache hit · 32k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. AutomationBench is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · agentic profile
  1. 1.DeepSeek V4 Flash 073141,987
  2. 2.Gemini 3.7 Flash1,373
  3. 3.Qwen3.8 Max Preview636
  4. 4.Kimi K3582
  5. 5.Claude Opus 5164
7 models
#
🥇GPT-6 Astraopenai
41.4
$10$50$0.3121331.05M
🥈Kimi K3moonshotai
30.8
$2$8$0.0535821.05M
🥉Gemini 3.7 Flashgoogle
30.4
$0.75$3.75$0.0221,3731.05M
4Qwen3.8 Max Previewalibaba
27.3
$2$6$0.0436361M
5Claude Opus 5anthropic
26
$5$25$0.1581641M
6DeepSeek V4 Flash 0731deepseek
25.1
$0.035$0.070.060¢41,9871.31M
7Claude Fable 5anthropic
17.4
$3$18.50$0.1111561M

Score reports: openai.com · kimi.com · deepmind.google · alibabacloud.com · anthropic.com · api-docs.deepseek.com — full source URLs are linked on each model page.

Cheapest scorer: DeepSeek V4 Flash 0731 at $0.035 input /M.