Skip to content

Leaderboard

Terminal-Bench Hardofficial site ↗

22 models scored · metric: success rate. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the agentic token profile — tool loops re-send a growing transcript and think between calls.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 12k in · 2k out · 3k reasoning · 70% cache hit · 32k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. Terminal-Bench Hard is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · agentic profile
  1. 1.Step 3.5 Flash 260314,764
  2. 2.Step 3.5 Flash12,674
  3. 3.Qwen3-Coder 30B-A3B Instruct8,801
  4. 4.GLM-4.5-Air5,284
  5. 5.Step 3.7 Flash5,190
22 models
#
🥇GPT-5-Codexopenai
37.9
$1.07$8.50$0.049771400K
🥈Step 3.7 Flashstepfun
35.6
$0.185$1.110.686¢5,190262K
🥉Step 3.5 Flash 2603stepfun
32.6
$0.1$0.30.221¢14,764262K
4Step 3.5 Flashstepfun
27.3
$0.09$0.30.215¢12,674262K
5Gemini 2.5 Progoogle
26.5
$0.87$7$0.0406551.07M
6GLM-4.6zhipuai
25
$0.286$1.140.966¢2,589205K
7GLM-4.5zhipuai
22
$0.286$1.140.966¢2,278131K
8Qwen3 Maxalibaba
20.5
$0.36$1.430.970¢2,114262K
9GLM-4.5-Airzhipuai
20.5
$0.1$0.50.388¢5,284131K
10Devstral 2mistral
18.9
$0.4$2$0.0121,512262K
11Mistral Small 4mistral
17.4
$0.143$0.5680.373¢4,665262K
12Mistral Large 3mistral
15.9
$0.5$1.50$0.0111,497262K
13Qwen3-Coder 30B-A3B Instructalibaba
15.2
$0.06$0.250.173¢8,801262K
14Gemini 2.5 Flashgoogle
13.6
$0.09$0.710.479¢2,8381.07M
15GPT-4o (2024-08-06)openai
8.3
$2.50$10$0.074112128K
16GPT-4o (2024-11-20)openai
8.3
$2.50$10$0.074112128K
17DeepSeek-R1deepseek
6.1
$0.4$1.70$0.012494164K
18Mistral Large 2.1mistral
6.1
$2.01$6$0.043143131K
19GLM-4.5Vzhipuai
5.3
$0.29$0.860.830¢638128K
20Gemini 2.5 Flash-Litegoogle
4.5
$0.1$0.10.188¢2,3941.05M

Score reports: openrouter.ai · openrouter.ai · openrouter.ai · openrouter.ai · openrouter.ai · openrouter.ai — full source URLs are linked on each model page.

Cheapest scorer: Llama-3.3-70B-Instruct at $0.05 input /M.