Leaderboard
Terminal-Bench Hardofficial site ↗
22 models scored · metric: success rate. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the agentic token profile — tool loops re-send a growing transcript and think between calls.
A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 12k in · 2k out · 3k reasoning · 70% cache hit · 32k context.
Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. Terminal-Bench Hard is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.
- 1.Step 3.5 Flash 260314,764
- 2.Step 3.5 Flash12,674
- 3.Qwen3-Coder 30B-A3B Instruct8,801
- 4.GLM-4.5-Air5,284
- 5.Step 3.7 Flash5,190
Score reports: openrouter.ai · openrouter.ai · openrouter.ai · openrouter.ai · openrouter.ai · openrouter.ai — full source URLs are linked on each model page.
Cheapest scorer: Llama-3.3-70B-Instruct at $0.05 input /M.