Leaderboard
Agents' Last Exam
5 models scored · metric: score. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the chat token profile — short turns, light history, little caching — an assistant in a product.
This benchmark has not been calibrated to a token profile, so the figures use the site default. Treat the ordering as indicative rather than tuned to Agents' Last Exam. Profile: 1.5k in · 500 out · 20% cache hit · 8k context.
Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. Agents' Last Exam is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.
- 1.DeepSeek V4 Flash 0731308,351
- 2.GPT-5.6 Luna179,964
- 3.GPT-5.6 Terra17,041
- 4.GPT-5.6 Sol6,891
- 5.GPT-6 Astra1,558
Score reports: openai.com · openai.com · api-docs.deepseek.com — full source URLs are linked on each model page.
Cheapest scorer: DeepSeek V4 Flash 0731 at $0.035 input /M.