Leaderboard
GDPval-AA
9 models scored · metric: Elo. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the chat token profile — short turns, light history, little caching — an assistant in a product.
A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 1.5k in · 500 out · 20% cache hit · 8k context.
Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. GDPval-AA is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.
- 1.Gemini 3.5 Flash1,949,204
- 2.Gemini 3.5 Flash Lite1,402,861
- 3.Gemini 3.6 Flash1,014,774
- 4.Grok 4.3533,808
- 5.Gemini 3.1 Pro Preview309,609
Score reports: anthropic.com · anthropic.com · kimi.com · deepmind.google · artificialanalysis.ai · deepmind.google — full source URLs are linked on each model page.
Cheapest scorer: Gemini 3.5 Flash Lite at $0.15 input /M.