Leaderboard
Aider Polyglotofficial site ↗
28 models scored · metric: percent correct. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the agentic token profile — tool loops re-send a growing transcript and think between calls.
A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 12k in · 2k out · 3k reasoning · 70% cache hit · 32k context.
Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. Aider Polyglot is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.
- 1.DeepSeek Reasoner29,633
- 2.DeepSeek Chat28,808
- 3.Qwen3 32B15,974
- 4.GPT-511,992
- 5.Gemini 2.5 Flash11,498
Score reports: aider.chat — full source URLs are linked on each model page.
Cheapest scorer: GPT-4o mini at $0.075 input /M.