Skip to content

Leaderboard

CharXiv Reasoning

9 models scored · metric: accuracy. Prices are the cheapest listed across all serving providers. Cost/run and Pts/$ are priced for the heavy reasoning token profile — thinking tokens outnumber the answer several times over.

A benchmark is priced for the workload it exercises, not a fixed input:output blend: an agent loop and a single hard question bill very differently on the same rate card. Profile: 2k in · 1.5k out · 8k reasoning · 0% cache hit · 12k context.

Scores are facts as publicly reported by labs (see report links below and on each model page) and aggregated via models.dev. CharXiv Reasoning is maintained by its own project — Model Pulse is not affiliated with or endorsed by it.

Best value · top 5 by pts per dollar · heavy reasoning profile
  1. 1.Muse Glimmer 30B9,850
  2. 2.Gemini 3.5 Flash7,685
  3. 3.Gemini 3.5 Flash Lite6,283
  4. 4.Gemini 3.6 Flash4,816
  5. 5.Muse Spark 1.12,062
9 models
#
🥇Kimi K3moonshotai
91.3
$2$8$0.0801,1411.05M
🥈Gemini 3.6 Flashgoogle
89.4
$0.375$1.88$0.0194,8161.05M
🥉Muse Spark 1.1meta
88.4
$1.25$4.25$0.0432,0621.05M
4Fugu Ultrasakana
86.6
$5$30$0.2952941.05M
5Fugusakana
85.1
1M
6Gemini 3.5 Flashgoogle
84.2
$0.186$1.11$0.0117,6851.05M
7Gemini 3.1 Pro Previewgoogle
83.3
$1$6$0.0591,4121.05M
8Muse Glimmer 30Bmeta
78.8
$0.2$0.80.800¢9,850131K
9Gemini 3.5 Flash Litegoogle
76.5
$0.15$1.25$0.0126,2831.05M

Score reports: kimi.com · deepmind.google · ai.meta.com · console.sakana.ai · deepmind.google · huggingface.co — full source URLs are linked on each model page.

Cheapest scorer: Gemini 3.5 Flash Lite at $0.15 input /M.