Skip to content

Plan your spend

What the workload costs

A sticker price is a rate, not a bill. Pick the shape of your traffic and see what each model would actually cost per month at its cheapest listed provider — including the cache writes, reasoning tokens and long-context tiers a headline $/M figure leaves out.

Workload

Short turns, light history, little caching — an assistant in a product.

input
1.5k tok
output
500 tok
reasoning
cache hit
20%
context
8k tok
Requests / month
1224 models
#RelativePer request
1Whisper 3 Largeopenai$0.47
0.000¢$0.0030 / $0
2KB Whisperunattributed$0.48
0.000¢$0.0023 / $0.0023
3Voxtral Small 24B 2507mistral$0.48
0.000¢$0.0023 / $0.0023
4Whisper Large v3 Turboopenai$0.48
0.000¢$0.0023 / $0.0023
5Green Sunattributed$0.69
0.001¢$0.0044 / $0
6Green S Prounattributed$0.69
0.001¢$0.0044 / $0
7Voxtral Mini 3Bunattributed$0.73
0.001¢$0.0046 / $0
8All-MiniLM-L6-v2unattributed$1
0.001¢$0.0090 / $0
9Multi-QA-mpnet-base-dot-v1unattributed$1
0.001¢$0.0090 / $0
10Qwen 3 Embedding 4Bunattributed$2
0.002¢$0.01 / $0
11Qwen 3 Embedding 8Bunattributed$2
0.002¢$0.01 / $0
12BGE Reranker v2 M3unattributed$2
0.002¢$0.01 / $0
13Qwen3-Embedding-8Bunattributed$2
0.002¢$0.01 / $0
14Llama 3.2 1B Instructunattributed$2
0.002¢$0.01 / $0.01
15Meta Llama Prompt Guard 2 22Munattributed$2
0.002¢$0.01 / $0.01
16Meta Llama Prompt Guard 2 86Munattributed$2
0.002¢$0.01 / $0.01
17Qwen3 Embedding 0.6Bunattributed$2
0.002¢$0.01 / $0.01
18Qwen3 Reranker 0.6Bunattributed$2
0.002¢$0.01 / $0.01
19Google Gemma 2unattributed$3
0.003¢$0.01 / $0.03
20BGE M3unattributed$3
0.003¢$0.02 / $0

Each estimate uses the model's cheapest listed provider and bills every line the workload touches: uncached input, cache reads, cache writes, reasoning tokens and output. Where a provider publishes no cache or reasoning rate, those tokens are billed at the full input or output rate — we never assume a discount nobody offers.

Sticker prices are shown for reference only — they are not what the workload costs.