Skip to content

The other half of the price question

Self-host or buy?

67 of the models tracked here are permissively licensed and download without asking, so paying per token is a choice. A rented GPU costs the same whether you serve one token or a billion, and the API has no fixed cost — so the question is simply whether your bill clears the rent. Set your volume and see.

100M tokens/mo

At 100M tokens a month, buying the API wins for every model here. A rented GPU costs more than your whole bill would.

79 models
Verdict
MiniCPM5-1B
Openbmb · 1B · $0.283/M
$28$803
1× A100 40GB
−$775Keep buying
Nemotron 3.5 Content Safety
NVIDIA · 4B · $0.207/M
$21$803
1× A100 40GB
−$782Keep buying
Pixtral 12B
Mistral · 13B · $0.156/M
$16$803
1× A100 40GB
−$787Keep buying
Phi-4-mini
Microsoft · 4B · $0.134/M
$13$803
1× A100 40GB
−$790Keep buying
Qwen2.5-Coder-0.5B
Alibaba · 0B · $0.104/M
$10$803
1× A100 40GB
−$793Keep buying
Llama-3.2-1B
Meta · 1B · $0.104/M
$10$803
1× A100 40GB
−$793Keep buying
Llama-3.2-3B
Meta · 3B · $0.104/M
$10$803
1× A100 40GB
−$793Keep buying
GLM-4.6V-Flash
Zhipu AI · 10B · $0.066/M
$7$803
1× A100 40GB
−$796Keep buying
Gemma 3 12B IT
Google DeepMind · 12B · $0.064/M
$6$803
1× A100 40GB
−$797Keep buying
Qwen3.5 9B
Alibaba · 10B · $0.064/M
$6$803
1× A100 40GB
−$797Keep buying
Llama-3.2-11B-Vision-Instruct
Meta · 11B · $0.057/M
$6$803
1× A100 40GB
−$797Keep buying
Llama-Guard-3-8B
Meta · 8B · $0.057/M
$6$803
1× A100 40GB
−$797Keep buying
Gemma 3 4B IT
Google DeepMind · 4B · $0.052/M
$5$803
1× A100 40GB
−$798Keep buying
Granite-4.0-H-Micro
Ibm · 3B · $0.041/M
$4$803
1× A100 40GB
−$799Keep buying
DeepSeek OCR 2
DeepSeek · 3B · $0.031/M
$3$803
1× A100 40GB
−$800Keep buying
Llama-3.1-8B-Instruct
Meta · 8B · $0.026/M
$3$803
1× A100 40GB
−$800Keep buying
Whisper Large v3 Turbo
OpenAI · 1B · $0.0024/M
$0$803
1× A100 40GB
−$803Keep buying
Qwen3.6 27B
Alibaba · 28B · $0.509/M
$51$1,314
1× A100 80GB
−$1,263Keep buying
Codestral-22B-v0.1
Mistral · 22B · $0.461/M
$46$1,314
1× A100 80GB
−$1,268Keep buying
Gemma-SEA-LION-v4-27B-IT
Aisingapore · 27B · $0.415/M
$42$1,314
1× A100 80GB
−$1,272Keep buying

Not your decision to make

Licence says API only

These publish their weights, but under a licence that forbids commercial use. However large you get, buying the API is the only route for a product.

Bigger than one box

Beyond a single node

At bf16 these need more memory than eight GPUs hold, so a meaningful rent figure would mean pricing a cluster and the networking around it. Quantising to 8- or 4-bit roughly halves or quarters the requirement and can bring some of them back into range — at some cost to quality.

How this is worked out

Method

Throughput is deliberately not modelled. Tokens per second depends on hardware, batch size, quantisation, framework and sequence length. Any single figure would be invented precision, so instead of guessing we compare the one cost that is certain: the rent.

Memory is parameters × 2 bytes (bf16) plus 25% for KV cache, activations and framework overhead. Long contexts and large batches need more. We then pick the cheapest configuration of up to eight GPUs whose combined VRAM holds that.

GPU rates are representative on-demand list prices as of 2026-08A100 40GB $1.1/hr, A100 80GB $1.8/hr, H100 80GB $3/hr, H200 141GB $4/hr — billed at 730 hours a month. Reserved and spot capacity run materially cheaper.

The comparison flatters self-hosting. It assumes a permanently saturated GPU and counts no engineering time, no redundancy, no idle capacity and no traffic peaks. Real deployments are worse on every one of those. So treat a positive difference as permission to price it out properly — not as a saving.

licence and parameter counts from huggingface.co model cards · API prices are the cheapest listed blended rate across every provider we track