- Abacus.AI Smaug Models Cut Agent Costs 100x [2026] ↗
- The TechBeat: AI Coding Tip 035 - Split Every Skill Description Into Three Sentences (9/12/2026) ↗
- DeepSeek Releases V4.1 Flash, a 748-Billion Flash That Now Sees ↗
- Qwen2.5-Coder-32B-Instruct vs Qwen3.8 Max 0902 - AI Model ... ↗
- DeepSeek's new model sets a template for powerful LLMs that run lean ↗theregister.com· 3d ago
- I Built a 256GB “E-Waste” AI Rig to Escape Token Fees. Here’s What Nobody Tells You. ↗
Eden AI raised its DeepSeek V4 Flash 0731 markup by 487%
The lab has not moved — this is a reseller changing its margin, from $0.03 to $0.176 per million input tokens. Worth knowing if you buy through them.
Independent price and benchmark analysis for AI models — what a workload actually costs, and where to buy it.
The wire
What was written · what actually changed
- 01DeepSeek V4 Flash$0.14 → $0.15 /M · DeepSeek▲ 7%
- 02DeepSeek V4 Flash Vision Exp$0.14 → $0.15 /M · DeepSeek▲ 7%
- 01Solar Pro 4$0.03 → $0.09 /M · OpenRouter▲ 200%
- 02Kimi K2.6$0.22 → $0.6 /M · DevPass (LLM Gateway)▲ 173%
- 03Qwen3 Coder Next$0.22 → $0.5 /M · Amazon Bedrock▲ 127%
- 04DeepSeek V3.2$0.28 → $0.62 /M · Vercel AI Gateway▲ 121%
- 05Qwen3-Coder 480B-A35B Instruct$0.22 → $0.45 /M · Amazon Bedrock▲ 105%
Same weights. Wildly different bill.
The gap nobody else publishes · chat workload
55 of 61 providers publish a price for GPT OSS 120B. The cheapest full-context option is LLM Gateway at $0.032/$0.14 per M. Listed prices span 50.3× for identical weights.
- 01GPT OSS 120B78 priced listings · cheapest full context at LLM Gateway$0.06 → $3.0350.3×
- 02Qwen3.6 35B-A3B27 priced listings · cheapest full context at SaladCloud AI Gateway$0.221 → $6.4441.6×
- 03GPT-547 priced listings · cheapest full context at NanoGPT$0.432 → $11.4426.5×
- 04Llama-3.3-70B-Instruct35 priced listings · cheapest full context at Meganova$0.154 → $3.0319.7×
- 05GPT OSS 20B43 priced listings · cheapest full context at Kilo Gateway$0.041 → $0.76518.8×
- 06GLM-4.734 priced listings · cheapest full context at NanoGPT$0.343 → $6.2618.3×
Prices use the chat workload: 2K input + 500 output tokens, 8K context. A model is not one price — it is a spread across every venue that sells it. Listings more than 20× from the model’s own median are dropped as feed errors before anything reaches this board, and a cheapest quote counts only when a second venue is within 1.5× of it.
The boards
Four cuts from today's data · prices use the chat workload
New in the last three weeks
Release dates as the labs published them, priced at the cheapest live listing.
- DeepSeek V4.1 FlashDeepSeek · Sep 10, 2026$0.246
- GPT-6 AstraOpenAI · Sep 4, 2026$19.03
- GPT-6 Astra (Fast)OpenAI · Sep 4, 2026$38.24
- Qwen3.8 Max 0902Alibaba · Sep 2, 2026$2.42
- Gemini 3.8 FlashGoogle DeepMind · Sep 2, 2026$1.40
Value on Aider Polyglot
Score per dollar, priced for the agentic profile. Independently measured results only.
- DeepSeek ReasonerDeepSeek · scored 74.229,633
- DeepSeek ChatDeepSeek · scored 70.228,808
- Qwen3 32BAlibaba · scored 4015,974
- GPT-5OpenAI · scored 8811,992
- Gemini 2.5 FlashGoogle DeepMind · scored 55.111,498
Where the gateway markup is
Median reseller quote against the lab's own counter. Positive means you pay more on the street.
- DeepSeek V4 Pro71 quotes vs DeepSeek+271%
- Mistral Large (latest)7 quotes vs Mistral+265%
- Seed 2.0 Lite6 quotes vs Volcengine Ark+261%
- Ministral 3B7 quotes vs Mistral+132%
- GLM-5.3-Flash50 quotes vs Zhipu AI+105%
Biggest windows, cheapest first
Models carrying a million tokens or more, with what a request actually costs there.
- Llama 4 Scout 17B InstructMeta · 10M$0.289
- Gemini 2.0 Flash-LiteGoogle DeepMind · 2M$0.093
- Grok 4.1 Fast (Reasoning)xAI · 2M$0.254
- Grok 4.1 FastxAI · 2M$0.26
- Grok 4.20 (Non-Reasoning)xAI · 2M$1.45
Work the numbers yourself
Three questions · each answered with a figure computed today
Sticker price is a rate, not a bill.
Pick the shape of your traffic — chat, RAG, agentic, long-context — and get the monthly number, cache writes and context tiers included.
$0.306
cheapest workhorse · Qwen3.5 122B-A10B
Price a workload →
79% of published scores are not independently measured.
We rank on independently measured results only, and price each benchmark for the token profile it actually exercises.
29,633
pts/$ leader · DeepSeek Reasoner on Aider Polyglot
See the evidence →
The rent is the floor, and it is higher than people expect.
We compute the smallest configuration the weights fit on. Throughput is not knowable from a model name, so every break-even here is a lower bound.
$1,314
lowest break-even · Qwen3.6 27B on 1× A100 80GB
Price out self-hosting →
The ecosystem watch
GitHub activity + npm usage · discovery signals, not recommendations
Weekly edition
One email. The changes that actually cost you money.
1159 catalog changes landed in the last 24 days. Most are inventory noise. The digest separates first-party repricing and retirements from reseller noise, names the models whose bill moved, and says plainly when the week was quiet.
Auto-written from the auditable diff log · no manufactured story · unsubscribe in one click
Email edition coming; RSS carries the same digest today.