LLM API cost calculator
Pick a workload, set your token counts and request volume, and see the monthly and annual API bill for each major model — sorted cheapest first.
3,162
100100,000
| Model | Monthly | Annual |
|---|---|---|
| DeepSeek V4 (Flash)Cheapest | $21.25 | $254.98 |
| Gemini 2.5 Flash-Lite | $22.77 | $273.20 |
| Mistral Small 4 | $34.15 | $409.80 |
| GPT-5.6 Luna | $60.71 | $728.52 |
| Claude Haiku 4.5 | $265.61 | $3.2K |
| Claude Sonnet 5 | $531.22 | $6.4K |
| Gemini 3.1 Pro | $607.10 | $7.3K |
| GPT-5.6 | $1.5K | $18.2K |
Estimates use published list API pricing and assume 30 days per month. Batch discounts, prompt caching and committed-use pricing are not included. Verify current rates with each provider before committing.
Read the model comparisons
- Claude Haiku 4.5 vs Gemini 2.5 Flash-Lite (2026)Claude Haiku 4.5 vs Gemini 2.5 Flash-Lite on price, context, latency and instruction following for high-volume support and chat workloads. August 2026.
- Claude Sonnet 5 vs Llama 4 (2026) — Frontier API vs Open WeightsClaude Sonnet 5 vs Llama 4 on coding, agents, long context, cost and control. When a hosted frontier model beats an open one you can self-host. August 2026.
- DeepSeek V4 vs GPT-5.6 (2026) — Cost, Coding & Where Each One WinsDeepSeek V4 vs GPT-5.6 compared on coding, cost, tool use and self-hosting. When the cheaper model is genuinely good enough, and when it isn't. August 2026.
- Gemini vs DeepSeek V4 (2026) — Context, Cost & RAG Fit ComparedGemini 3.1 Pro and 2.5 Flash-Lite vs DeepSeek V4 on context window, price, latency and self-hosting. Which fits RAG and long-document work in 2026. August 2026.