rates verified 2026-08-31
API token pricing
All 129 models across 14 providers, priced per 1M tokens. Tap a column to sort; cell color runs from cheap (green) to expensive (orange) on a log scale. Prices come from each provider's official price list.
| Tier | |||||
|---|---|---|---|---|---|
Llama Prompt Guard 2 22Mpreview Groq | Budget | $0.03 | — | $0.03 | 1K |
Command R7B Cohere | Budget | $0.04 | — | $0.15 | 128K |
Llama Prompt Guard 2 86Mpreview Groq | Budget | $0.04 | — | $0.04 | 1K |
Qwen Flash Alibaba Cloud (Qwen) | Budget | $0.05 | — | $0.40 | 1M |
Qwen Turbo Alibaba Cloud (Qwen) | Budget | $0.05 | — | $0.20 | — |
GPT-OSS 20B Together AI | Budget | $0.05 | — | $0.20 | 128K |
GLM-4.7-FlashX Z.ai | Budget | $0.07 | $0.01 | $0.40 | 200K |
GPT OSS 20B Groq | Budget | $0.07 | — | $0.30 | 131K |
Safety GPT OSS 20Bpreview Groq | Budget | $0.07 | — | $0.30 | 131K |
GLM-5.3-Flash Z.ai | Budget | $0.07 | $0.01 | $0.25 | — |
Ministral 3 3B Mistral | Budget | $0.10 | — | $0.10 | — |
DeepSeek V4 Flash 0731 Together AI | Budget | $0.14 | — | $0.28 | 1M |
Command R Cohere | Mid | $0.15 | — | $0.60 | 128K |
GPT OSS 120B Groq | Mid | $0.15 | — | $0.60 | 131K |
Mistral Small 4 Mistral | Budget | $0.15 | — | $0.60 | — |
Ministral 3 8B Mistral | Budget | $0.15 | — | $0.15 | — |
Qwen 3.8 Flash Together AI | Budget | $0.15 | — | $0.47 | 1M |
GLM-5.3 Flash Together AI | Budget | $0.15 | — | $0.50 | 1M |
GPT-OSS 120B Together AI | Mid | $0.15 | — | $0.60 | 128K |
Qwen 3.5 9B Together AI | Budget | $0.17 | — | $0.25 | 262K |
Ministral 3 14B Mistral | Budget | $0.20 | — | $0.20 | — |
GPT-5.6 Luna OpenAI | Budget | $0.20 | $0.02 | $1.20 | 1.1M |
GPT-5.4 nano OpenAI | Budget | $0.20 | $0.02 | $1.25 | — |
Gemini 3.1 Flash-Lite Google | Budget | $0.25 | $0.03 | $1.50 | — |
Qwen3 Coder Flash Alibaba Cloud (Qwen) | Budget | $0.30 | — | $1.50 | 1M |
Gemini 3.5 Flash-Lite Google | Budget | $0.30 | $0.03 | $2.50 | — |
MiniMax M3 MiniMax | Frontier | $0.30 | $0.06 | $1.20 | 1M |
MiniMax M2.7 MiniMax | Mid | $0.30 | $0.06 | $1.20 | — |
Codestral Mistral | Mid | $0.30 | — | $0.90 | — |
MiniMax M3 Together AI | Mid | $0.30 | — | $1.20 | 524K |
Qwen 3.7 Plus Together AI | Mid | $0.32 | — | $1.28 | 1M |
Gemma 4 31B Instruct Together AI | Mid | $0.39 | — | $0.97 | 262K |
Qwen Plus Alibaba Cloud (Qwen) | Mid | $0.40 | — | $1.20 | 1M |
DeepSeek V4 Flash DeepSeek | Budget | $0.44 | $0.01 | $1.32 | 1M |
DeepSeek V4 Flash Vision (Experimental)preview DeepSeek | Budget | $0.44 | $0.01 | $1.32 | 1M |
Mistral Large 3 Mistral | Frontier | $0.50 | — | $1.50 | — |
Qwen 3.6-27Bpreview Groq | Mid | $0.60 | — | $3 | 131K |
MiniMax M2.7 High-Speed MiniMax | Mid | $0.60 | $0.06 | $2.40 | — |
GLM-4.7 Z.ai | Mid | $0.60 | $0.11 | $2.20 | 200K |
Gemini 3.7 Flash Google | Mid | $0.75 | $0.07 | $3.75 | — |
Gemini 3.6 Flash Google | Mid | $0.75 | $0.07 | $3.75 | — |
GPT-5.4 mini OpenAI | Mid | $0.75 | $0.07 | $4.50 | — |
Qwen 3.8-27Bpreview Groq | Mid | $0.80 | — | $4 | 131K |
Kimi K2.7 Code Moonshot AI | Mid | $0.95 | $0.19 | $4 | 262K |
Kimi K2.6 Moonshot AI | Mid | $0.95 | $0.16 | $4 | 262K |
Qwen3 Coder Plus Alibaba Cloud (Qwen) | Mid | $1 | — | $5 | 1M |
Claude Haiku 4.5 Anthropic | Budget | $1 | $0.10 | $5 | 200K |
Sonar Perplexity | Budget | $1 | — | $1 | 128K |
Inkling Together AI | Frontier | $1 | — | $4.05 | 524K |
Grok Build 0.1 xAI | Budget | $1 | $0.20 | $2 | 256K |
GLM-5 Z.ai | Mid | $1 | $0.20 | $3.20 | — |
Llama 3.3 70B Instruct Turbo Together AI | Mid | $1.04 | — | $1.04 | 131K |
GPT-5.1 OpenAI | Frontier | $1.25 | $0.13 | $10 | 400K |
Qwen 3.7 Max Together AI | Frontier | $1.25 | — | $3.75 | — |
Grok 4.3 xAI | Mid | $1.25 | $0.20 | $2.50 | 1M |
Grok 4.20 Reasoning (0309) xAI | Mid | $1.25 | $0.20 | $2.50 | 1M |
Grok 4.20 Non-Reasoning (0309) xAI | Mid | $1.25 | $0.20 | $2.50 | 1M |
Grok 4.20 Multi-Agent (0309) xAI | Mid | $1.25 | $0.20 | $2.50 | 1M |
DeepSeek V4 Pro DeepSeek | Frontier | $1.32 | $0.04 | $3.96 | 1M |
DeepSeek V4 Pro 0813 Together AI | Frontier | $1.32 | — | $3.96 | 1.0M |
GLM-5.3 Together AI | Frontier | $1.40 | — | $4.40 | 1M |
GLM-5.3 Z.ai | Frontier | $1.40 | $0.26 | $4.40 | 1M |
GLM-5.2 Z.ai | Frontier | $1.40 | $0.26 | $4.40 | 1M |
GLM-5.1 Z.ai | Frontier | $1.40 | $0.26 | $4.40 | — |
Gemini 3.5 Flash Google | Mid | $1.50 | $0.15 | $9 | — |
Mistral Medium 3.5 Mistral | Frontier | $1.50 | — | $7.50 | — |
GPT-5.2 OpenAI | Frontier | $1.75 | $0.17 | $14 | — |
Kimi K2.7 Code High-Speed Moonshot AI | Mid | $1.90 | $0.38 | $8 | 262K |
Qwen3.8 Max Alibaba Cloud (Qwen) | Frontier | $2 | — | $6 | 1M |
Claude Sonnet 5 Anthropic | Mid | $2 | $0.20 | $10 | 1M |
Gemini 3.1 Pro Previewpreview Google | Frontier | $2 | $0.20 | $12 | — |
GPT-5.6 Terra OpenAI | Mid | $2 | $0.20 | $12 | 1.1M |
Sonar Reasoning Pro Perplexity | Mid | $2 | — | $8 | 128K |
Sonar Deep Research Perplexity | Frontier | $2 | — | $8 | 128K |
Qwen 3.8-2.4T-A95B Together AI | Frontier | $2 | — | $6 | — |
Grok 4.6 xAI | Frontier | $2 | $0.50 | $6 | 500K |
Grok 4.5 xAI | Frontier | $2 | $0.30 | $6 | 500K |
Qwen3.7 Max Alibaba Cloud (Qwen) | Frontier | $2.50 | — | $7.50 | 1M |
GPT-5.4 OpenAI | Frontier | $2.50 | $0.25 | $15 | — |
Kimi K3 Moonshot AI | Frontier | $3 | $0.30 | $15 | 1.0M |
Sonar Pro Perplexity | Frontier | $3 | — | $15 | 200K |
Kimi K3 Together AI | Frontier | $3 | — | $15 | 1.0M |
GPT-5.6 Sol OpenAI | Frontier | $4 | $0.40 | $20 | 1.1M |
Claude Opus 5 Anthropic | Frontier | $5 | $0.50 | $25 | 1M |
GPT-5.5 OpenAI | Frontier | $5 | $0.50 | $30 | — |
Claude Fable 5 Anthropic | Frontier | $10 | $1 | $50 | 1M |
GPT-5.2 Pro OpenAI | Frontier | $21 | — | $168 | — |
GPT-5.5 Pro OpenAI | Frontier | $30 | — | $180 | — |
GPT-5.4 Pro OpenAI | Frontier | $30 | — | $180 | — |
Cost of 1M input + 1M output tokens. Bar length is square-root scaled so cheap models stay visible next to premium ones.
Per-token rates only matter multiplied by your traffic. The calculators turn these rates into monthly bills — including cached-input discounts and subscription break-evens.
questions
Frequently asked
›Where do these prices come from?
Every rate is read from the provider's own published price list (the pages linked on each provider profile). We record standard pay-as-you-go rates — not batch, off-peak, or enterprise discounts.
›What does the cached-input column mean?
Most providers charge a reduced rate when a prompt prefix repeats and hits their cache. If your app reuses a long system prompt, the cached rate can matter more than the headline input rate.
›How current is the data?
Rates on this page were last verified on 2026-08-31. Providers change prices without notice, so confirm on the provider's page before committing a budget.
›Can I get this data as JSON?
Yes — the full dataset is available at /api/pricing.json, free to use with attribution. See the API docs page for the schema.
›Why are some models marked legacy?
Legacy models are still purchasable but superseded by a newer generation from the same provider. They're hidden by default; tick the checkbox above the table to include them.