← All providers
Together AI pricing
verified 2026-08-31Broad catalog of open-weight models on serverless endpoints — a one-stop shop for Llama, DeepSeek, and Qwen inference.
current models · $ per 1M tokens
| Model | Input /1M | Cached in | Output /1M | Context |
|---|---|---|---|---|
Kimi K3 Kimi | $3 | — | $15 | 1.0M |
Qwen 3.8-2.4T-A95B Qwen 3.8 | $2 | — | $6 | — |
GLM-5.3 GLM | $1.40 | — | $4.40 | 1M |
DeepSeek V4 Pro 0813 DeepSeek V4 | $1.32 | — | $3.96 | 1.0M |
Qwen 3.7 Max Qwen 3.7 | $1.25 | — | $3.75 | — |
Llama 3.3 70B Instruct Turbo Llama 3.3 | $1.04 | — | $1.04 | 131K |
Inkling Inkling | $1 | — | $4.05 | 524K |
Gemma 4 31B Instruct Gemma 4 | $0.39 | — | $0.97 | 262K |
Qwen 3.7 Plus Qwen 3.7 | $0.32 | — | $1.28 | 1M |
MiniMax M3 MiniMax | $0.30 | — | $1.20 | 524K |
Qwen 3.5 9B Qwen 3.5 | $0.17 | — | $0.25 | 262K |
Qwen 3.8 Flash Qwen 3.8 | $0.15 | — | $0.47 | 1M |
GLM-5.3 Flash GLM | $0.15 | — | $0.50 | 1M |
GPT-OSS 120B GPT-OSS | $0.15 | — | $0.60 | 128K |
DeepSeek V4 Flash 0731 DeepSeek V4 | $0.14 | — | $0.28 | 1M |
GPT-OSS 20B GPT-OSS | $0.05 | — | $0.20 | 128K |
Together AI totals · 1M in + 1M out
Qwen 3.5 9B
$0.4200
Qwen 3.8 Flash
$0.6200
Gemma 4 31B Instruct
$1.36
MiniMax M3
$1.50
Qwen 3.7 Plus
$1.60
Llama 3.3 70B Instruct Turbo
$2.08
Qwen 3.7 Max
$5.00
Inkling
$5.05
DeepSeek V4 Pro 0813
$5.28
GLM-5.3
$5.80
Qwen 3.8-2.4T-A95B
$8.00
Kimi K3
$18.00
Cost of 1M input + 1M output tokens. Bar length is square-root scaled so cheap models stay visible next to premium ones.