llmproviders.ai
← All providers

Together AI pricing

verified 2026-08-31

Broad catalog of open-weight models on serverless endpoints — a one-stop shop for Llama, DeepSeek, and Qwen inference.

current models · $ per 1M tokens
ModelInput /1MCached inOutput /1MContext
Kimi K3
Kimi
$3$151.0M
Qwen 3.8-2.4T-A95B
Qwen 3.8
$2$6
GLM-5.3
GLM
$1.40$4.401M
DeepSeek V4 Pro 0813
DeepSeek V4
$1.32$3.961.0M
Qwen 3.7 Max
Qwen 3.7
$1.25$3.75
Llama 3.3 70B Instruct Turbo
Llama 3.3
$1.04$1.04131K
Inkling
Inkling
$1$4.05524K
Gemma 4 31B Instruct
Gemma 4
$0.39$0.97262K
Qwen 3.7 Plus
Qwen 3.7
$0.32$1.281M
MiniMax M3
MiniMax
$0.30$1.20524K
Qwen 3.5 9B
Qwen 3.5
$0.17$0.25262K
Qwen 3.8 Flash
Qwen 3.8
$0.15$0.471M
GLM-5.3 Flash
GLM
$0.15$0.501M
GPT-OSS 120B
GPT-OSS
$0.15$0.60128K
DeepSeek V4 Flash 0731
DeepSeek V4
$0.14$0.281M
GPT-OSS 20B
GPT-OSS
$0.05$0.20128K
Together AI totals · 1M in + 1M out
Qwen 3.5 9B
$0.4200
Qwen 3.8 Flash
$0.6200
Gemma 4 31B Instruct
$1.36
MiniMax M3
$1.50
Qwen 3.7 Plus
$1.60
Llama 3.3 70B Instruct Turbo
$2.08
Qwen 3.7 Max
$5.00
Inkling
$5.05
DeepSeek V4 Pro 0813
$5.28
GLM-5.3
$5.80
Qwen 3.8-2.4T-A95B
$8.00
Kimi K3
$18.00

Cost of 1M input + 1M output tokens. Bar length is square-root scaled so cheap models stay visible next to premium ones.