← All providers
Groq pricing
verified 2026-08-31Custom LPU hardware serving open models at very high speed. Often the cheapest way to run Llama-class models fast.
current models · $ per 1M tokens
| Model | Input /1M | Cached in | Output /1M | Context |
|---|---|---|---|---|
Qwen 3.8-27B Qwen 3.8 | $0.80 | — | $4 | 131K |
Qwen 3.6-27B Qwen 3.6 | $0.60 | — | $3 | 131K |
GPT OSS 120B GPT-OSS | $0.15 | — | $0.60 | 131K |
GPT OSS 20B GPT-OSS | $0.07 | — | $0.30 | 131K |
Safety GPT OSS 20B GPT-OSS | $0.07 | — | $0.30 | 131K |
Llama Prompt Guard 2 86M Llama Guard | $0.04 | — | $0.04 | 1K |
Llama Prompt Guard 2 22M Llama Guard | $0.03 | — | $0.03 | 1K |
Groq totals · 1M in + 1M out
Llama Prompt Guard 2 22M
$0.0600
Llama Prompt Guard 2 86M
$0.0800
GPT OSS 20B
$0.3750
Safety GPT OSS 20B
$0.3750
GPT OSS 120B
$0.7500
Qwen 3.6-27B
$3.60
Qwen 3.8-27B
$4.80
Cost of 1M input + 1M output tokens. Bar length is square-root scaled so cheap models stay visible next to premium ones.