Pricing¶
Blaick charges per token, and input and output are priced separately — output typically costs more. Everything is billed from your token balance ($1 = 625,000 tokens); the tables below convert each model's rate to US dollars per 1M tokens so you can compare directly.
Live rates are always in the API
The tables here are a snapshot (last updated September 2026). The authoritative, always-current rate for every model your account can use is on GET /models — each model reports a cost_multiplier and supports_tools. Use the API for billing-critical calculations.
How a request is priced¶
Example — 10,000 input + 1,000 output tokens on Llama 3.3 70B:
Every completion returns the exact billed amount in usage.cost_tokens — see Usage & Tokens.
Economy models¶
Fast, open-weight models — the best value for classification, extraction, chat, and high-volume workloads.
| Model | $ / 1M input | $ / 1M output |
|---|---|---|
| Llama 3.1 8B Instant | $0.09 | $0.15 |
| Llama 4 Scout | $0.20 | $0.62 |
| Llama 3.1 8B Turbo | $0.33 | $0.33 |
| Gemma 2 9B | $0.36 | $0.36 |
| Mistral 7B | $0.36 | $0.36 |
| Mixtral 8x7B | $0.44 | $0.44 |
| Gemma 2 27B | $0.49 | $0.49 |
| Qwen QwQ 32B | $0.53 | $0.71 |
| Qwen 2.5 72B | $0.67 | $0.67 |
| DeepSeek V3 | $0.89 | $1.62 |
| DeepSeek R1 | $1.00 | $3.98 |
| Llama 3.3 70B | $1.07 | $1.44 |
| DeepSeek R1 Distill 70B | $1.36 | $1.80 |
| Llama 3.3 70B Turbo | $1.60 | $1.60 |
Frontier models¶
Premium models — GPT-4o, GPT-4.1, and Claude Haiku / Sonnet / Opus — are available for the hardest reasoning, long context, and vision.
Live rates
Per-token rates for the frontier tier are being refreshed. For the current rate of any model, call GET /models, or set auto_route: true and let Blaick select and price the right model for each request (Auto-routing).
Spend less¶
- Auto-route. Set
auto_route: trueand Blaick picks the cheapest model that can handle each prompt. See Chat Completions. - Right-size the model. Reserve frontier models for hard reasoning; economy models handle most classification, extraction, and chat.
- Watch output tokens. Output is the pricier side of every model — cap it with
max_tokenswhen you don't need long responses. - Subscribe. Monthly plans include large token grants at the same rate; unused tokens roll over up to 3 months. See Subscription plans.