Skip to content

Pricing

Blaick charges per token, and input and output are priced separately — output typically costs more. Everything is billed from your token balance ($1 = 625,000 tokens); the tables below convert each model's rate to US dollars per 1M tokens so you can compare directly.

Live rates are always in the API

The tables here are a snapshot (last updated September 2026). The authoritative, always-current rate for every model your account can use is on GET /models — each model reports a cost_multiplier and supports_tools. Use the API for billing-critical calculations.

How a request is priced

cost (USD) = (input_tokens  ÷ 1,000,000 × input rate)
           + (output_tokens ÷ 1,000,000 × output rate)

Example — 10,000 input + 1,000 output tokens on Llama 3.3 70B:

(10,000 / 1M × $1.07) + (1,000 / 1M × $1.44) = $0.0107 + $0.0014 ≈ $0.012

Every completion returns the exact billed amount in usage.cost_tokens — see Usage & Tokens.

Economy models

Fast, open-weight models — the best value for classification, extraction, chat, and high-volume workloads.

Model $ / 1M input $ / 1M output
Llama 3.1 8B Instant $0.09 $0.15
Llama 4 Scout $0.20 $0.62
Llama 3.1 8B Turbo $0.33 $0.33
Gemma 2 9B $0.36 $0.36
Mistral 7B $0.36 $0.36
Mixtral 8x7B $0.44 $0.44
Gemma 2 27B $0.49 $0.49
Qwen QwQ 32B $0.53 $0.71
Qwen 2.5 72B $0.67 $0.67
DeepSeek V3 $0.89 $1.62
DeepSeek R1 $1.00 $3.98
Llama 3.3 70B $1.07 $1.44
DeepSeek R1 Distill 70B $1.36 $1.80
Llama 3.3 70B Turbo $1.60 $1.60

Frontier models

Premium models — GPT-4o, GPT-4.1, and Claude Haiku / Sonnet / Opus — are available for the hardest reasoning, long context, and vision.

Live rates

Per-token rates for the frontier tier are being refreshed. For the current rate of any model, call GET /models, or set auto_route: true and let Blaick select and price the right model for each request (Auto-routing).

Spend less

  • Auto-route. Set auto_route: true and Blaick picks the cheapest model that can handle each prompt. See Chat Completions.
  • Right-size the model. Reserve frontier models for hard reasoning; economy models handle most classification, extraction, and chat.
  • Watch output tokens. Output is the pricier side of every model — cap it with max_tokens when you don't need long responses.
  • Subscribe. Monthly plans include large token grants at the same rate; unused tokens roll over up to 3 months. See Subscription plans.