Chat Completions¶
POST /chat/completions is the core of the Blaick API. You send a list of messages and a model, and you get back a model response — either all at once (JSON) or streamed token by token (Server-Sent Events).
Request body¶
{
"model_id": "claude-sonnet-4-20250514",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the capital of France?"}
],
"stream": false,
"temperature": 0.7,
"max_tokens": 4096,
"conversation_id": null,
"auto_route": false
}
| Field | Type | Required | Description |
|---|---|---|---|
model_id |
string | Yes | The model to use. Accepts a model UUID or a provider model id from GET /models. |
messages |
array | Yes | The conversation so far. Each item has a role (system, user, or assistant) and content. |
stream |
boolean | No | Stream the response as SSE. Defaults to false. See Streaming. |
temperature |
number | No | Sampling temperature. Higher is more random. Defaults to 0.7. |
max_tokens |
integer | No | Maximum tokens to generate in the response. Defaults to 4096. |
conversation_id |
string | null | No | Attach this exchange to an existing conversation. If omitted or null, a new conversation is created and its id is returned. |
auto_route |
boolean | No | Let Blaick pick the cheapest capable model instead of model_id. See Auto-routing. |
tools |
array | No | OpenAI-format tool/function definitions the model may call. See Tool calling. |
tool_choice |
string | object | No | How the model selects tools: "auto" (default), "none", "required", or a specific tool object. |
Messages¶
messages is an ordered list. Use roles to structure the exchange:
system— instructions that shape behavior. Optional; put it first if you use it.user— input from the user.assistant— a previous model reply, when you're managing history yourself.
"messages": [
{"role": "system", "content": "You translate English to French. Reply only with the translation."},
{"role": "user", "content": "Good morning"}
]
Two ways to keep context
You can send the full messages history on every request, or let Blaick store it for you with a conversation. With a conversation_id, you only need to send the newest user message.
Non-streaming response¶
With stream omitted or false, you get a single JSON object:
{
"id": "c0ffee00-0000-4000-8000-000000000000",
"conversation_id": "11111111-2222-4333-8444-555555555555",
"model_id": "claude-sonnet-4-20250514",
"content": "The capital of France is Paris.",
"usage": {
"input_tokens": 18,
"output_tokens": 8,
"total_tokens": 26,
"cost_tokens": 52
},
"created_at": "2026-08-04T10:00:00Z",
"auto_routed": false,
"provider_metadata": {
"provider": "anthropic",
"failback_used": false,
"latency_ms": 640
}
}
| Field | Description |
|---|---|
id |
Unique id for this response message. |
conversation_id |
The conversation this exchange belongs to (created for you if you didn't pass one). |
model_id |
The model that actually served the request — may differ from your request if auto-routing or failback kicked in. |
content |
The generated text. Empty when the model returns tool_calls instead. |
tool_calls |
Present when the model wants to call one or more tools you supplied. See Tool calling. |
usage |
Token counts. cost_tokens is what was billed. See Usage & Tokens. |
created_at |
ISO 8601 timestamp. |
auto_routed |
Whether the auto-router chose the model. |
provider_metadata |
Which upstream provider served it, whether a failback occurred, and latency. |
Full example¶
curl https://api.blaick.ai/api/v1/chat/completions \
-H "Authorization: Bearer $BLAICK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model_id": "claude-sonnet-4-20250514",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the capital of France?"}
],
"max_tokens": 100
}'
import os, requests
BASE = "https://api.blaick.ai/api/v1"
headers = {"Authorization": f"Bearer {os.environ['BLAICK_API_KEY']}"}
resp = requests.post(
f"{BASE}/chat/completions",
headers=headers,
json={
"model_id": "claude-sonnet-4-20250514",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the capital of France?"},
],
"max_tokens": 100,
},
)
resp.raise_for_status()
data = resp.json()
print(data["content"])
const BASE = "https://api.blaick.ai/api/v1";
const resp = await fetch(`${BASE}/chat/completions`, {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.BLAICK_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model_id: "claude-sonnet-4-20250514",
messages: [
{ role: "system", content: "You are a concise assistant." },
{ role: "user", content: "What is the capital of France?" },
],
max_tokens: 100,
}),
});
const data = await resp.json();
console.log(data.content);
Auto-routing¶
Set auto_route: true and Blaick analyzes your prompt and picks the cheapest model that can handle it — a small model for a simple question, a larger one for hard reasoning. You still pass a model_id as the ceiling/fallback.
{
"model_id": "claude-sonnet-4-20250514",
"messages": [{"role": "user", "content": "What's 2 + 2?"}],
"auto_route": true
}
When the router changes the model, the response reflects it:
{
"content": "4",
"model_id": "a-cheaper-model",
"auto_routed": true,
"routing_reason": "Simple arithmetic — handled by a lower-tier model.",
"task_type": "simple_qa"
}
To preview what the router would do without spending tokens on a full completion, use POST /chat/triage or POST /chat/auto-route.
Tool calling¶
Give the model tools (functions) it can call, and it will respond with a structured tool_calls request instead of prose. You run the tool in your own code and send the result back — the standard OpenAI-compatible function-calling loop. Blaick never executes your tools; it only relays the model's request and your result.
Model support
Tool calling is available on tool-capable models — check supports_tools on GET /models. It works today on Groq- and Together-served models (e.g. Llama 3.3 70B, Llama 4, Qwen, Kimi). Models that don't support it ignore tools and answer normally.
1. Send a request with tools¶
{
"model_id": "llama-3.3-70b-versatile",
"messages": [
{"role": "user", "content": "What's the weather in Paris?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name"}
},
"required": ["city"]
}
}
}
],
"tool_choice": "auto"
}
tool_choice accepts "auto" (default — the model decides), "none" (never call a tool), "required" (must call one), or a specific tool: {"type": "function", "function": {"name": "get_weather"}}.
2. The model asks to call a tool¶
When the model chooses a tool, content is empty and tool_calls is populated:
{
"content": "",
"tool_calls": [
{
"id": "call_abc123",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"city\": \"Paris\"}"
}
}
],
"usage": { "input_tokens": 82, "output_tokens": 20, "total_tokens": 102, "cost_tokens": 140 }
}
arguments is a JSON string — parse it before use.
3. Run the tool and send the result back¶
Append the assistant's tool-call turn and a tool message with your result, keyed by the tool_call_id, then call the endpoint again:
{
"model_id": "llama-3.3-70b-versatile",
"messages": [
{"role": "user", "content": "What's the weather in Paris?"},
{"role": "assistant", "content": "", "tool_calls": [
{"id": "call_abc123", "type": "function",
"function": {"name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}}
]},
{"role": "tool", "tool_call_id": "call_abc123", "content": "18°C, partly cloudy"}
],
"tools": [ /* same tools array */ ]
}
The model now replies with a normal answer (e.g. "It's 18°C and partly cloudy in Paris."), or requests another tool call — repeat until tool_calls is absent.
Streaming
With stream: true, tool calls arrive as SSE events of type: "tool_calls" (in addition to the usual delta and usage events). See Streaming.
Errors to handle¶
Chat completions can return:
402— out of tokens or subscription past due403— the model isn't available on your plan (common on trial)429— rate limited503— the model is temporarily unavailable
See Errors & Rate Limits for the full list and retry guidance.