Skip to content

Chat Completions

POST /chat/completions is the core of the Blaick API. You send a list of messages and a model, and you get back a model response — either all at once (JSON) or streamed token by token (Server-Sent Events).

POST https://api.blaick.ai/api/v1/chat/completions

Request body

{
  "model_id": "claude-sonnet-4-20250514",
  "messages": [
    {"role": "system", "content": "You are a concise assistant."},
    {"role": "user", "content": "What is the capital of France?"}
  ],
  "stream": false,
  "temperature": 0.7,
  "max_tokens": 4096,
  "conversation_id": null,
  "auto_route": false
}
Field Type Required Description
model_id string Yes The model to use. Accepts a model UUID or a provider model id from GET /models.
messages array Yes The conversation so far. Each item has a role (system, user, or assistant) and content.
stream boolean No Stream the response as SSE. Defaults to false. See Streaming.
temperature number No Sampling temperature. Higher is more random. Defaults to 0.7.
max_tokens integer No Maximum tokens to generate in the response. Defaults to 4096.
conversation_id string | null No Attach this exchange to an existing conversation. If omitted or null, a new conversation is created and its id is returned.
auto_route boolean No Let Blaick pick the cheapest capable model instead of model_id. See Auto-routing.
tools array No OpenAI-format tool/function definitions the model may call. See Tool calling.
tool_choice string | object No How the model selects tools: "auto" (default), "none", "required", or a specific tool object.

Messages

messages is an ordered list. Use roles to structure the exchange:

  • system — instructions that shape behavior. Optional; put it first if you use it.
  • user — input from the user.
  • assistant — a previous model reply, when you're managing history yourself.
"messages": [
  {"role": "system", "content": "You translate English to French. Reply only with the translation."},
  {"role": "user", "content": "Good morning"}
]

Two ways to keep context

You can send the full messages history on every request, or let Blaick store it for you with a conversation. With a conversation_id, you only need to send the newest user message.

Non-streaming response

With stream omitted or false, you get a single JSON object:

{
  "id": "c0ffee00-0000-4000-8000-000000000000",
  "conversation_id": "11111111-2222-4333-8444-555555555555",
  "model_id": "claude-sonnet-4-20250514",
  "content": "The capital of France is Paris.",
  "usage": {
    "input_tokens": 18,
    "output_tokens": 8,
    "total_tokens": 26,
    "cost_tokens": 52
  },
  "created_at": "2026-08-04T10:00:00Z",
  "auto_routed": false,
  "provider_metadata": {
    "provider": "anthropic",
    "failback_used": false,
    "latency_ms": 640
  }
}
Field Description
id Unique id for this response message.
conversation_id The conversation this exchange belongs to (created for you if you didn't pass one).
model_id The model that actually served the request — may differ from your request if auto-routing or failback kicked in.
content The generated text. Empty when the model returns tool_calls instead.
tool_calls Present when the model wants to call one or more tools you supplied. See Tool calling.
usage Token counts. cost_tokens is what was billed. See Usage & Tokens.
created_at ISO 8601 timestamp.
auto_routed Whether the auto-router chose the model.
provider_metadata Which upstream provider served it, whether a failback occurred, and latency.

Full example

curl https://api.blaick.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $BLAICK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model_id": "claude-sonnet-4-20250514",
    "messages": [
      {"role": "system", "content": "You are a concise assistant."},
      {"role": "user", "content": "What is the capital of France?"}
    ],
    "max_tokens": 100
  }'
import os, requests

BASE = "https://api.blaick.ai/api/v1"
headers = {"Authorization": f"Bearer {os.environ['BLAICK_API_KEY']}"}

resp = requests.post(
    f"{BASE}/chat/completions",
    headers=headers,
    json={
        "model_id": "claude-sonnet-4-20250514",
        "messages": [
            {"role": "system", "content": "You are a concise assistant."},
            {"role": "user", "content": "What is the capital of France?"},
        ],
        "max_tokens": 100,
    },
)
resp.raise_for_status()
data = resp.json()
print(data["content"])
const BASE = "https://api.blaick.ai/api/v1";

const resp = await fetch(`${BASE}/chat/completions`, {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.BLAICK_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model_id: "claude-sonnet-4-20250514",
    messages: [
      { role: "system", content: "You are a concise assistant." },
      { role: "user", content: "What is the capital of France?" },
    ],
    max_tokens: 100,
  }),
});

const data = await resp.json();
console.log(data.content);

Auto-routing

Set auto_route: true and Blaick analyzes your prompt and picks the cheapest model that can handle it — a small model for a simple question, a larger one for hard reasoning. You still pass a model_id as the ceiling/fallback.

{
  "model_id": "claude-sonnet-4-20250514",
  "messages": [{"role": "user", "content": "What's 2 + 2?"}],
  "auto_route": true
}

When the router changes the model, the response reflects it:

{
  "content": "4",
  "model_id": "a-cheaper-model",
  "auto_routed": true,
  "routing_reason": "Simple arithmetic — handled by a lower-tier model.",
  "task_type": "simple_qa"
}

To preview what the router would do without spending tokens on a full completion, use POST /chat/triage or POST /chat/auto-route.

Tool calling

Give the model tools (functions) it can call, and it will respond with a structured tool_calls request instead of prose. You run the tool in your own code and send the result back — the standard OpenAI-compatible function-calling loop. Blaick never executes your tools; it only relays the model's request and your result.

Model support

Tool calling is available on tool-capable models — check supports_tools on GET /models. It works today on Groq- and Together-served models (e.g. Llama 3.3 70B, Llama 4, Qwen, Kimi). Models that don't support it ignore tools and answer normally.

1. Send a request with tools

{
  "model_id": "llama-3.3-70b-versatile",
  "messages": [
    {"role": "user", "content": "What's the weather in Paris?"}
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city.",
        "parameters": {
          "type": "object",
          "properties": {
            "city": {"type": "string", "description": "City name"}
          },
          "required": ["city"]
        }
      }
    }
  ],
  "tool_choice": "auto"
}

tool_choice accepts "auto" (default — the model decides), "none" (never call a tool), "required" (must call one), or a specific tool: {"type": "function", "function": {"name": "get_weather"}}.

2. The model asks to call a tool

When the model chooses a tool, content is empty and tool_calls is populated:

{
  "content": "",
  "tool_calls": [
    {
      "id": "call_abc123",
      "type": "function",
      "function": {
        "name": "get_weather",
        "arguments": "{\"city\": \"Paris\"}"
      }
    }
  ],
  "usage": { "input_tokens": 82, "output_tokens": 20, "total_tokens": 102, "cost_tokens": 140 }
}

arguments is a JSON string — parse it before use.

3. Run the tool and send the result back

Append the assistant's tool-call turn and a tool message with your result, keyed by the tool_call_id, then call the endpoint again:

{
  "model_id": "llama-3.3-70b-versatile",
  "messages": [
    {"role": "user", "content": "What's the weather in Paris?"},
    {"role": "assistant", "content": "", "tool_calls": [
      {"id": "call_abc123", "type": "function",
       "function": {"name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}}
    ]},
    {"role": "tool", "tool_call_id": "call_abc123", "content": "18°C, partly cloudy"}
  ],
  "tools": [ /* same tools array */ ]
}

The model now replies with a normal answer (e.g. "It's 18°C and partly cloudy in Paris."), or requests another tool call — repeat until tool_calls is absent.

Streaming

With stream: true, tool calls arrive as SSE events of type: "tool_calls" (in addition to the usual delta and usage events). See Streaming.

Errors to handle

Chat completions can return:

  • 402 — out of tokens or subscription past due
  • 403 — the model isn't available on your plan (common on trial)
  • 429 — rate limited
  • 503 — the model is temporarily unavailable

See Errors & Rate Limits for the full list and retry guidance.