Chat Completions
POSThttps://api.stdcmpt.com/v1/chat/completions
Generates a response from a list of messages. This is the most widely supported endpoint and works with any OpenAI-compatible client.
Request
curl https://api.stdcmpt.com/v1/chat/completions \
-H "Authorization: Bearer $STANDARDCOMPUTE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "standardcompute",
"messages": [
{ "role": "system", "content": "You are a concise assistant." },
{ "role": "user", "content": "Explain what a mutex is." }
],
"max_tokens": 500
}'
Parameters
The request body follows the OpenAI Chat Completions format. The most common fields:
| Field | Type | Description |
|---|---|---|
model | string | standardcompute for smart routing, or an exact model ID. Defaults to smart routing if omitted. |
messages | array | The conversation. Supports system, user, assistant and tool roles, and image content parts on vision models. |
max_tokens / max_completion_tokens | integer | Maximum output tokens. Capped at your plan's output limit. |
stream | boolean | Stream the response as server-sent events. |
tools, tool_choice | array, string/object | Function calling. |
response_format | object | Request JSON output. |
reasoning_effort, reasoning | string, object | Control reasoning. See Reasoning. |
temperature, top_p, stop | Standard sampling controls. |
Response
{
"id": "chatcmpl-...",
"object": "chat.completion",
"model": "StandardCompute",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "A mutex is..." },
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 24, "completion_tokens": 88, "total_tokens": 112 }
}
With smart routing, model is StandardCompute. With a pinned model, it is the full model ID.
Streaming
Set "stream": true to receive server-sent events. Each event is a chat.completion.chunk, and the stream ends with data: [DONE]. The final chunk includes usage.
stream = client.chat.completions.create(
model="standardcompute",
messages=[{"role": "user", "content": "Write a haiku about compilers."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
If an error happens after streaming has started, it arrives as an error event inside the stream rather than as an HTTP status code.