Skip to main content

Chat Completions

POSThttps://api.stdcmpt.com/v1/chat/completions

Generates a response from a list of messages. This is the most widely supported endpoint and works with any OpenAI-compatible client.

Request​

curl https://api.stdcmpt.com/v1/chat/completions \
-H "Authorization: Bearer $STANDARDCOMPUTE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "standardcompute",
"messages": [
{ "role": "system", "content": "You are a concise assistant." },
{ "role": "user", "content": "Explain what a mutex is." }
],
"max_tokens": 500
}'

Parameters​

The request body follows the OpenAI Chat Completions format. The most common fields:

FieldTypeDescription
modelstringstandardcompute for smart routing, or an exact model ID. Defaults to smart routing if omitted.
messagesarrayThe conversation. Supports system, user, assistant and tool roles, and image content parts on vision models.
max_tokens / max_completion_tokensintegerMaximum output tokens. Capped at your plan's output limit.
streambooleanStream the response as server-sent events.
tools, tool_choicearray, string/objectFunction calling.
response_formatobjectRequest JSON output.
reasoning_effort, reasoningstring, objectControl reasoning. See Reasoning.
temperature, top_p, stopStandard sampling controls.

Response​

{
"id": "chatcmpl-...",
"object": "chat.completion",
"model": "StandardCompute",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "A mutex is..." },
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 24, "completion_tokens": 88, "total_tokens": 112 }
}

With smart routing, model is StandardCompute. With a pinned model, it is the full model ID.

Streaming​

Set "stream": true to receive server-sent events. Each event is a chat.completion.chunk, and the stream ends with data: [DONE]. The final chunk includes usage.

stream = client.chat.completions.create(
model="standardcompute",
messages=[{"role": "user", "content": "Write a haiku about compilers."}],
stream=True,
)

for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")

If an error happens after streaming has started, it arrives as an error event inside the stream rather than as an HTTP status code.