Prerequisites
- A Token Factory API key. See Authentication.
- Python with the
anthropicpackage, or Node.js with the@anthropic-ai/sdkpackage. - A model ID from the live catalog.
Configure the client
/v1 suffix; the SDK appends /v1/messages itself and sends the key
in the x-api-key header. max_tokens is required by the Messages API.
Both Token Factory models reason before answering, so content starts with a
thinking block followed by the text block. Select the block by type rather
than by position, as the examples do. Reasoning tokens count toward
max_tokens.
Run the repository examples
2xx response raises the SDK’s API error type with the Anthropic-shaped
Token Factory error body. See
Errors: Anthropic endpoints.
Streaming and tool use
The endpoint acceptsstream: true and returns the Anthropic Server-Sent Events
sequence (message_start, content_block_delta, …, message_stop), so the
SDK’s streaming helpers work as they do against Anthropic. Tool definitions and
tool_use / tool_result content blocks are accepted for both Token Factory
models.
Count tokens
POST /v1/messages/count_tokens validates the key and request shape without
running inference:
Compatibility notes
- Model IDs. Use the Token Factory
namespace/modelID exactly asGET /v1/modelsreturns it. Anthropic model names are not served. - Prompt caching. The endpoint accepts
cache_controlbut Token Factory does not treat it as a request to cache content. See Errors: Prompt caching. - Input types. Neither model accepts image input. Unsupported inputs return a structured error.
Related documentation
- Connect Claude Code — the same Messages API from the Claude Code CLI.
- Use the OpenAI SDK — the Chat Completions surface for OpenAI clients.
- API reference — supported Messages API request and response fields.