code values below are stable for the Token Factory
Alpha contract. New codes may be added without changing existing code meanings.
Quick reference
An HTTP
502 uses service_unavailable when the gateway cannot reach the model
backend and server_error when the backend returns 502. Both are retryable.
Categories
Error envelopes
The endpoint family decides the envelope shape (no content negotiation): the OpenAI-compatible endpoints return the OpenAI-style envelope below, and the Anthropic-compatible endpoints return the Anthropic-style envelope. The stablecode catalog is shared by both.
OpenAI endpoints
All Token Factory OpenAI-family inference endpoints (/v1/chat/completions,
/v1/completions, /v1/embeddings, /v1/models, /v1/responses) return errors as
Content-Type: application/json with this shape:
invalid_request_errors where the failing field is identifiable, the
gateway additionally emits an optional param:
code— stable, snake_case, machine-readable slug. Match on this in your client code. Codes do not change once published.message— human-readable string safe to surface in UIs or logs. The wording may change between releases — do not parse it.type— OpenAI-style category enum:invalid_request_error,authentication_error,permission_error,not_found_error,rate_limit_error,server_error,service_unavailable. Used by the OpenAI Python and TypeScript SDKs to pick the exception class (AuthenticationError,RateLimitError, etc.). Safe to ignore if you match oncodedirectly.param(optional) — forinvalid_request_errors, the request field at fault ("model","input","messages", etc.). Omitted when no specific field is identifiable, and omitted on non-validation errors. Mirrors OpenAI’s error-object shape so the OpenAI Python and TypeScript SDKs surface it onBadRequestError.param.- HTTP status — conveys the broad category (auth, throttling, server,
etc.). The
codeis the precise identity within that category. x-request-id(response header) — present on every response, including all errors (401/403/404 and every 4xx/5xx). Log it and quote it in support requests so a failure can be correlated end-to-end.
These endpoints return an OpenAI-compatible JSON envelope under
application/json, not an RFC 7807 application/problem+json document.Anthropic endpoints (/v1/messages)
The Anthropic-compatible endpoints (POST /v1/messages,
POST /v1/messages/count_tokens) return every 4xx/5xx as
Content-Type: application/json with the Anthropic-style envelope,
extended with the same stable code catalog:
error.type— Anthropic’s published category enum:invalid_request_error,authentication_error,permission_error,not_found_error,request_too_large,rate_limit_error,api_error,overloaded_error. The Anthropic Python and TypeScript SDKs use it (with the HTTP status) to pick the exception class (AuthenticationError,RateLimitError, etc.). Relative to the OpenAI-styletype, two values map:server_error→api_errorandservice_unavailable→overloaded_error; the rest are identical.error.code— the same stable, snake_case catalog documented on this page. Match on this for precise handling — e.g.rate_limitedvstoken_limitedboth surface asrate_limit_error, and onlycodedistinguishes them.error.message— human-readable; wording may change between releases. There is noparamfield on this surface — the failing request field is named in the message text instead.request_id— echoes yourX-Request-Idrequest header (also returned as a response header) so you can correlate to gateway logs. Omitted when the request carried noX-Request-Id.- HTTP status — unchanged from the table below; Token Factory does not adopt
Anthropic’s 529 (
overloaded_errorships on 503) or 413 (oversized requests surface as 400bad_request).
502/504 surfaces as api_error and a 503 as overloaded_error:
stream: true requests, failures that occur before the stream starts
return this JSON envelope with the error status (not an SSE body).
Codes
Each entry below documents:- Cause — what triggered this error.
- Remediation — what the API consumer should do.
- Retry — whether the SDK retries this automatically, and how.
- Where to look — dashboards, settings, or support flow to investigate.
invalid_api_key (401)
Cause. NoAuthorization: Bearer <key> header, malformed header, or
the key value is unknown to the gateway.
Remediation. Create or rotate an API key (sk-corvex-*) in the
dashboard and send it as Authorization: Bearer sk-corvex-... (or as an
x-api-key header on the Anthropic /v1/messages surface). See
Authentication for the key
types.
Retry. No. This is a non-retryable client error.
Where to look. Dashboard → API Keys. Revoked keys appear with status
disabled; expired keys must be replaced.
virtual_key_blocked (403)
Cause. The API key has been deactivated. Remediation. Create a new API key in the dashboard. If access remains blocked, contact Token Factory support. Retry. No. Where to look. Dashboard → API Keys → key status. Audit log will show the deactivation event.model_blocked (403)
Cause. The key is active, but the requested model is not among the models allowed for your account or key. Remediation. Use a model your account is allowed to call (list them withGET /v1/models), or contact
Token Factory support about model access.
Retry. No.
Where to look. Dashboard → API Keys → key detail → Allowed models.
model_unavailable (404)
The HTTP status determines how to handle this code. A 404 means the model is
not in the public catalog and is not retryable. A 503 means the model is in the
catalog but temporarily not serving; retry that response with bounded
exponential backoff.
some-org/Nonexistent-Model) is not in the public catalog.
Remediation. List available models via GET /v1/models and pick a
model returned by the gateway.
Retry. No for this 404 response.
Where to look. The models list endpoint: GET /v1/models.
not_found (404)
Cause. The request reached a valid inference surface prefix (/v1/…
or /openai/v1/…) but no endpoint matched — e.g. GET /v1/models/{id}
or a typo like GET /v1/bogus. Distinct from model_unavailable, which
means the endpoint matched but the model is not serveable.
Remediation. Check the path against the API
reference. Unknown /v1/* paths return
this structured envelope rather than a bare string, so SDK error
parsing stays intact.
Retry. No. This is a non-retryable client error.
Where to look. Your request URL. The x-request-id response header
correlates the 404 in the gateway logs.
rate_limited (429)
Cause. Your account’s request-rate limit or an API key’s concurrent-request limit was reached, or the requested model is temporarily at capacity. Remediation. Slow down. If the response includes aRetry-After
header (in seconds), wait that long before retrying. Otherwise back off
exponentially.
Retry. Yes. Honor Retry-After when present and otherwise use
exponential backoff. SDK defaults vary by package and version; configure the
attempt count explicitly for production clients.
Where to look. Check your account limits
in the dashboard. To request higher limits, contact
Token Factory support.
token_limited (429)
Cause. A configured token-usage limit for your account was exceeded. This limit is separate from the request-rate limit. Remediation. Same asrate_limited. The response code is
token_limited (not rate_limited), so you can branch on which limit was
hit; the SDKs surface both as a RateLimitError.
Retry. Yes, same retry policy as rate_limited. Honor
Retry-After.
Where to look. Review token usage in the dashboard. For questions about
a configured token limit, contact
Token Factory support.
insufficient_credits (402)
Cause. Your account’s available credit balance has reached zero, a hard quota has been exhausted, or a configured monthly spend cap has been reached. Remediation. For an exhausted credit balance, add credits through the dashboard. For quota or spend-cap exhaustion, contact Token Factory support or wait for the configured reset window. Retry. No. Retrying before the billing or quota condition is resolved will return the same error. Match oncode (insufficient_credits), not type: on the OpenAI surface
this carries type: permission_error (there is no billing-specific type).
Where to look. Check the dashboard’s billing and usage sections, or contact
Token Factory support.
bad_request (400)
Cause. The request body is malformed, missing required fields, or fails server-side validation. Examples: missingmodel, missing
messages for chat completions, or invalid JSON.
Common validation cases:
Integer-valued numbers such as
5.0 and 1e3, and the value 0, are accepted.
Remediation. Read the message for the specific field at fault and
correct the request. For the legacy-completions case, call
/v1/chat/completions (recommended), or switch the model to a
text-completion model. The API reference documents the
public request schemas.
Retry. No. Retrying the same payload will return the same error.
Where to look. Validate the request against the public OpenAPI schema and
call GET /v1/models to check model capabilities.
context_length_exceeded (400)
Cause. The request’s input (or input +max_tokens) exceeds the model’s
context window. The gateway surfaces this as a stable code whenever the
upstream engine reports a context-length rejection — on both endpoint
families, and including engine failures initially reported as a 5xx that the
gateway downgrades to the spec-correct 400. The message carries the
engine’s own token accounting (e.g. the model’s maximum context length and
your requested totals) where the engine reports it.
Remediation. Reduce the input: compact the conversation, drop or truncate
older turns or tool outputs, or lower max_tokens. For agentic clients
(Claude Code and similar harnesses), treat this code as the compact-and-retry
signal. To size a session up front, read the model’s limits from
GET /v1/models: context_window (total), max_output_tokens (the model’s
completion cap), and — only for a model whose output cap sits below its
window — max_input_tokens (context_window − max_output_tokens). That last
value is the input ceiling for a request that reserves the full output cap; a
request with a smaller max_tokens has correspondingly more input room, since
the engine enforces input + max_tokens ≤ context_window.
Retry. No. Retrying the same payload returns the same error. Retry only
after compacting or otherwise shrinking the input.
Where to look. The model’s limits on GET /v1/models (context_window,
max_output_tokens, max_input_tokens), and the message field for the
engine’s token accounting.
OpenAI envelope:
server_error (500, 502)
Cause. An unexpected server-side failure: a downstream component crashed, a database call failed, or the gateway hit an internal panic. On HTTP 502, this code means the upstream engine itself answered 502 and the gateway normalized its error body into the standard envelope. Remediation. Retry with exponential backoff. If the error persists after the configured retries, capture a sample response and itsx-request-id header
before contacting support.
Retry. Yes. Use exponential backoff and a bounded retry count. SDK
defaults vary by package and version.
Where to look. Quote the x-request-id response header in a support
request so the failure can be correlated.
service_unavailable (502, 503)
Cause. The platform is temporarily over capacity, or the requested model is still starting up and not yet ready to serve. On HTTP 502, the gateway could not reach the model backend at all. Remediation. Retry with exponential backoff. HonorRetry-After when
present.
Retry. Yes, same policy as server_error.
Where to look. Confirm that GET /v1/models lists the model. If it does not,
choose an available model; if it does, retry after the temporary failure.
Retry semantics
The official OpenAI and Anthropic SDKs include automatic retry behavior, but the exact statuses, delays, and attempt counts depend on the SDK and version. For Token Factory responses:- Retry: HTTP
429(rate_limited,token_limited) and5xxresponses, including a temporary503 model_unavailable. - Do not retry: Other
4xxresponses, includingbad_request,context_length_exceeded,invalid_api_key,virtual_key_blocked,model_blocked, the404form ofmodel_unavailable, andinsufficient_credits. - Backoff. Use exponential backoff with a bounded maximum delay and attempt count. Configure these values for the latency and reliability needs of your application.
Retry-Afterheader. When the server returnsRetry-After, the SDK waits at least that long before retrying. The header takes precedence over the default exponential backoff.
rate_limited and token_limited are surfaced by the SDKs as a
RateLimitError (a subclass of APIError). To tell which limit was hit,
match on the wire error.code (rate_limited vs token_limited) — the
OpenAI error.type is rate_limit_error for both. Use the Retry-After
response header to decide how long to wait.
Prompt caching (cache_control)
The Anthropic Messages surface (POST /v1/messages,
POST /v1/messages/count_tokens) accepts cache_control annotations on system
and message content blocks. Token Factory does not interpret the annotation as a
request to cache content.
When the model backend reports cache usage, Token Factory passes those values
through. When the backend omits them, the response includes explicit zero values
in the non-streaming usage object or streaming message_start event:
Related documentation
- DeepSeek Harness — setup and troubleshooting.
- API Reference — endpoint-by-endpoint request/response schemas, including error responses per endpoint.
- Authentication — virtual
keys (
sk-corvex-*) for inference. - Quickstart — first inference call end-to-end.