Skip to main content
Corvex Token Factory exposes an OpenAI Responses API endpoint, which is the protocol Codex requires from a custom model provider. Add Token Factory as a provider in ~/.codex/config.toml, point it at the API key in your environment, and select a Token Factory model. The Codex CLI and the Codex IDE extension read the same configuration file.

Prerequisites

  • A Token Factory API key. See Authentication.
  • Codex CLI installed (npm install -g @openai/codex). Codex requires a Responses API provider; a provider configured with wire_api = "chat" fails at startup.
  • A model ID from the Token Factory catalog.
1

Store the key in your shell

Codex reads the API key from the environment variable named by env_key. Export it in your shell profile so both the CLI and the IDE extension can read it.
2

Add the provider to config.toml

Create or edit ~/.codex/config.toml:
Codex does not read context limits from the endpoint, so model_context_window declares the selected model’s window. Use the value for the model you chose:Enter the model ID exactly as the catalog lists it, including the namespace/ prefix. Do not add a corvex/ prefix.
3

Run Codex

Start Codex in your repository. The provider and model from config.toml are used unless you override them on the command line.
To run a single task non-interactively:
To switch models for one session without editing the file:
4

Verify the connection

In Codex, run /status to confirm the provider shows Corvex Token Factory and the model ID you configured. Then ask Codex to read a file in the repository. A completed tool call confirms the endpoint, key, and model.

Multiple models as profiles

Profiles let you keep one provider and switch models with --profile:

How Token Factory serves the Responses API

POST /v1/responses accepts the Codex request shape, including instructions, input items, tools, tool_choice, max_output_tokens, temperature, and stream. The gateway translates each request to the model’s Chat Completions interface and returns Responses-shaped output, with streaming events in the Responses event order. The model’s reasoning is returned as a reasoning output item ahead of the message item. Conversation persistence is not exposed: store and previous_response_id have no effect, and there is no /v1/conversations endpoint. Codex resends the full conversation on each turn, so this does not affect normal use. Request fields that control OpenAI-hosted reasoning (reasoning, include) are accepted and ignored; the model reasons with its default settings. See the API reference for the documented request and response fields.

Compatibility notes

  • Context length. If a session exceeds the model’s window, the gateway returns context_length_exceeded. Run /compact in Codex, or start a new session. See Errors.
  • Output budget. Reasoning tokens count toward max_output_tokens. If a response is truncated, shorten the prompt or split the task.
  • Input types. Neither model accepts image input. Do not attach images with -i; the request returns a structured error.
  • Use the OpenAI SDK — Chat Completions from Python or TypeScript.
  • Connect OpenCode — another terminal coding agent on the OpenAI-compatible surface.
  • Errors — authentication, rate-limit, and validation responses.
Last modified on September 28, 2026