~/.codex/config.toml, point it at the API key in your environment,
and select a Token Factory model. The Codex CLI and the Codex IDE extension read
the same configuration file.
Prerequisites
- A Token Factory API key. See Authentication.
- Codex CLI
installed (
npm install -g @openai/codex). Codex requires a Responses API provider; a provider configured withwire_api = "chat"fails at startup. - A model ID from the Token Factory catalog.
1
Store the key in your shell
Codex reads the API key from the environment variable named by
env_key.
Export it in your shell profile so both the CLI and the IDE extension can
read it.2
Add the provider to config.toml
Create or edit Codex does not read context limits from the endpoint, so
~/.codex/config.toml:model_context_window declares the selected model’s window. Use the value
for the model you chose:Enter the model ID exactly as the catalog lists it, including the
namespace/ prefix. Do not add a corvex/ prefix.3
Run Codex
Start Codex in your repository. The provider and model from To run a single task non-interactively:To switch models for one session without editing the file:
config.toml
are used unless you override them on the command line.4
Verify the connection
In Codex, run
/status to confirm the provider shows Corvex Token
Factory and the model ID you configured. Then ask Codex to read a file in
the repository. A completed tool call confirms the endpoint, key, and model.Multiple models as profiles
Profiles let you keep one provider and switch models with--profile:
How Token Factory serves the Responses API
POST /v1/responses accepts the Codex request shape, including instructions,
input items, tools, tool_choice, max_output_tokens, temperature, and
stream. The gateway translates each request to the model’s Chat Completions
interface and returns Responses-shaped output, with streaming events in the
Responses event order. The model’s reasoning is returned as a reasoning output
item ahead of the message item.
Conversation persistence is not exposed: store and previous_response_id
have no effect, and there is no /v1/conversations endpoint. Codex resends the
full conversation on each turn, so this does not affect normal use. Request
fields that control OpenAI-hosted reasoning (reasoning, include) are
accepted and ignored; the model reasons with its default settings.
See the API reference for the documented request and
response fields.
Compatibility notes
- Context length. If a session exceeds the model’s window, the gateway
returns
context_length_exceeded. Run/compactin Codex, or start a new session. See Errors. - Output budget. Reasoning tokens count toward
max_output_tokens. If a response is truncated, shorten the prompt or split the task. - Input types. Neither model accepts image input. Do not attach images with
-i; the request returns a structured error.
Related documentation
- Use the OpenAI SDK — Chat Completions from Python or TypeScript.
- Connect OpenCode — another terminal coding agent on the OpenAI-compatible surface.
- Errors — authentication, rate-limit, and validation responses.