CLAUDE_CONFIG_DIR to keep Token Factory sessions and settings in
~/.claude-corvex, separate from your default Claude Code profile.
Prerequisites
- A Token Factory API key (
sk-corvex-...) from the dashboard. See Authentication. - A model name from the Models catalog (this guide uses
zai-org/GLM-5.3as the example). - Claude Code installed and
jqavailable.
1
Store your key and endpoint in a private env file
Replace
sk-corvex-... with your key from the dashboard.2
Create the launcher
The
claude-corvex launcher sources the private environment file, selects
the Token Factory profile, and passes its arguments to claude.The script installs to ~/.local/bin so you can invoke it as claude-corvex
from any directory. That directory is on PATH by default on most Linux
distributions; on macOS, add export PATH="${HOME}/.local/bin:${PATH}" to your
shell profile (e.g. ~/.zshrc) and restart your shell.3
Verify the API key
Call the Anthropic-compatible token-counting endpoint. This validates the
key and request shape without running inference.
4
Run Claude Code through Token Factory
Context length
Claude Code does not auto-detect the context window of a model reached through a custom endpoint. It limits these sessions to 200,000 tokens, even when the model supports a larger native window. Two limits bound a session, and the compaction window must respect both:- Claude Code’s 200K session limit —
CLAUDE_CODE_AUTO_COMPACT_WINDOWcan never usefully exceed200000. - The model’s context window minus the request’s output reservation. The
engine enforces
input + max_tokens ≤ context_windowon every request, and Claude Code sendsmax_tokens= its max output setting (32,000 by default;CLAUDE_CODE_MAX_OUTPUT_TOKENS). Tool-heavy turns also overshoot the compaction trigger by roughly 15% before compaction runs, so size the window as(context_window − max_tokens) / 1.15, then cap it at200000.
For both models, the 200K session limit binds with the default 32,000-token output reservation. If you raise
CLAUDE_CODE_MAX_OUTPUT_TOKENS, recalculate the compaction window using the formula above.
Set CLAUDE_AUTOCOMPACT_PCT_OVERRIDE to the percentage of the window at which
compaction should begin; the example uses 90. Confirm the effective window
with /context. To use the full native window, call the API directly. When you
switch models, update CORVEX_CLAUDE_MODEL and re-derive the compaction
window from the table or the formula above.
GET /v1/models reports each model’s context_window and max_output_tokens.
If a model carries an output cap below its window, the response also includes
max_input_tokens (context_window − max_output_tokens): the input ceiling for
a request that reserves the full output cap. Use the smaller of that value and
context_window − max_tokens as the input room in the formula. If a session
does hit the model’s limit anyway, the gateway returns context_length_exceeded
(see Errors) — compact and
retry.
Optional: migrate plugins, settings, and MCP servers
Optional: migrate plugins, settings, and MCP servers
The launcher starts with a clean Token Factory profile. Run this
Run it:
claude-corvex-migrate script to copy your
plugins, settings, CLAUDE.md, and MCP servers from your normal Claude profile
into ~/.claude-corvex. Run it again after changing your default profile.
A failed optional copy step warns and continues.The first part copies named paths from ~/.claude/ via an rsync include-list.
It needs rsync and is skipped with a warning if rsync is absent. The
second part merges only the mcpServers key from ~/.claude.json into the
Token Factory profile. It needs jq. The first run may be slow if your
plugin tree is large.Subcommands
The launcher passes arguments straight through toclaude and exports
CLAUDE_CONFIG_DIR, so mcp, plugin, config, and other subcommands target
~/.claude-corvex automatically.
claude mcp addrun through this launcher writes to the Token Factory profile (~/.claude-corvex/.claude.json), not the default profile.- Use
-s user(orlocal) for personal config.-s projectwrites a.mcp.jsoninto the current project dir, which is shared across both profiles.
Compatibility notes
- Prompt caching. The Messages API accepts
cache_control, but Token Factory does not treat it as a request to cache content. Engine-reported cache usage is passed through; missing cache-usage fields are reported as zero. See Errors: Prompt caching for details. - Model capabilities. Choose a model whose input types and features match the Claude Code task. Unsupported inputs return a structured error; compare the available models on the Models page.