Prerequisites
- A Token Factory API key. See Authentication.
- Cline installed in VS Code, Cursor, JetBrains, or another supported IDE.
- A model ID from the Token Factory catalog.
1
Choose your own API key
On first launch, Cline shows an onboarding screen titled How will you
use Cline?. Select Bring my own API key and then Continue. The
settings (gear) icon is inactive until onboarding is complete.If Cline is already configured, open the Cline panel, select the settings
(gear) icon, and go to API Configuration.
2
Enter the Token Factory endpoint, key, and model
Enter the model ID exactly as the catalog lists it, including the
namespace/ prefix. Do not add a corvex/ prefix.3
Set the model limits
Cline does not read context limits from the endpoint and defaults to a
128,000-token window. Expand Model Configuration and set the values for
the model you selected.
Supports Images is checked by default; clear it, because neither model
accepts image input. Leave Enable R1 messages format off. The Max
Output Tokens value is a per-response budget; keep it well below the
context window so there is room for input and tool results. The price
fields only affect Cline’s local cost display; the published rates are on
the Pricing page.
4
Verify the connection
Return to the chat, confirm the context bar shows the configured window
(for example
393.2k), and start a task in Act mode that reads a file
in your repository. A completed tool call confirms the endpoint, key, and
model ID.Compatibility notes
- Model IDs. Use the exact
idfromGET /v1/models. IDs are case-sensitive. See Models. - Context length. If a long task exceeds the model’s window, the gateway
returns
context_length_exceeded. Start a new task or shorten the context. See Errors. - Reasoning. Both models reason before answering; Cline shows the reasoning under Thinking. Reasoning tokens count toward the Max Output Tokens value you configured. The Reasoning Effort setting is sent with the request; Token Factory applies the model’s default reasoning behavior.
Related documentation
- Use the OpenAI SDK — the same endpoint from Python or TypeScript.
- Connect Kilo Code — the same provider settings in a related extension.
- Errors — authentication, rate-limit, and validation responses.