Prerequisites
- A Token Factory API key. See Authentication.
- Kilo Code installed in VS Code, JetBrains, or as the CLI.
- A model from the Token Factory catalog.
1
Add a custom provider
Open the Kilo Code settings (gear icon), choose Global Config, select
the Providers tab, scroll to the bottom, and choose Custom provider.
2
Enter the Token Factory details
After you enter the Base URL, Kilo Code fetches available models from
GET /v1/models. Under Models, choose zai-org/GLM-5.3,
deepseek-ai/DeepSeek-V4-Flash-0731, or both. Leave the reasoning
toggle on and the image toggle off for each model, then select
Submit. If a model does not appear, add its catalog ID manually.3
Set the model limits
Kilo Code does not read context limits from the endpoint and assumes a
256K window. Open the global config at Kilo Code reads this file when it starts, so reload the VS Code window
(Developer: Reload Window) after saving. The output limit is a
per-response budget and is sent as
~/.config/kilo/kilo.jsonc. The
provider you saved in the previous step is already there under
provider.corvex; add limit to each model you selected:max_tokens; leave room in the
context window for the request.Neither model accepts image input.4
Select the model and verify the connection
New sessions start on Auto Free, Kilo Code’s own model router, which
does not use Token Factory. Select the model button next to the mode
selector under the chat box, find the Corvex Token Factory group, and
choose
zai-org/GLM-5.3. The button then shows the model ID and the
context bar shows the configured window (393.2K).Start a task that reads a file in your repository. A completed Read
tool call confirms the provider configuration and key.Compatibility notes
- Model IDs. Use the exact
idfromGET /v1/models. IDs are case-sensitive and usenamespace/modelformat. Do not add acorvex/prefix. - Context length. If a long task exceeds the model’s window, the gateway
returns
context_length_exceeded. Start a new task or shorten the context. See Errors. - Reasoning. Both models reason before answering. Reasoning tokens count toward the output budget.
Related documentation
- Connect Cline — the same provider settings in a related extension.
- Use the OpenAI SDK — the same endpoint from Python or TypeScript.
- Errors — authentication, rate-limit, and validation responses.