Continue.dev Setup: config.yaml for a Custom Endpoint
Point Continue at any OpenAI-compatible base URL using config.yaml — model roles, apiBase, request options, and why the autocomplete slot needs its own model.
Continue is an open-source assistant for VS Code and JetBrains that keeps its configuration in a file rather than behind a settings panel. That makes it one of the easier tools to point at a custom OpenAI-compatible endpoint, and one of the easier ones to check into a repository so a whole team gets the same setup.
The one thing to get right up front: Continue moved from JSON to YAML. config.json still loads, but it is the legacy format, and if a config.yaml is present it is used instead. Most stale tutorials you will find show the JSON shape. Write YAML.
Where configuration lives
There are three places, and they layer:
~/.continue/config.yaml— your global configuration. On Windows,%USERPROFILE%\.continue\config.yaml.~/.continue/assistants/— one YAML file per assistant, globally available. The filename becomes the assistant name in the picker, soBackend Reviewer.yamlshows up as "Backend Reviewer"..continue/assistants/in a repository root — the same thing, scoped to that workspace. This is the one to commit.
All three use the same schema, so you can develop a configuration globally and move the file into a repo later without editing it.
A minimal working configuration
Every config file starts with three header keys, then a models list:
name: cozy-setup
version: 0.0.1
schema: v1
models:
- name: Kimi K3
provider: openai
model: kimi-k3
apiBase: https://api.cozyapi.com/v1
apiKey: cozy_your_key_here
Three of those keys carry the weight. provider: openai tells Continue to speak the OpenAI chat-completions shape — it does not mean the request goes to OpenAI. apiBase overrides the default endpoint for that model. model is the string sent in the request body, so it must be an alias your endpoint actually recognises.
name is only a label for the model picker. Make it readable; nothing depends on its value.
Roles decide which model does what
Continue does not have one model slot. Each entry declares which jobs it is eligible for through roles. The available roles are chat, edit, apply, autocomplete, embed, rerank and summarize. If you omit roles entirely, the model defaults to [chat, edit, apply, summarize].
That default is the source of the most common complaint about Continue, which is that inline completion feels sluggish. A large reasoning model in the autocomplete role is the wrong tool: autocomplete fires on almost every keystroke pause and needs a response in tens of milliseconds, not seconds. Split the roles explicitly:
name: cozy-setup
version: 0.0.1
schema: v1
models:
- name: Kimi K3
provider: openai
model: kimi-k3
apiBase: https://api.cozyapi.com/v1
apiKey: cozy_your_key_here
roles: [chat, edit, apply]
defaultCompletionOptions:
contextLength: 200000
maxTokens: 8192
- name: Fast helper
provider: openai
model: minimax-m2.7
apiBase: https://api.cozyapi.com/v1
apiKey: cozy_your_key_here
roles: [summarize]
Note what is missing: an autocomplete entry. Inline completion wants a small model tuned for fill-in-the-middle, and running it locally against a hosted chat endpoint is usually the wrong trade. If your latency to a hosted endpoint is good and you want to try it, add a fast alias in the autocomplete role and judge it on how often you accept the suggestion, not on how good the suggestions look when you read them.
Tuning the request
Two blocks control what actually goes over the wire.
defaultCompletionOptions sets sampling and sizing: contextLength, maxTokens, temperature, topP, topK, stop, and reasoning controls including reasoning and reasoningBudgetTokens. Setting contextLength matters for custom aliases, because Continue otherwise has to assume a window and will trim your context earlier than necessary.
requestOptions handles transport: headers, extraBodyProperties, timeout, verifySsl and proxy. If your gateway wants an authentication header that is not a bearer token, this is where it goes:
requestOptions:
headers:
X-Auth-Token: cozy_your_key_here
timeout: 120
There is also a capabilities list — tool_use and image_input — which is worth setting explicitly for custom model names. Continue infers capability from known model identifiers, and an alias it does not recognise may be assumed to lack tool calling, which quietly disables agent behaviour.
One more flag to know about: useLegacyCompletionsEndpoint: true falls back to /completions rather than /chat/completions. You almost certainly do not want this against a modern gateway.
The rest of the file
Beyond models, a config file can carry rules, context, prompts, docs, mcpServers and data. Two of those are worth setting up on day one.
rules are always-on instructions — your conventions, your framework choices, the things you would otherwise retype into every chat. Putting them in a workspace assistant file means every developer on the repo gets the same behaviour without configuring anything.
mcpServers connects Model Context Protocol servers, which is how you give the assistant access to your database schema, issue tracker or internal documentation. That is a larger topic, but the configuration hook is here rather than in a separate file.
Committing a team configuration
The pattern that works: put a file at .continue/assistants/project-name.yaml in the repository with the models, roles and rules, and keep the API key out of it. Continue supports secret interpolation for hub-sourced blocks, and for local configuration the pragmatic approach is a per-developer global ~/.continue/config.yaml holding credentials, with the committed workspace file carrying everything else.
Test the endpoint independently before blaming the configuration. A single curl against your base URL with the same model string rules out half the possible problems in ten seconds.
Troubleshooting
- Model does not appear in the picker — YAML indentation, or a missing
schema: v1header. Continue silently ignores a file it cannot parse. - 404 —
apiBaseshould end at/v1. Continue appends the path itself. - Agent mode does nothing — add
capabilities: [tool_use]so Continue knows the alias supports tool calling. - Context trimmed too aggressively — set
contextLengthunderdefaultCompletionOptions. - Old settings keep applying — you have a leftover
config.json. A presentconfig.yamltakes precedence, so if changes are not landing, check which file the extension is actually reading.
Because everything above is a base URL, a key and a model string, moving to a different provider is a three-line diff in one file. Keep the model names in configuration rather than habit and re-evaluating a new model is an afternoon rather than a project.
Common questions
Does Continue still support config.json?
Yes, but it is the legacy format and YAML is current. If a config.yaml exists it is loaded instead of config.json, which is a common reason edits appear to have no effect.
Why is my inline autocomplete slow in Continue?
Usually because a large chat model landed in the autocomplete role. Set roles explicitly on each model entry, and give autocomplete a small fast model — or leave the role unassigned rather than routing keystrokes to a reasoning model.
How do I share one Continue configuration across a team?
Put a YAML file in .continue/assistants/ at the repository root and commit it. The filename becomes the assistant name. Keep API keys in each developer global ~/.continue/config.yaml rather than in the committed file.