Continue.dev Setup: config.yaml for a Custom Endpoint
Guides

Continue.dev Setup: config.yaml for a Custom Endpoint

Point Continue at any OpenAI-compatible base URL using config.yaml — model roles, apiBase, request options, and why the autocomplete slot needs its own model.

Continue is an open-source assistant for VS Code and JetBrains that keeps its configuration in a file rather than behind a settings panel. That makes it one of the easier tools to point at a custom OpenAI-compatible endpoint, and one of the easier ones to check into a repository so a whole team gets the same setup.

The one thing to get right up front: Continue moved from JSON to YAML. config.json still loads, but it is the legacy format, and if a config.yaml is present it is used instead. Most stale tutorials you will find show the JSON shape. Write YAML.

Where configuration lives

There are three places, and they layer:

  • ~/.continue/config.yaml — your global configuration. On Windows, %USERPROFILE%\.continue\config.yaml.
  • ~/.continue/assistants/ — one YAML file per assistant, globally available. The filename becomes the assistant name in the picker, so Backend Reviewer.yaml shows up as "Backend Reviewer".
  • .continue/assistants/ in a repository root — the same thing, scoped to that workspace. This is the one to commit.

All three use the same schema, so you can develop a configuration globally and move the file into a repo later without editing it.

A minimal working configuration

Every config file starts with three header keys, then a models list:

name: cozy-setup
version: 0.0.1
schema: v1

models:
  - name: Kimi K3
    provider: openai
    model: kimi-k3
    apiBase: https://api.cozyapi.com/v1
    apiKey: cozy_your_key_here

Three of those keys carry the weight. provider: openai tells Continue to speak the OpenAI chat-completions shape — it does not mean the request goes to OpenAI. apiBase overrides the default endpoint for that model. model is the string sent in the request body, so it must be an alias your endpoint actually recognises.

name is only a label for the model picker. Make it readable; nothing depends on its value.

Roles decide which model does what

Continue does not have one model slot. Each entry declares which jobs it is eligible for through roles. The available roles are chat, edit, apply, autocomplete, embed, rerank and summarize. If you omit roles entirely, the model defaults to [chat, edit, apply, summarize].

That default is the source of the most common complaint about Continue, which is that inline completion feels sluggish. A large reasoning model in the autocomplete role is the wrong tool: autocomplete fires on almost every keystroke pause and needs a response in tens of milliseconds, not seconds. Split the roles explicitly:

name: cozy-setup
version: 0.0.1
schema: v1

models:
  - name: Kimi K3
    provider: openai
    model: kimi-k3
    apiBase: https://api.cozyapi.com/v1
    apiKey: cozy_your_key_here
    roles: [chat, edit, apply]
    defaultCompletionOptions:
      contextLength: 200000
      maxTokens: 8192

  - name: Fast helper
    provider: openai
    model: minimax-m2.7
    apiBase: https://api.cozyapi.com/v1
    apiKey: cozy_your_key_here
    roles: [summarize]

Note what is missing: an autocomplete entry. Inline completion wants a small model tuned for fill-in-the-middle, and running it locally against a hosted chat endpoint is usually the wrong trade. If your latency to a hosted endpoint is good and you want to try it, add a fast alias in the autocomplete role and judge it on how often you accept the suggestion, not on how good the suggestions look when you read them.

Tuning the request

Two blocks control what actually goes over the wire.

defaultCompletionOptions sets sampling and sizing: contextLength, maxTokens, temperature, topP, topK, stop, and reasoning controls including reasoning and reasoningBudgetTokens. Setting contextLength matters for custom aliases, because Continue otherwise has to assume a window and will trim your context earlier than necessary.

requestOptions handles transport: headers, extraBodyProperties, timeout, verifySsl and proxy. If your gateway wants an authentication header that is not a bearer token, this is where it goes:

    requestOptions:
      headers:
        X-Auth-Token: cozy_your_key_here
      timeout: 120

There is also a capabilities list — tool_use and image_input — which is worth setting explicitly for custom model names. Continue infers capability from known model identifiers, and an alias it does not recognise may be assumed to lack tool calling, which quietly disables agent behaviour.

One more flag to know about: useLegacyCompletionsEndpoint: true falls back to /completions rather than /chat/completions. You almost certainly do not want this against a modern gateway.

The rest of the file

Beyond models, a config file can carry rules, context, prompts, docs, mcpServers and data. Two of those are worth setting up on day one.

rules are always-on instructions — your conventions, your framework choices, the things you would otherwise retype into every chat. Putting them in a workspace assistant file means every developer on the repo gets the same behaviour without configuring anything.

mcpServers connects Model Context Protocol servers, which is how you give the assistant access to your database schema, issue tracker or internal documentation. That is a larger topic, but the configuration hook is here rather than in a separate file.

Committing a team configuration

The pattern that works: put a file at .continue/assistants/project-name.yaml in the repository with the models, roles and rules, and keep the API key out of it. Continue supports secret interpolation for hub-sourced blocks, and for local configuration the pragmatic approach is a per-developer global ~/.continue/config.yaml holding credentials, with the committed workspace file carrying everything else.

Test the endpoint independently before blaming the configuration. A single curl against your base URL with the same model string rules out half the possible problems in ten seconds.

Troubleshooting

  • Model does not appear in the picker — YAML indentation, or a missing schema: v1 header. Continue silently ignores a file it cannot parse.
  • 404 — apiBase should end at /v1. Continue appends the path itself.
  • Agent mode does nothing — add capabilities: [tool_use] so Continue knows the alias supports tool calling.
  • Context trimmed too aggressively — set contextLength under defaultCompletionOptions.
  • Old settings keep applying — you have a leftover config.json. A present config.yaml takes precedence, so if changes are not landing, check which file the extension is actually reading.

Because everything above is a base URL, a key and a model string, moving to a different provider is a three-line diff in one file. Keep the model names in configuration rather than habit and re-evaluating a new model is an afternoon rather than a project.

Common questions

Does Continue still support config.json?

Yes, but it is the legacy format and YAML is current. If a config.yaml exists it is loaded instead of config.json, which is a common reason edits appear to have no effect.

Why is my inline autocomplete slow in Continue?

Usually because a large chat model landed in the autocomplete role. Set roles explicitly on each model entry, and give autocomplete a small fast model — or leave the role unassigned rather than routing keystrokes to a reasoning model.

How do I share one Continue configuration across a team?

Put a YAML file in .continue/assistants/ at the repository root and commit it. The filename becomes the assistant name. Keep API keys in each developer global ~/.continue/config.yaml rather than in the committed file.

Similar articles

Cline Setup Guide: Pointing It at a Custom Endpoint
Guides
Guides·8 min read

Cline Setup Guide: Pointing It at a Custom Endpoint

Configure Cline against any OpenAI-compatible base URL — the provider fields, the model configuration block that people skip, and separate Plan and Act models.

Read
Kilo Code Setup Guide: Custom Endpoints and Profiles
Guides
Guides·8 min read

Kilo Code Setup Guide: Custom Endpoints and Profiles

Configure Kilo Code against your own OpenAI-compatible endpoint, set the model metadata it cannot infer, and assign different models per mode with profiles.

Read
Roo Code Setup: Custom Endpoint and Profiles Per Mode
Guides
Guides·8 min read

Roo Code Setup: Custom Endpoint and Profiles Per Mode

Configure Roo Code against an OpenAI-compatible base URL, then use API configuration profiles to give each mode its own model. Includes the tool-calling caveat.

Read