Roo Code Setup: Custom Endpoint and Profiles Per Mode
Guides

Roo Code Setup: Custom Endpoint and Profiles Per Mode

Configure Roo Code against an OpenAI-compatible base URL, then use API configuration profiles to give each mode its own model. Includes the tool-calling caveat.

Roo Code is a VS Code agent with a feature most of its peers lack: named API configuration profiles that can be bound to individual modes. That turns model routing from a manual decision into something the tool does for you, which is the main reason to spend more than five minutes on its settings.

Before any of that, one caveat that determines whether Roo Code will work with your endpoint at all.

The tool-calling requirement

Roo Code uses native tool calling exclusively. Its documentation is blunt about it: this is the only supported tool protocol, there is no XML-based fallback, and a model that does not support native tool calling cannot be used with Roo Code.

For a gateway or proxy, that means OpenAI function calling must be implemented properly end to end — the tools parameter passed upstream, tool_calls returned in the response, and tool-call deltas handled correctly when streaming. A provider that accepts tools and silently ignores it will produce an agent that appears to start and then does nothing.

Test this before configuring anything else. A single curl with a tools array, checking that the response contains a tool_calls block, takes thirty seconds and saves an hour of confusion.

Connecting to a custom endpoint

Open the Roo Code settings panel and, under Providers:

  1. Set API Provider to OpenAI Compatible.
  2. Set Base URL to https://api.cozyapi.com/v1. The documentation flags this field as crucial; it should end at /v1 exactly once.
  3. Set API Key to your key, for example cozy_your_key_here. Keys are held in the VS Code secret storage rather than in plain-text settings.
  4. Set Model to an alias your endpoint serves — kimi-k3, glm-5.2, deepseek-v4-pro and so on.

Then the Model Configuration block underneath, which is the part that actually shapes behaviour: Max Output Tokens, Context Window Size, Image Support, Prompt Caching, Input Price, Output Price, and cache read and write prices. There is a Reset to Defaults button if you make a mess of it.

Set Context Window Size to the real window of the model you selected. Roo Code cannot know it for an arbitrary alias, and an incorrect value means either premature truncation of your conversation or upstream errors partway through a task.

Max Output Tokens accepts -1 to let the server decide, which is often the right answer against a gateway that already enforces sensible limits.

The switches that fix odd behaviour

Several other toggles live on the same screen, and each exists because some provider needed it:

  • Enable streaming — on by default. Turn it off only to isolate a streaming bug.
  • Include max output tokens — controls whether the max output parameter is sent at all, since some providers reject it.
  • Enable R1 model parameters — documented as required for R1-style reasoning models such as QWQ, to prevent 400 errors. If a reasoning model returns 400 on every request, this is the first thing to try.
  • Enable Reasoning Effort — exposes a reasoning effort setting for models that accept one.
  • Custom Headers — name and value pairs added to every request, for gateways that want something other than a bearer token.
  • Use Azure and Set Azure API version — ignore unless you are on Azure.

One thing to be aware of: the published provider documentation still lists a Computer Use checkbox, but recent releases have removed it from this screen. If you are following an older guide and cannot find the field, that is why.

Profiles are the feature worth using

An API Configuration Profile is a named bundle: provider, credentials, model, temperature, thinking budget, diff settings, rate limits. Create one with the plus button in the Providers section, rename with the pencil, delete with the bin, and pin your favourites.

You switch profiles from the dropdown in Settings or, more usefully, from the API Configuration dropdown in the chat interface itself — so changing model mid-task is one click rather than a settings excursion.

A setup that works well in practice is three profiles pointing at the same base URL and key, differing only in model and sampling:

  • Heavy — a strong model for architecture and hard debugging.
  • Standard — your default for feature work.
  • Fast — a cheap model for renames, boilerplate, docstrings and commit messages.

Because they share credentials, creating them costs nothing beyond a few seconds each, and the routing discipline they enable is where the actual saving comes from.

Binding profiles to modes

Roo Code ships five built-in modes — code, architect, ask, debug and orchestrator — and lets you associate a specific configuration profile with each one from the Prompts tab. It also remembers the last profile you used per mode, so the association forms on its own if you never configure it explicitly.

The obvious mapping: your Heavy profile on architect and debug, Standard on code, Fast on ask. Architecture and debugging are where model quality most changes the outcome; answering a question about the codebase rarely is.

One limitation worth knowing before you plan around it: the mode-to-profile binding lives in your global VS Code state, not in a repository file. You cannot commit it and have teammates inherit it.

Custom modes and what they can carry

Custom modes are defined in a .roomodes file at your repository root, in YAML or JSON, and take precedence over globally defined modes. The fields are slug (letters, numbers and hyphens only), name, roleDefinition, optional whenToUse, description and customInstructions, plus groups.

groups is the permission model: read, edit, command and mcp. The edit entry can carry a file regex, so a mode can be restricted to only touch certain paths:

# .roomodes
customModes:
  - slug: test-writer
    name: Test Writer
    roleDefinition: >-
      You write tests against the documented contract of a
      module, never against its implementation details.
    groups:
      - read
      - - edit
        - fileRegex: \.(test|spec)\.(ts|js)$
      - command

A mode that can only edit test files is a genuinely useful safety property, and it is the sort of constraint that is much more reliable than asking the model nicely in a prompt.

Note that .roomodes has no field for an API configuration. Modes are commitable; the model bound to them is not.

Moving settings between machines

Settings has Export, Import and a red Reset. Export writes a JSON file containing your provider profiles and global settings. Import merges rather than replaces — new profiles are added, existing ones updated, nothing is deleted.

The exported file contains API keys in plain text. Treat it as a credential, not a config file: do not commit it, do not put it in shared storage, and delete it once imported. Reset is irreversible and wipes profiles, secret-storage keys, custom modes and task history.

A working order

  1. Verify your endpoint returns tool_calls for a request with tools. Nothing works otherwise.
  2. Configure one OpenAI Compatible profile with base URL, key and model.
  3. Set Context Window Size correctly; leave Max Output Tokens at -1 unless you have a reason.
  4. Clone it into Heavy, Standard and Fast profiles.
  5. Bind profiles to modes in the Prompts tab.
  6. Add a .roomodes file with any restricted modes your repository benefits from, and commit it.

Common questions

Does Roo Code work with a provider that lacks native tool calling?

No. Roo Code supports native tool calling only, with no XML fallback, so a model or gateway that does not implement OpenAI function calling properly cannot be used. Verify that a request with a tools array returns tool_calls before configuring anything else.

Can I commit my Roo Code model configuration to the repository?

Partly. Custom modes go in a .roomodes file at the repo root and are commitable, but the binding between a mode and an API configuration profile lives in global VS Code state and cannot be shared that way.

Why does a reasoning model return 400 on every request in Roo Code?

Try the Enable R1 model parameters toggle on the provider screen. It exists specifically to prevent 400 errors with R1-style reasoning models, and it is the first thing to check before suspecting the endpoint.

Similar articles

Cline Setup Guide: Pointing It at a Custom Endpoint
Guides
Guides·8 min read

Cline Setup Guide: Pointing It at a Custom Endpoint

Configure Cline against any OpenAI-compatible base URL — the provider fields, the model configuration block that people skip, and separate Plan and Act models.

Read
Continue.dev Setup: config.yaml for a Custom Endpoint
Guides
Guides·8 min read

Continue.dev Setup: config.yaml for a Custom Endpoint

Point Continue at any OpenAI-compatible base URL using config.yaml — model roles, apiBase, request options, and why the autocomplete slot needs its own model.

Read
Kilo Code Setup Guide: Custom Endpoints and Profiles
Guides
Guides·8 min read

Kilo Code Setup Guide: Custom Endpoints and Profiles

Configure Kilo Code against your own OpenAI-compatible endpoint, set the model metadata it cannot infer, and assign different models per mode with profiles.

Read