Cline Setup Guide: Pointing It at a Custom Endpoint
Guides

Cline Setup Guide: Pointing It at a Custom Endpoint

Configure Cline against any OpenAI-compatible base URL — the provider fields, the model configuration block that people skip, and separate Plan and Act models.

Cline is an autonomous coding agent that runs inside VS Code. It reads files, writes edits, runs terminal commands and iterates, asking for approval at whatever granularity you configure. Pointing it at a custom OpenAI-compatible endpoint takes about a minute, and the part that goes wrong is not the part people expect.

The base URL and key are trivial. The section that determines whether the agent behaves is the model configuration block underneath them, and it is easy to scroll past.

Connecting to the endpoint

Open the Cline panel and click the gear icon to reach settings. Then:

  1. Set API Provider to OpenAI Compatible.
  2. Set Base URL to https://api.cozyapi.com/v1. The documentation calls this a crucial step, and it is the field most often filled in wrongly — it should end at /v1 once, with no /chat/completions appended.
  3. Set API Key to your key, for example cozy_your_key_here.
  4. Set Model — referred to as the Model ID in the documentation — to an alias your endpoint recognises, such as kimi-k3, glm-5.2 or qwen-3.5-coder. This string is passed through verbatim, so a typo produces a 404 rather than a fallback.
  5. Click Verify to confirm the connection before doing anything else.

There is also a Use Azure Identity Authentication checkbox on this screen. Leave it off unless you are actually on Azure; it reuses an existing identity rather than triggering a new sign-in, so enabling it by accident produces confusing auth failures.

The model configuration block matters more than you think

Below the credentials is a Model Configuration section: Max Output Tokens, Context Window size, Image Support capabilities, Computer Use, Input Price and Output Price.

These exist because Cline has no way to know anything about a model behind a custom endpoint. For built-in providers it ships metadata; for an arbitrary alias it has to be told. Three of these fields change behaviour rather than just display:

  • Context Window size drives when Cline compacts or truncates conversation history. Set it too low and long sessions get summarised aggressively for no reason, losing detail the agent needed. Set it too high and you hit upstream errors mid-task. Use the real window for the model you selected — Kimi K3, GLM 5.2 and DeepSeek V4 Pro all publish a one-million-token context, but check the alias you are actually calling rather than assuming.
  • Max Output Tokens caps a single response. Too low truncates a large file write halfway, which surfaces as a mangled edit rather than an error.
  • Computer Use is described in the documentation as applying to models with tool and function calling. Cline is an agent: it lives on tool calls. If this is off for a model that supports them, you get a chatbot that describes changes instead of making them.

The price fields only feed the in-panel cost readout. They are worth filling in if you are on per-token billing and want the number to mean something, and safe to leave alone otherwise.

Different models for Plan and Act

Cline separates planning from execution. In Plan mode it reads the codebase and proposes an approach without editing; in Act mode it carries the plan out. The split exists because those two jobs want different models.

In settings there is a toggle labelled Use different models for Plan and Act. Enable it and you can select a model per mode, with the switch happening automatically when you change modes.

The useful configuration is a strong reasoning model for Plan and a fast, cheap one for Act. Planning is where mistakes are expensive and cheap to fix — a bad plan costs you a paragraph of reading, a bad execution costs you a broken branch. Execution, once the plan is agreed, is largely mechanical file editing.

A reasonable pairing to start from: kimi-k3 or glm-5.2 on Plan, something faster such as qwen-3.5-coder or minimax-m2.7 on Act. Then measure whether the Act model is producing edits you have to correct; if it is, move it up a tier rather than moving the Plan model down.

Rules are where the real configuration lives

Model choice sets the ceiling. Rules determine whether you get there.

Cline reads project rules from a .clinerules/ directory in your workspace, where each .md or .txt file contributes to a combined rule set. It also auto-detects AGENTS.md, and picks up .cursorrules and .windsurfrules if you have migrated from another tool. Global rules live under ~/Documents/Cline/Rules. Where they conflict, workspace rules win.

YAML frontmatter in a rule file lets you scope it, so a rule about database migrations can apply only when the agent touches migration files rather than being present in every request. The toggle UI for enabling and disabling rule files is behind the scale icon at the bottom of the panel.

What belongs in rules: your package manager, your test command, your error-handling convention, the directories that are generated and must not be edited by hand, and an explicit instruction about how much to change at once. Agents over-reach by default, and a rule that says "make the smallest change that satisfies the request, and ask before refactoring adjacent code" pays for itself within a day.

Approval settings deserve a deliberate decision

Cline can auto-approve reads, edits and terminal commands independently. The temptation after a good session is to approve everything, which is exactly when it deletes something.

A defensible baseline: auto-approve file reads, because they are harmless and constant. Review edits until you trust the model on this codebase. Never auto-approve arbitrary terminal commands on a machine with credentials, and if you do relax that, work inside a container or a sandbox rather than on your host.

Work on a branch and commit frequently. Cline moves fast, and git is the only undo that reliably works across a multi-file agent session.

When it does not connect

  • Verify fails with 404 — a doubled /v1, or an unknown model alias. Curl the endpoint directly to tell the two apart.
  • 401 — the key belongs to a different provider, or has a trailing space from the paste. Retype the last character.
  • Agent explains instead of editing — tool calling is not being negotiated. Check that Computer Use is enabled for the model.
  • Edits truncate mid-file — Max Output Tokens is too low for the files you are asking it to write.
  • History compacts constantly — Context Window size is set below the model's real window.

Cline stores its configuration under a ~/.cline directory alongside project-level .cline state, so a broken setup can be inspected on disk rather than only through the panel. The provider settings themselves are simplest to change through the UI.

The takeaway

Base URL, key and model take a minute. Spend the next ten on the model configuration block, a Plan and Act split, and a .clinerules file that encodes how your team works. That is where the difference between a useful agent and an expensive one is decided.

Common questions

Why does Cline describe changes instead of making them?

Tool calling is not being used. For a custom OpenAI-compatible provider, check that Computer Use is enabled in the model configuration block — Cline cannot infer capability for an alias it does not recognise, and without tool calls it degrades to chat.

Do I have to set the context window manually?

For a custom endpoint, yes. Cline ships metadata only for built-in providers, so an arbitrary alias needs the window set explicitly. Too low causes constant history compaction; too high causes upstream errors mid-task.

Is it worth using different models for Plan and Act?

Often. Planning rewards reasoning quality and is cheap to correct; execution is largely mechanical editing. A strong model on Plan and a fast one on Act cuts usage noticeably without much quality cost — measure by how often you have to fix the Act output.

Similar articles

Continue.dev Setup: config.yaml for a Custom Endpoint
Guides
Guides·8 min read

Continue.dev Setup: config.yaml for a Custom Endpoint

Point Continue at any OpenAI-compatible base URL using config.yaml — model roles, apiBase, request options, and why the autocomplete slot needs its own model.

Read
Kilo Code Setup Guide: Custom Endpoints and Profiles
Guides
Guides·8 min read

Kilo Code Setup Guide: Custom Endpoints and Profiles

Configure Kilo Code against your own OpenAI-compatible endpoint, set the model metadata it cannot infer, and assign different models per mode with profiles.

Read
Roo Code Setup: Custom Endpoint and Profiles Per Mode
Guides
Guides·8 min read

Roo Code Setup: Custom Endpoint and Profiles Per Mode

Configure Roo Code against an OpenAI-compatible base URL, then use API configuration profiles to give each mode its own model. Includes the tool-calling caveat.

Read