Flat-rate billing
One payment covers the whole pass. Prompts are unlimited, so a long agent run costs exactly what a one-line question costs — nothing extra. No token counting, no overage line, no bill shock at month end.
ProjectCOZY is a managed gateway that puts Kimi, Claude, DeepSeek, GLM, Qwen and MiniMax behind a single OpenAI-compatible endpoint — billed at a flat rate instead of per token.
One payment covers the whole pass. Prompts are unlimited, so a long agent run costs exactly what a one-line question costs — nothing extra. No token counting, no overage line, no bill shock at month end.
Kimi K3, Kimi K2.6, Claude 5 Opus, Claude 5 Sonnet, DeepSeek V4 Pro, GLM 5.2, GLM 5.1, MiniMax M3, MiniMax M2.7 and Qwen-3.5-Coder — all on the same key, all at the same price.
A drop-in chat-completions endpoint. Change the base URL and the model alias and your existing code, SDKs and tooling keep working — including streaming and tool calling.
A single cozy_ key drives your editor, your terminal agent and your CI job at once. Revoke and regenerate it from the dashboard any time without touching billing.
Server-sent events pass straight through, so token-by-token output arrives exactly as your client expects. Function and tool calling are supported on the models that implement them.
Request counts, token totals and per-model daily trends in the dashboard. You get the visibility of a metered API without being billed like one.
Only a SHA-256 hash of your key is stored. The raw value is shown once at generation and never retrievable — if it leaks, revoke and reissue in a click.
Pay with Binance Pay in any coin — USDT, BTC, ETH, BNB or SOL — plus KHQR where enabled. No card, no billing address, no per-seat minimum.
Your key is live the moment payment clears. No approval queue, no sales call, no waiting list — register and you are making requests in under a minute.
Every pass includes every model. Switching is a one-line change to the model alias in your client — no separate key, no tier upgrade, no renegotiation.
Long-context agentic coding and deep multi-step tool use.
Fast general-purpose coding with a large context window.
Hardest reasoning, refactors and architecture work.
Balanced speed and quality for day-to-day development.

Strong reasoning and maths at high throughput.

Reliable instruction following and structured output.

Long-horizon agent runs and high-volume batch work.

Code completion, test generation and translation.
Anything that speaks the OpenAI chat-completions format works unmodified. Set the base URL, paste your cozy_ key, pick an alias.
OpenCode
Roo CodeFull setup instructions for each client live in the documentation.
Sign up with email or Telegram. Every new account starts on the free trial — no card required.
One click in the dashboard issues a cozy_ key. It is displayed once, so copy it straight into your tool.
Set the base URL to our OpenAI-compatible endpoint and paste the key. No other code changes are needed.
Name the model you want in the request. Switch between Kimi, Claude, DeepSeek, GLM, Qwen and MiniMax whenever you like — same key, same price.
No card required. See the pass lengths and prices on the pricing page, or start with the trial and test the models on your own codebase first.