Open-Weight vs Closed Models: What the Gap Costs You Now
Closed models still lead on aggregate benchmarks, but the gap narrowed sharply through 2026. What open weights buy, what they still cost, and how to decide.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Closed models still lead on aggregate benchmarks, but the gap narrowed sharply through 2026. What open weights buy, what they still cost, and how to decide.
ReadConfigure the official Node client for any OpenAI-compatible endpoint: retries, timeouts, abort signals, streaming, typed errors and the runtime traps in serverless.
ReadConfigure the official Python client against a custom base URL: retries, timeouts, async, streaming, tool calls and the error hierarchy you should be catching.
ReadConfigure opencode against your own OpenAI-compatible endpoint — the provider block, model naming, per-mode model choice, and how to verify tool calls work.
ReadThe orchestrator-worker shape is the one multi-agent design that reliably pays for itself. Here is how to build the lead, the brief and the synthesis step.
ReadPaged attention stores the KV cache in fixed-size blocks instead of one contiguous slab, which is what lets a server hold far more concurrent sessions.
ReadRunning tool calls concurrently cuts wall-clock time but not tokens, and it is only safe for some operations. How to decide what to parallelise.
ReadSix workers, four returned, two timed out. Whether to answer, retry or abort is a design decision — here is how to make it before the incident, not during.
ReadPer-seat billing meters people, per-token billing meters work. Here is what each optimises for, which workloads each punishes, and how to model both.
ReadPersonal data reaches model context through stack traces, fixtures and pasted tickets rather than through design. Detection, tokenisation and what to ask a vendor.
ReadA stable model alias can point at different weights over time. How that breaks tuned prompts without any error, and what to pin instead.
ReadHow to size disk for open-weight models: what drives checkpoint size, why quantisations and versions multiply it, and why storage is rarely the real constraint.
ReadShowing 193–204 of 404 articles