Vendor Lock-In Costs: Measure It Instead of Fearing It
The HTTP call is trivially portable. Prompt tuning, tool-call formats, evaluation baselines and caching semantics are not. How to measure and contain the expensive parts.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
The HTTP call is trivially portable. Prompt tuning, tool-call formats, evaluation baselines and caching semantics are not. How to measure and contain the expensive parts.
ReadA launch-post score and a third-party score measure different things. Where the gap comes from, and how to normalise figures before you compare them.
ReadPoint the AI SDK at your own base URL, stream to a React client, keep tool calls working, and avoid the edge runtime traps that break streaming in production.
ReadThe most reliable agent improvement is refusing to accept completion until something verifies it. How to build a verifier that is worth trusting.
ReadNative image input is now common in open-weight models, but published multimodal benchmarks predict almost nothing. Three probes that separate them.
ReadA bigger vocabulary shortens sequences and costs parameters and rare-token quality. What the dial actually controls and how it reaches your bill.
ReadA provider-agnostic guide to wiring any VS Code AI extension to your own endpoint — where settings live, what the extension cannot infer, and how to debug it.
ReadA newer model at a lower rate can still raise your monthly bill. Regression testing, prompt drift, output length and tool-call changes are where the cost actually lands.
ReadNew models ship constantly and switching is never free. The signals that justify a migration, the ones that do not, and how to run the decision.
ReadQuantisation degrades capabilities unevenly rather than uniformly. Which ones erode first, why a benchmark delta hides it, and how to measure the loss on your own work.
ReadTotal parameters set the memory floor, active parameters set the speed. How to work out what your hardware can actually serve before reading a benchmark.
ReadTwo credible evaluations can rank the same models differently without either being wrong. How to read that disagreement instead of picking the flattering one.
ReadShowing 289–300 of 404 articles