Local Models vs API: Running the Numbers
Guides

Local Models vs API: Running the Numbers

Local inference is free at the margin and expensive everywhere else. Work through the memory arithmetic, the real break-even, and where each option genuinely wins.

Running a model locally has an obvious appeal: no per-token cost, no rate limits, no data leaving the machine, and it works on a plane. The pitch is strong enough that people frequently buy hardware before doing the arithmetic.

The arithmetic is worth doing, because local inference trades a variable cost you can see for several fixed costs you cannot. This is a comparison of what each option actually gives you, starting with the constraint that decides most of it: memory.

The memory arithmetic

Two things must fit in memory: the weights and the KV cache. The weights are simple.

weight_bytes = parameters × bytes_per_parameter

FP16    2 bytes/param
FP8     1 byte/param
INT4    ~0.5 bytes/param

Real 4-bit quantisation formats use mixed precision and land somewhat above the theoretical figure — commonly cited around 4.5 bits per weight for the popular k-quant variants — so treat 0.5 bytes as a floor rather than a promise.

The KV cache is what people forget. It stores the attention keys and values for every token in the context, and it grows linearly with context length:

kv_bytes ≈ 2 × layers × kv_heads × head_dim
              × context_length × bytes_per_element

The practical consequence: a model that fits at 4k context may not fit at 64k. This is the single most common surprise in local deployment. Advertised context windows are a property of the model; usable context is a property of your remaining memory after the weights.

Two mitigations are standard. Quantise the KV cache to 8 bits, which most current runtimes support and which roughly halves it. Or simply configure a shorter context, which is what most local setups end up doing.

What local inference actually costs

The marginal token is free. Everything else is not.

  • Hardware. Amortise the purchase over its useful life, not over the first month. A machine bought for this is a capital cost whether or not you enjoy owning it.
  • Electricity. Modest for intermittent use, non-trivial for a GPU held at load. Compute it from your own tariff and the card power draw.
  • Your time. Runtime updates, quantisation choices, model upgrades, driver problems. This is the largest hidden cost by a wide margin and it recurs.
  • Capability gap. The strongest model that fits on your hardware is not the strongest model available. For hard tasks that gap costs hours, which is real money at any professional rate.

Set against that, the honest comparison to metered API access:

break_even_months =
    (hardware_cost + setup_hours × your_rate)
  / (monthly_api_spend - monthly_electricity)

Run this with your real numbers. For light users the denominator is small enough that the break-even is measured in years, by which time the hardware is obsolete. For sustained heavy inference on a task a local model handles well, it can be genuinely short.

Where local genuinely wins

There are cases where this is not close, and they are mostly not about cost.

  • Data that cannot leave. Regulated environments, client code under strict agreements, air-gapped networks. No API arrangement fully substitutes for the data never being transmitted.
  • High-volume, narrow tasks. Classification, extraction, embedding generation, bulk summarisation. Small models handle these well and the volumes are where marginal cost adds up.
  • Latency-critical short completions. No network round trip. For autocomplete-style workloads, local can be perceptibly faster despite weaker hardware.
  • Offline and unreliable connectivity. Straightforward and decisive.
  • Learning how models behave. Being able to inspect logits, swap quantisations and break things without a bill is genuinely valuable.

Where the API wins

Equally, be honest about the other direction.

  • Frontier capability. The largest models are not runnable on any single-machine setup you are likely to own, and for hard reasoning the gap is not subtle.
  • Long context. A million-token window is a memory problem you cannot solve locally at reasonable cost. KV cache scaling makes this decisive.
  • Bursty load. Hardware sized for your peak sits idle the rest of the time. Elasticity is exactly what a hosted API sells.
  • Agentic coding. Long-horizon multi-step work is where per-step reliability compounds hardest, and where the capability gap between a local model and a frontier one is widest.
  • Not being the operator. Somebody else handles capacity, updates and uptime.

The hybrid setup most people should run

This is not a binary, and treating it as one is the actual mistake.

  1. Local small model for the mechanical band: explanations, boilerplate, rewrites, commit messages, anything on private code. Zero marginal cost, no data egress.
  2. Hosted frontier model for hard reasoning, long context and agent loops.
  3. An abstraction between them. Most local runtimes expose an OpenAI-compatible endpoint, so switching is a base URL and a model name. Keep both in configuration and the choice stays cheap.

Route by difficulty, not by ideology. The failure mode of pure local is spending an evening on a problem a better model solves in a minute; the failure mode of pure API is paying frontier rates to reformat a JSON blob.

A checklist before buying hardware

  1. Identify the largest model you actually need, and compute weights plus KV cache at your required context length.
  2. Add 15 to 20 percent overhead for the runtime.
  3. Test that exact model through a hosted provider first, on your real tasks. Confirm it is good enough before buying a machine to run it on.
  4. Compute the break-even against your current spend, including your own setup time at your real hourly rate.
  5. Decide what happens when a better model ships in six months and does not fit.

Point five is the one that catches people. Hardware is fixed; models are not. If your reason for going local is per-token cost on exploratory work rather than data residency, a flat-rate hosted arrangement solves the same problem with no capital outlay and no hardware to outgrow. If the reason is that the data genuinely cannot leave your network, none of the arithmetic above matters and you should run it locally.

Common questions

How much VRAM do I need to run a model locally?

Multiply parameters by bytes per parameter, then add the KV cache, then add 15 to 20 percent overhead. At 4-bit a 70B model needs roughly 40 GB for weights alone, and long context can add tens of gigabytes more.

Is running a local model cheaper than paying for an API?

Only at sustained high volume on tasks a local model handles well. Divide hardware cost plus your setup time by monthly API spend minus electricity to get a break-even in months. For light users it usually exceeds the useful life of the hardware.

Why does my local model run out of memory at long context?

The KV cache grows linearly with context length and is stored on top of the weights. A model that fits at 4k context can fail at 64k. Quantising the KV cache to 8 bits or reducing the configured context length are the usual fixes.

Similar articles

Aider Setup Guide: Any OpenAI-Compatible Endpoint
Guides
Guides·8 min read

Aider Setup Guide: Any OpenAI-Compatible Endpoint

Configure Aider against a custom base URL — the openai/ prefix, .aider.conf.yml, model metadata for unknown models, and picking the right edit format.

Read
Building a Changelog Generator People Actually Read
Guides
Guides·9 min read

Building a Changelog Generator People Actually Read

Restating commit subjects is not a changelog. How to pick the right input, separate classification from writing, handle reverts, and keep regeneration deterministic.

Read
Building a Chatbot From Scratch: The Parts Nobody Mentions
Guides
Guides·9 min read

Building a Chatbot From Scratch: The Parts Nobody Mentions

The model call is twenty lines. The other ninety percent is conversation state, idempotency, abuse limits and knowing when a reply was wrong. A build order that works.

Read