Open Weights vs Open Source: The Distinction That Matters
Downloadable weights are not the same thing as open source. Here is what the licences actually permit, and which differences change what you can ship.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
How large language models actually work, in plain terms.
Downloadable weights are not the same thing as open source. Here is what the licences actually permit, and which differences change what you can ship.
ReadPrompt caching reuses the prefill work for a repeated prefix, cutting cost and first-token latency. The whole technique comes down to one ordering rule.
ReadQuantization shrinks model weights to fewer bits so they fit in less memory. Here is how the formats differ, what accuracy you lose, and how to pick one without guessing.
ReadLarge context windows were supposed to kill retrieval. They did not. Here is how the two actually compare on cost, accuracy and latency — and when to use each.
ReadReasoning models spend extra tokens working before they answer. Here is what that changes mechanically, which tasks it helps, and where it is just an expensive delay.
ReadSpeculative decoding drafts several tokens cheaply and verifies them in one pass, cutting latency without changing the output distribution. How it works.
ReadA system prompt is not a control panel and not a security boundary. Here is what it really is, where it sits in the instruction hierarchy, and how to write one.
ReadTemperature is not a creativity dial and top-p is not a quality setting. Here is what each sampling parameter changes mathematically, and how to set them for real work.
ReadTokenizers decide what your prompt costs, how long your context really is, and why models miscount letters. Here is how subword tokenization works in practice.
ReadParameter counts are the most quoted and least understood model spec. What they measure, why total and active differ, and when the number predicts anything useful.
ReadHallucination is not a bug that will be patched out. It follows from the training objective and from how we grade models. Here is the mechanism and the mitigations that work.
ReadTemperature zero is not determinism, and a seed is only best effort. Here is where LLM nondeterminism really comes from and how to build around it.
ReadShowing 61–72 of 72 articles