Transformer Architecture Explained for Working Developers
What a transformer block actually contains, why the design won, and which parts of it explain the behaviour you observe when using a model.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
How large language models actually work, in plain terms.
What a transformer block actually contains, why the design won, and which parts of it explain the behaviour you observe when using a model.
ReadEmoji that cost five tokens, accented text that costs double, and truncation that produces broken characters. How Unicode meets the tokenizer, and what breaks.
ReadA bigger vocabulary shortens sequences and costs parameters and rare-token quality. What the dial actually controls and how it reaches your bill.
ReadQuantisation degrades capabilities unevenly rather than uniformly. Which ones erode first, why a benchmark delta hides it, and how to measure the loss on your own work.
ReadA million-token context window sounds like the end of chunking. In practice models get slower, pricier and less accurate long before you fill it. Here is why.
ReadEmbeddings turn text into vectors so you can search by meaning instead of keywords. Here is how they work, what similarity really measures, and the failure modes to expect.
ReadThree ways to make a general model do your specific job, with very different costs. Here is a decision rule based on what each one can and cannot actually change.
ReadA large language model does one thing repeatedly. Here is what happens between your prompt and the first token out, and what that mechanism explains about behaviour.
ReadLLM latency is two different problems wearing one name. Understanding prefill and decode tells you which knob to turn when a request feels slow.
ReadMoE models decouple parameter count from compute per token, which is why a trillion-parameter model can be cheap to serve. Here is the mechanism and what it costs you.
ReadHow a small model inherits the behaviour of a much larger one, what gets lost along the way, and why distillation is the reason cheap models got good so fast.
ReadA multimodal model does not see an image. It sees patch embeddings projected into the same space as text tokens. What that implies for cost and accuracy.
ReadShowing 49–60 of 72 articles