Sparse vs Dense Activation: What Runs on Every Token
Sparse activation means most of a model sits idle for any given token. That single fact explains model pricing, memory bills and misleading spec sheets.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
How large language models actually work, in plain terms.
Sparse activation means most of a model sits idle for any given token. That single fact explains model pricing, memory bills and misleading spec sheets.
ReadSpeculative decoding is standard in serving stacks now, but the speedup is workload-dependent. What decides whether it helps you, and how to tell.
ReadHow stop sequences work, why they interact badly with tokenisation and streaming, and the finish_reason check most integrations forget to make.
ReadHow generated data is used to train modern models, what distillation and self-instruct pipelines look like, and what the model-collapse evidence actually supports.
ReadSpending more computation at inference can substitute for a larger model. How the trade works, where it pays off, and what it does to your latency budget.
ReadModels retrieve facts from the start and end of a long prompt far more reliably than from the middle. Why that happens and how to arrange prompts around it.
ReadServing more tokens per second and serving them faster are opposing goals. The knob that reconciles them is batch size, and it decides both speed and price.
ReadTTFT is dominated by prefill compute, queueing and network distance rather than by model speed. What each contributes, and how to measure it without fooling yourself.
ReadThe same sentence costs more in some languages than in English, and code has its own profile. Where the multiplier comes from and how to measure yours.
ReadTwo models quoting the same price per million tokens can cost different amounts for identical text. Why tokenizers vary and what it does to your bill.
ReadTokens per second means different things depending on who is measuring. What drives generation speed, why quoted figures disagree, and how to measure yours.
ReadTop-k cuts the tail at a fixed count, which is right sometimes and wrong often. How nucleus sampling adapts, and which knob to actually touch.
ReadShowing 37–48 of 72 articles