A/B Testing Two Models Without Fooling Yourself
Comparing two models on live traffic sounds simple and usually is not. Sample sizes, paired designs, and the metrics that actually settle the question.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Picking between models — and testing them on your own work.
Comparing two models on live traffic sounds simple and usually is not. Sample sizes, paired designs, and the metrics that actually settle the question.
ReadA 2.8T model and a 27B model can be two-to-one apart on the figure that governs thinking. How to read parameter counts across the 2026 field.
ReadChat benchmarks say little about a model driven in a loop for forty turns. What agentic performance actually measures, and how the 2026 field ranks on it.
ReadArena ratings come from blind pairwise votes, not from tests. What the number means, why gaps under fifty points are noise, and where it misleads.
ReadContamination inflates scores when test data leaks into training. How it happens, why it is over-diagnosed, and the practical defences that actually work.
ReadAn agent that is 95 percent reliable per step fails most thirty-step tasks. Why compounding error, not peak capability, decides which model to run in a loop.
ReadBackend work is the one domain where SWE-bench scores roughly mean what you want them to mean — and the places where they still mislead you.
ReadWhen nobody is waiting for the answer, latency stops mattering and unit cost dominates. How to pick and operate a model for offline high-volume work.
ReadThe code runs, the chart renders, and the conclusion is unsound. Why data science needs a different evaluation than general code generation.
ReadInfrastructure code fails differently: rarely, and expensively. Why terminal benchmarks are the relevant signal and what to gate before applying.
ReadDocumentation is the task where fluent output is most dangerous, because readers cannot verify it. How to pick and constrain a model for docs that stay true.
ReadAt enterprise scale the constraints are licensing, data residency, auditability and version stability. How those narrow the field before capability matters.
ReadShowing 1–12 of 119 articles