Generating Tests With LLMs: Coverage Is Not the Goal
Models are very good at writing tests that pass and prove nothing. How to get suites that actually catch regressions, and how mutation testing tells you which ones do.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Models are very good at writing tests that pass and prove nothing. How to get suites that actually catch regressions, and how mutation testing tells you which ones do.
ReadZ.ai kept the architecture and the price list identical and changed what the model is for. The context jump, the long-horizon gains, and the token cost nobody mentions.
ReadBoth are MIT-licensed open models, so the licence is a wash. The split is agentic endurance against algorithmic power, at a three-to-five times price gap.
ReadZ.ai shipped a 744B MoE with 40B active, MIT-licensed weights and the first open-weight Terminal-Bench 2.1 score above 80. A technical read on what that means.
ReadA large language model does one thing repeatedly. Here is what happens between your prompt and the first token out, and what that mechanism explains about behaviour.
ReadSix properties decide whether a model is good at your codebase — and only one of them shows up on a leaderboard. A practical framework for picking.
ReadMost approval gates fail not from too little oversight but from too much. Here is how to design agent checkpoints people actually read before clicking yes.
ReadOutput tokens cost several times more per token, yet input usually dominates the bill. Here is how to compute your own blended rate and act on it.
ReadNo inference provider can sell genuinely uncapped compute. Here is what the word hides, why the limits exist, and exactly what to demand in writing before you buy.
ReadK3 costs about seventeen times more per output token than DeepSeek V4 Pro. Here is the break-even calculation that decides whether that is worth paying.
ReadOne is the most capable open-weight model and hard to serve. The other is MIT-licensed, leaner and built for long agent loops. How to choose between them.
ReadMoonshot shipped the largest open-weight model yet: 2.8T parameters, 1M context, and a licence that is not MIT. A technical read on where K3 earns its price.
ReadShowing 337–348 of 404 articles