Best Model for Solo Developers: One Model, Every Job
Working alone means no routing layer and no ops budget. Why breadth and predictable cost beat peak capability when one model has to do everything.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Picking between models — and testing them on your own work.
Working alone means no routing layer and no ops budget. Why breadth and predictable cost beat peak capability when one model has to do everything.
ReadA wrong query returns rows rather than an error, which is why SQL generation fails quietly. What to feed the model and how to verify before trusting.
ReadEarly-stage teams change their mind quarterly. Why the model decision that matters is how cheaply you can replace it, not which one wins today.
ReadModel choice matters less than constrained decoding for reliable JSON. What actually guarantees valid output, and where model quality still decides.
ReadA smarter model that emits malformed calls is worse than a weaker one that never does. What to test before choosing a model for an agent loop.
ReadTypeScript gives you a free verifier, which changes which model you should pay for. How to build the tsc loop and where models still fall down.
ReadVision in a developer workflow means screenshots, diagrams and scanned documents. Which models take image input, and where visual understanding still fails.
ReadKnowing which model wrote an answer changes how you judge it. A practical protocol for blind pairwise comparison that a small team can run in a day.
ReadA practical design for a local evaluation runner: fixture format, isolation, scoring, cost tracking and the CI wiring that keeps it from rotting.
ReadComments, identifiers and issues in another language change tokenisation, cost and accuracy. What to test and how to pick a model that handles it.
ReadFilling a million-token window costs between fourteen cents and three dollars depending on the model. The full comparison, including output ceilings.
ReadCapability leaderboards ignore price entirely. How to build a score that reflects what a model costs to run on your workload, and where it misleads.
ReadShowing 25–36 of 119 articles