Best Model for Frontend Code: Judgement Without a Compiler
Frontend work has no oracle that tells you the output is wrong. What that means for model choice, and how to close the loop with screenshots.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Picking between models — and testing them on your own work.
Frontend work has no oracle that tells you the output is wrong. What that means for model choice, and how to close the loop with screenshots.
ReadGo is the language models write most confidently, and confidence is the problem. Where generated Go compiles cleanly and still races under load.
ReadA million-token window does not make a model good at a million-line codebase. What actually determines whether a model can work in a large repository.
ReadLegacy work is mostly reading, not writing. Why the model that scores highest on code generation is often the wrong one for a twenty-year-old codebase.
ReadLong context is three different problems wearing one name. Which one you have decides whether window size, attention quality or cache pricing is the real constraint.
ReadFor interactive work, perceived speed is decided by time to first token, not throughput. Which models and settings actually make an interface feel fast.
ReadFramework and language migrations are hundreds of near-identical edits. The model property that decides success is consistency across files, not peak reasoning.
ReadMobile work punishes models differently to backend work. Slow builds, churning platform SDKs and visual output change which model is actually worth running.
ReadAn automated PR reviewer lives or dies on signal-to-noise, not raw capability. What the workflow demands of a model, and how to keep the bot from being muted.
ReadEvery model writes decent Python, which is exactly why choosing one is hard. The real differences show up in library recency and runtime failure.
ReadRust punishes weak models loudly rather than quietly. Why iterations-to-green is the metric that matters and where models predictably fail.
ReadSelf-hosting turns model selection into a memory problem. Which open-weight models fit on real hardware, and what you give up at each tier.
ReadShowing 13–24 of 119 articles