Vision-Language Models: What to Test Before You Commit
Native image input is now common in open-weight models, but published multimodal benchmarks predict almost nothing. Three probes that separate them.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Picking between models — and testing them on your own work.
Native image input is now common in open-weight models, but published multimodal benchmarks predict almost nothing. Three probes that separate them.
ReadNew models ship constantly and switching is never free. The signals that justify a migration, the ones that do not, and how to run the decision.
ReadTotal parameters set the memory floor, active parameters set the speed. How to work out what your hardware can actually serve before reading a benchmark.
ReadTwo credible evaluations can rank the same models differently without either being wrong. How to read that disagreement instead of picking the flattering one.
ReadAn automated reviewer that flags everything gets muted within a week. How to pick and evaluate a model on false-positive rate, and why review has odd economics.
ReadDebugging is a search problem, so the model that wins is the one that forms testable hypotheses and abandons them fast. How to evaluate that properly.
ReadModels are very good at producing tests that pass and prove nothing. Use mutation score to pick one, and treat test writing as your cheapest high-volume task.
ReadA 1.6T MoE that activates 49B per token, ships under MIT, and costs $0.435 per million input tokens. Where V4 Pro is strong, where it is not, and Pro versus Flash.
ReadZ.ai kept the architecture and the price list identical and changed what the model is for. The context jump, the long-horizon gains, and the token cost nobody mentions.
ReadBoth are MIT-licensed open models, so the licence is a wash. The split is agentic endurance against algorithmic power, at a three-to-five times price gap.
ReadZ.ai shipped a 744B MoE with 40B active, MIT-licensed weights and the first open-weight Terminal-Bench 2.1 score above 80. A technical read on what that means.
ReadSix properties decide whether a model is good at your codebase — and only one of them shows up on a leaderboard. A practical framework for picking.
ReadShowing 97–108 of 119 articles