JSON Schema for Tool Calls: What Providers Actually Support
Which JSON Schema features survive the trip to a model, why strict mode changes the rules, and how to write tool parameter schemas that validate and get used correctly.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Which JSON Schema features survive the trip to a model, why strict mode changes the rules, and how to write tool parameter schemas that validate and get used correctly.
ReadConfigure Kilo Code against your own OpenAI-compatible endpoint, set the model metadata it cannot infer, and assign different models per mode with profiles.
ReadMoonshot shipped K3 four months after K2.6, but the older model did not become useless. What K2.6 is, what it costs, and the jobs it is still the right pick for.
ReadTwo models released three days apart with opposite design goals. Active parameters, context length, licensing and vision decide which one fits your workload.
ReadK2.6 sees images at 256K context under a custom licence. V4 Pro is text-only, MIT, 1M context, and reasons harder. Two models that barely overlap.
ReadK2.6 sees images and stops at 256K context under a custom licence. GLM-5.2 is text-only, MIT, and takes 1M tokens. Two clean trade-offs, no overlap.
ReadBoth take images natively, but one gives you 256K of context at a known price and the other 1M at a price sources disagree on. How to pick between them.
ReadA trillion-parameter vision model you rent against a dense 27B you can own. The comparison is about deployment shape, not a few points of benchmark difference.
ReadThe strongest open-weight model against one of the cheapest, both with 1M context. When a twenty-fold price difference is worth paying and when it is waste.
ReadK3 is the stronger model. K2.6 costs roughly a quarter as much per output token and activates a third of the parameters. When the older model is still the right call.
ReadOne is the strongest open-weight model available. The other is among the cheapest capable ones and natively multimodal. The gap between them is a budget decision.
ReadK3 needs roughly 1.6TB of weights and multi-node serving. Qwen 3.6 27B is dense and fits on one accelerator. The comparison is about deployment, not benchmarks.
ReadShowing 145–156 of 404 articles