GLM-5.2 vs MiniMax M3: Long-Horizon Code or Cheap Multimodal
Both ship 1M context and target coding. GLM-5.2 is built for unattended agent runs; M3 is multimodal and cheaper. Which one your workload actually needs.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Both ship 1M context and target coding. GLM-5.2 is built for unattended agent runs; M3 is multimodal and cheaper. Which one your workload actually needs.
ReadBoth are permissively licensed and strong at code. One is a 744B mixture-of-experts, the other a dense 27B on a single GPU. The architecture is the decision.
ReadConfigure Goose against a custom OpenAI-compatible endpoint, understand what its extensions cost you per request, and set approvals you can live with.
ReadWeights are only the first line of the memory budget. Where the rest goes, why concurrency runs out before compute does, and how to size a deployment.
ReadA rented accelerator bills by the hour whether or not you use it. How to work out effective cost per token from your own traffic shape.
ReadGQA shrinks the KV cache by sharing keys and values across query heads. What it costs in quality, and why it is on the spec sheet of nearly every model.
ReadStreamed completions fail silently after a 200. How to detect a truncated SSE stream, set idle timeouts, survive proxy hops, and decide when a retry is safe.
ReadModel cards mix hard facts, marketing and careful omissions. Which fields are reliable, which need checking, and what an absent section tells you.
ReadAgents retry, resume and duplicate calls constantly. Idempotency keys, natural keys and check-then-act patterns that stop one action happening three times.
ReadTraining is a one-off capital cost you never pay. Inference is a recurring cost you pay per request. How the two differ and why the distinction shapes pricing.
ReadA pretrained model continues text; it does not answer questions. Instruction tuning is the small, cheap stage that turns one into the other.
ReadHow to run a model of your choosing inside IntelliJ, PyCharm or GoLand — what the bundled assistant will and will not do, and the plugin route that does.
ReadShowing 133–144 of 404 articles