curl Recipes for LLM APIs: Debug Before You Write Code
A working set of curl commands for OpenAI-compatible endpoints: streaming, timing, tool calls, error bodies, and building JSON safely with jq.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
A working set of curl commands for OpenAI-compatible endpoints: streaming, timing, tool calls, error bodies, and building JSON safely with jq.
ReadHow to run Cursor against a custom OpenAI-compatible base URL, which features stop using your key, and how to tell whether the override took effect.
ReadHow benchmark data leaks into training sets, why it is hard to prove, and what a contaminated score actually costs you when you pick a model for real work.
ReadStuck is four different failures with four different fixes. How to tell hung from looping from stalled from quietly truncated, and what to do about each one.
ReadTwo budget models with 1M context. One is cheaper and text-only under MIT, the other sees images. The choice is almost entirely about input type.
ReadA 13B-active MoE at fourteen cents per million tokens against a dense 27B you can run yourself. The crossover point is lower than most teams assume.
ReadThe cheap half of the DeepSeek V4 family keeps the million-token window and drops active parameters to 13B. What that trade buys, and where it stops working.
ReadV4 Pro reports 80.6 percent on SWE-bench Verified, Qwen 3.6 27B reports 77.2. One is sixty times larger. What that tells you about model size in 2026.
ReadBoth ship 1M context and an MIT licence. Pro activates 49B parameters per token, Flash 13B. Where that single difference decides which one you should run.
ReadMoE models are cheap to compute and expensive to hold. Dense models are the reverse. The architecture decides your deployment more than your benchmark scores.
ReadMost agent plans are flat lists executed top to bottom. Modelling dependencies instead unlocks parallelism, better retries and honest progress reporting.
ReadMost LLM cost dashboards show token counts nobody acts on. The four or five panels worth keeping, why unit cost beats totals, and how to avoid vanity metrics.
ReadShowing 97–108 of 404 articles