Tool Call Retries: Retry the Transport, Not the Judgement
Agents retry constantly, and most of it is wasted. How to tell a transport failure from a wrong decision, and what each one actually needs.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Agents retry constantly, and most of it is wasted. How to tell a transport failure from a wrong decision, and what each one actually needs.
ReadTool output is the largest block of prompt content you never wrote. How to shape it so agents stay accurate, cheap and able to decide what to do next.
ReadTool definitions are prompt content the model reads on every turn. Design choices there change agent reliability more than model selection does.
ReadTop-k cuts the tail at a fixed count, which is right sometimes and wrong often. How nucleus sampling adapts, and which knob to actually touch.
ReadA flat log of model calls cannot explain a twenty-step run. How to shape spans, propagate context through subagents and async tools, and replay a run from its trace.
ReadWhat a transformer block actually contains, why the design won, and which parts of it explain the behaviour you observe when using a model.
ReadWhy a flat-rate AI plan needs a fair-use boundary at all, what abuse looks like in the usage data, and how a heavy user stays clearly on the right side of it.
ReadEmoji that cost five tokens, accented text that costs double, and truncation that produces broken characters. How Unicode meets the tokenizer, and what breaks.
ReadHard caps stop runaway spend and stop your pipeline with it. How to design caps, warning paths and overage pricing that protect budget without outages.
ReadA model will produce a plausible schema in seconds. Plausible is the problem. How to brief it with queries, review the parts it gets wrong, and keep migrations safe.
ReadConfig generation is one of the safest LLM tasks if you supply the schema and validate mechanically. How to set it up, and the plausible-key failure mode to guard against.
ReadPoint LangChain at a non-OpenAI base URL, keep tool calling and streaming working, and know which abstractions quietly assume the official API.
ReadShowing 277–288 of 404 articles