Rate Limits and Retries: Backoff That Does Not Make It Worse
Token buckets, jitter, retry budgets and circuit breakers for LLM APIs — how to stay under the limit instead of discovering it, and why naive retries amplify outages.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Hands-on setup for the tools developers actually use.
Token buckets, jitter, retry budgets and circuit breakers for LLM APIs — how to stay under the limit instead of discovering it, and why naive retries amplify outages.
ReadConfigure Roo Code against an OpenAI-compatible base URL, then use API configuration profiles to give each mode its own model. Includes the tool-calling caveat.
ReadPutting a model call inside CI adds a non-deterministic network dependency to your build. How to keep it cheap, secret-safe on forks, and incapable of blocking a merge.
ReadStreaming is what makes an AI feature feel fast. It is also where proxies, buffers and framework defaults quietly break things. A practical guide.
ReadMost AI tools speak one API shape. Understanding what compatibility covers — and the four places it usually leaks — saves hours of debugging a base URL swap.
ReadShowing 73–77 of 77 articles