Building an Incident Assistant On-Call Engineers Trust
Automate the first ten minutes of context gathering, not the diagnosis. Read-only tools, evidence-linked output, a latency budget, and why a wrong answer at 3am is expensive.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Automate the first ten minutes of context gathering, not the diagnosis. Read-only tools, evidence-linked output, a latency budget, and why a wrong answer at 3am is expensive.
ReadWiring a model into pull request review: what to send it, how to post comments that land, and the signal-to-noise threshold that decides whether the bot survives.
ReadA practical walkthrough of writing your first MCP server: choosing the tools, wiring the transport, handling errors, and testing it with a real agent.
ReadBPE is a compression algorithm that became the standard way to split text for language models. How it is trained, what it produces, and why it behaves oddly.
ReadA declarative pipeline stage that calls an OpenAI-compatible endpoint, handles credentials properly, fails loudly on errors and does not bankrupt a shared build server.
ReadAsking a model to reason step by step measurably improves accuracy on some tasks and wastes tokens on others. The mechanism, and when it is worth the cost.
ReadAllocating LLM spend back to the teams that caused it: tagging, per-key versus per-service attribution, shared costs, and when showback is the better step.
ReadComments, identifiers and issues in another language change tokenisation, cost and accuracy. What to test and how to pick a model that handles it.
ReadHow to build a breaker around an inference provider: what to count as a failure, where to set thresholds, how half-open probes work, and when to fall back instead of failing.
ReadPer-token, per-seat, credits, tiers and flat rate all quote different units. Here is how to normalise them onto one number you can actually compare.
ReadLong-lived streaming requests break the assumptions behind default HTTP pools. How to size keepalive connections, avoid pool starvation, and stop creating a client per call.
ReadInstead of paying humans to rank thousands of responses, write the principles down and have the model apply them. How the method works and where it strains.
ReadShowing 73–84 of 404 articles