Handling SSE Disconnects Without Losing the Response
Streamed completions fail silently after a 200. How to detect a truncated SSE stream, set idle timeouts, survive proxy hops, and decide when a retry is safe.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Hands-on setup for the tools developers actually use.
Streamed completions fail silently after a 200. How to detect a truncated SSE stream, set idle timeouts, survive proxy hops, and decide when a retry is safe.
ReadHow to run a model of your choosing inside IntelliJ, PyCharm or GoLand — what the bundled assistant will and will not do, and the plugin route that does.
ReadWhich JSON Schema features survive the trip to a model, why strict mode changes the rules, and how to write tool parameter schemas that validate and get used correctly.
ReadConfigure Kilo Code against your own OpenAI-compatible endpoint, set the model metadata it cannot infer, and assign different models per mode with profiles.
ReadConfigure LlamaIndex against a non-OpenAI base URL, split the LLM from the embedding model, and avoid the defaults that silently call OpenAI anyway.
ReadMost classification failures are label definition failures. How to design a taxonomy, handle abstention and imbalance, and check the model beats a cheap baseline.
ReadExtraction failures are usually schema failures. How to model absent, ambiguous and multi-valued fields, attach provenance, and validate what comes back.
ReadSummaries fail by leaving things out, not by making things up. How to control what gets kept, pick a chunking strategy, and evaluate faithfulness cheaply.
ReadMachine translation is solved enough. What breaks in software localisation is interpolation syntax, terminology drift and missing context, not language quality.
ReadAdding a model to search usually means query rewriting and reranking, not generation. Where each stage helps, what it costs in latency, and how it fails.
ReadThe fields that make a model request debuggable are mostly not the fields that make it risky. How to log traffic that stays useful and minimal by default.
ReadWire a Neovim LLM plugin to your own OpenAI-compatible endpoint — how adapters are shaped, where the key should come from, and why streaming often looks broken.
ReadShowing 25–36 of 77 articles