OpenAI Node SDK Setup Against a Custom Base URL
Configure the official Node client for any OpenAI-compatible endpoint: retries, timeouts, abort signals, streaming, typed errors and the runtime traps in serverless.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Hands-on setup for the tools developers actually use.
Configure the official Node client for any OpenAI-compatible endpoint: retries, timeouts, abort signals, streaming, typed errors and the runtime traps in serverless.
ReadConfigure the official Python client against a custom base URL: retries, timeouts, async, streaming, tool calls and the error hierarchy you should be catching.
ReadConfigure opencode against your own OpenAI-compatible endpoint — the provider block, model naming, per-mode model choice, and how to verify tool calls work.
ReadPersonal data reaches model context through stack traces, fixtures and pasted tickets rather than through design. Detection, tokenisation and what to ask a vendor.
ReadHow to wire a model into a git pre-commit hook, which checks actually belong there, and the latency and cost budget that decides whether the team keeps it.
ReadPrompts drift, break silently and get edited in production. How to template them safely, version them properly and know which version produced which output.
ReadPutting a gateway in front of an inference provider centralises keys, budgets and routing. Here is what breaks if you get streaming, headers or cancellation wrong.
ReadCredentials reach model context through error output, config files and git history. Pre-send scanning, why gitignore is no defence, and rotating on suspicion.
ReadConsuming SSE from a chat completions endpoint in TypeScript: reading the body stream, buffering partial frames, assembling tool-call deltas and cancelling cleanly.
ReadConsume an OpenAI-compatible token stream in Python — the raw wire format, parsing data lines, partial chunks, tool-call deltas, timeouts and what buffering breaks.
ReadConnect, first-token and idle timeouts do different jobs. How to pick each from your own latency data, propagate deadlines, and avoid the default ten-minute wait.
ReadA model will produce a plausible schema in seconds. Plausible is the problem. How to brief it with queries, review the parts it gets wrong, and keep migrations safe.
ReadShowing 37–48 of 77 articles