Building a Chatbot From Scratch: The Parts Nobody Mentions
The model call is twenty lines. The other ninety percent is conversation state, idempotency, abuse limits and knowing when a reply was wrong. A build order that works.
A working chatbot is about twenty lines of code. A chatbot you can run for real users, with more than one conversation at a time, that does not bankrupt you or silently degrade, is a normal web application with an unusual dependency.
This is a build order for the second thing. It assumes you already know how to send a chat completion request and read the response.
The server should be stateless; the conversation should not be
The single most consequential early decision is where conversation history lives. The API is stateless — every request carries the full message list — so somebody has to keep it.
Keeping it in the browser is tempting and wrong for anything real. You lose the conversation on refresh, you cannot moderate or audit it, and you have handed the client the ability to rewrite its own history, including the system prompt.
Keep messages in your database, keyed by conversation. The client sends a conversation ID and the new user message; the server loads history, appends, calls the model, appends the reply, and returns it. Your API surface ends up looking like this:
POST /conversations -> { id }
POST /conversations/:id/messages -> { message }
GET /conversations/:id/messages -> { messages[] }
A schema this size is enough to start:
conversations(id, user_id, title, created_at, updated_at)
messages(id, conversation_id, role, content, tool_calls,
prompt_tokens, completion_tokens, model, created_at)
Store token counts per message from the first commit. Retrofitting usage data onto a table that does not have it means you cannot answer any cost question about the past.
The system prompt is configuration, not a constant
Inline the system prompt as a string literal and you will regret it the first time a reply goes wrong, because you will not know which version produced it.
Version it. Store the identifier of the prompt version on each assistant message. When someone reports a bad answer, you can reproduce it exactly instead of guessing which deploy it came from.
Keep the system prompt server-side and never accept it from the client. A chat endpoint that lets the caller supply system instructions is an open proxy to your provider account, and it will be found.
Decide the history policy before you need one
Sending the entire conversation on every turn works until it does not. Costs grow quadratically with turn count, latency before the first token grows with input size, and eventually you hit the context limit mid-conversation, in front of a user.
Pick a strategy explicitly. A sliding window of recent turns is the simplest and is fine for support-style chat. Summarising older turns into a running digest preserves more but adds a call and a failure mode. Whatever you choose, count tokens before you send rather than catching the overflow error afterwards.
Always pin the system message and the first user turn if the conversation has a premise the model needs — dropping the opening message is how a bot forgets what it was asked to do halfway through.
Idempotency, or the double-send problem
Users double-click. Mobile clients retry on flaky networks. Your own retry layer replays requests. Without protection, each one produces a second inference call, a second billed response, and a duplicate message in the transcript.
Have the client generate a UUID per message and make it a unique key on insert. If the row already exists, return the stored assistant reply instead of calling the model again:
INSERT INTO messages (id, conversation_id, role, content)
VALUES ($1, $2, 'user', $3)
ON CONFLICT (id) DO NOTHING
RETURNING id;
If nothing was returned, this is a replay. Serve the existing response. This one pattern removes an entire category of support ticket and an entire category of surprise cost.
Concurrency within a single conversation
Two messages sent to the same conversation before the first reply arrives will interleave badly: both read the same history, both append, and the transcript ends up with two user turns and two replies that ignore each other.
Take a lock per conversation for the duration of a turn, or reject the second message with a clear conflict status the UI can display. Either is fine. Doing nothing produces transcripts that make no sense and are impossible to debug after the fact.
Cost controls belong in the first version
Three limits, all cheap to add, all painful to add later.
- Per-user rate limiting. Requests per minute and per day, enforced server-side. Redis with a sliding window is enough.
- Input caps. Reject an oversized message before it reaches the model. A paste of a 400-page PDF into a chat box should return a 413, not a bill.
- Output caps. Set a maximum completion length appropriate to the product. Unbounded output is unbounded cost.
Add a daily spend ceiling per account and an alert on aggregate spend. The failure you are guarding against is not a clever attacker; it is one enthusiastic user with a script, or your own bug in a retry loop.
Predictability is also why some teams put agent-style chat traffic behind flat-rate access rather than metered billing — it converts a variable risk into a fixed line item. That trade-off is worth evaluating against your actual token volumes rather than assuming either way.
Cancellation is a cost feature
When a user closes the tab or hits stop, abort the upstream request. If you only abort the browser fetch, generation continues on the provider side and you pay for tokens nobody will ever see.
Propagate the cancellation through your server to the provider call. Then persist whatever partial text was produced and mark the message as interrupted — users expect a stopped reply to stay on screen, and an interrupted flag also stops your evaluation pipeline from scoring truncated output as a bad answer.
Make failures legible in the UI
Every model call has four outcomes, not two: success, a refusal or filtered response, an infrastructure error, and a truncation. The last two are frequently rendered identically to the first, which is how users end up trusting incomplete answers.
Check the finish reason on every response. Show a distinct state for "the model stopped early" and for "something failed, retry". A generic red toast for all failures makes your own bug reports useless.
Evaluation, before you start tuning
You will change the system prompt. You will change the model. Without a fixed set of conversations to test against, every change is a vibe judgement and regressions ship quietly.
Capture 30 to 50 real conversations, redact them, and store them as fixtures. Write assertions for the mechanical things — the bot never reveals the system prompt, it always answers in the user language, it escalates when it should, tool arguments are valid. Run them in CI on every prompt change.
This does not need to be sophisticated to be useful. A test that catches "the refusal behaviour changed" before users do is worth more than an elaborate scoring rig you never look at.
A build order
- Persist conversations and messages server-side, with token counts.
- Version the system prompt and record which version produced each reply.
- Add client-generated message IDs and idempotent inserts.
- Lock or reject concurrent turns in a conversation.
- Add per-user rate limits and input and output caps.
- Stream, and propagate cancellation upstream.
- Render truncation and error states distinctly.
- Freeze a fixture set and assert on it in CI.
None of this is about the model. That is the point — the parts that determine whether a chatbot survives contact with users are ordinary application engineering, and they are the parts most prototypes skip.
Common questions
Should conversation history live in the client or the server?
The server. Client-side history is lost on refresh, cannot be audited or moderated, and lets the caller rewrite its own transcript including the system prompt.
How do I stop duplicate replies when a user double-clicks send?
Have the client generate a message UUID and make it a unique key. On conflict, return the stored assistant reply rather than calling the model a second time.
What is the cheapest way to keep chatbot costs predictable?
Cap input length, cap output length, rate limit per user, and abort upstream generation when a client disconnects. Those four together remove most runaway-cost scenarios.