System Prompts: What They Actually Do and How to Write One
Fundamentals

System Prompts: What They Actually Do and How to Write One

A system prompt is not a control panel and not a security boundary. Here is what it really is, where it sits in the instruction hierarchy, and how to write one.

Most teams treat the system prompt as a settings panel: write the rule, the model obeys the rule. It is closer to the opening paragraph of a document the model is trying to continue — text with unusually strong influence, but text all the same.

Understanding that distinction explains almost every surprise people hit. Why a rule gets ignored at turn forty. Why "never do X" sometimes produces X. Why a user message can talk the model out of something you thought was fixed.

What a system prompt actually is

At inference time there is one token sequence. Your system prompt, the conversation history, tool definitions and tool results are all concatenated into it, separated by role markers the model was trained to recognise.

The special status of the system role is entirely learned. During post-training, models are taught that content in that position carries more authority than content in the user position, and they are rewarded for following it under pressure. That training is strong, and it is not absolute.

So a system prompt is a very high-prior instruction, competing with everything else in context. Ten thousand tokens of tool output that contradict it will sometimes win.

The instruction hierarchy

Providers formalise this ordering. OpenAI publishes a Model Spec that defines a chain of command from root rules, through provider-set system instructions, then developer instructions supplied through the API, then user messages, then softer defaults it calls guidelines. Higher levels are meant to override lower ones.

Two practical consequences. First, you as an API caller usually sit at the developer level, not the top — you cannot instruct the model out of the provider policy above you. Second, everything downstream of you, including anything a user types and anything a tool returns, sits below you in that ordering and should be treated as lower-trust data.

Anthropic exposes this as a separate system parameter rather than a message with a role. Different shape, same idea: content placed there is weighted differently from conversation turns.

It is not a security boundary

This is the mistake that causes real incidents. A system prompt is a strong behavioural prior, not an access control. If your agent has a tool that can delete production data, "only delete when the user explicitly confirms" in the system prompt is a suggestion, not a permission check.

The boundary belongs in your code: a permission layer that decides whether a requested tool call is allowed, independent of what the model was told. Treat the system prompt as documentation of intent, and the harness as the thing that enforces it.

The same reasoning applies to secrets. Anything in the system prompt can surface in output, directly or through paraphrase. Do not put credentials there.

What belongs in it, and what does not

Belongs:

  • Role and scope. What this assistant is for, and explicitly what it is not for.
  • Hard constraints. The handful of rules that must hold on every single turn.
  • Output contract. Format, length, whether to ask clarifying questions or proceed on assumptions.
  • Tool policy. When to reach for a tool, when to answer directly, what requires confirmation.
  • Stable domain context. Facts that are true for every request in this application.

Does not belong: anything that changes per request. Retrieved documents, the current user record, today's date computed at call time, the file the agent is editing. Those go in the conversation, for reasons covered in the next section.

Position, stability and cost

A system prompt sits at the very front of the sequence, which makes it the natural cache prefix. Providers cache the prefill work for a repeated prefix, so a system prompt that is byte-identical across requests is cheap and fast, while one containing a timestamp or a session ID invalidates the cache on every call.

Position also affects attention. The long-context literature — most famously the Lost in the Middle work on how models use long inputs — finds a U-shaped curve: material at the very start or the very end of context is used more reliably than material buried in the middle.

Early in a conversation your system prompt is at a strong position. Forty turns later it is a long way from the end. If a constraint absolutely must hold, restate it near the end of context — in the latest user message, or in a short reminder appended before each model call.

Writing one that holds up

  • Say what to do, not only what to avoid. "If the request is out of scope, reply with one sentence saying so and stop" beats "never answer out-of-scope questions".
  • Keep the hard rules short and countable. Five rules the model follows are worth more than thirty it averages over.
  • Resolve conflicts yourself. If two instructions can disagree, state which wins. The model will pick one otherwise, and not always the same one.
  • Show one example of the exact output shape. A single worked example does more than a paragraph of description.
  • Delete instructions that never fire. Prompts accumulate rules added for a bug that is long fixed. Every dead rule dilutes the live ones.

Treat it as code

System prompts drift because they are edited under pressure and nobody measures the result. Put yours in version control, not in a dashboard text box. Give it a version identifier you log alongside each request, so you can tell which prompt produced which behaviour.

Then keep a small set of cases that exercise each hard rule, and re-run them whenever you edit. Prompt changes are the least reviewed and most behaviour-altering deploys most teams ship — a twenty-case regression suite catches the majority of self-inflicted regressions before users do.

Common questions

Can a user override my system prompt?

Sometimes. Models are trained to weight system instructions above user ones, but that is a learned preference rather than an enforced rule. Anything that must hold has to be checked in your code.

How long should a system prompt be?

As long as the stable rules require and no longer. Length itself is not the problem; contradictory and dead rules are. Move anything that changes per request into the conversation instead.

Why does the model forget a system prompt rule in long conversations?

Attention favours the start and end of context, so an instruction that was near the front drifts into the weak middle as history grows. Restate critical constraints close to the latest turn.

Similar articles

Chain of Thought: Why Thinking Out Loud Actually Helps
Fundamentals
Fundamentals·8 min read

Chain of Thought: Why Thinking Out Loud Actually Helps

Asking a model to reason step by step measurably improves accuracy on some tasks and wastes tokens on others. The mechanism, and when it is worth the cost.

Read
The Lost-in-the-Middle Problem: Position Beats Relevance
Fundamentals
Fundamentals·8 min read

The Lost-in-the-Middle Problem: Position Beats Relevance

Models retrieve facts from the start and end of a long prompt far more reliably than from the middle. Why that happens and how to arrange prompts around it.

Read
Attention Mechanisms Explained Without the Linear Algebra
Fundamentals
Fundamentals·9 min read

Attention Mechanisms Explained Without the Linear Algebra

What attention actually computes, why it made transformers work, and why its cost scaling explains almost every practical limit you hit with long context.

Read