Prompt Engineering for Coding Agents Is Mostly Context Engineering
Clever phrasing barely moves the needle. What actually improves agent output is what you put in front of it, in what order, with what constraints.
Most prompt engineering advice is written for chat. Coding agents are a different problem: the model runs many turns, accumulates its own output, and acts on a codebase. Phrasing tricks matter far less than what you put in front of it.
Constraints beat instructions
"Write good code" does nothing. Models already try to write good code — they just have a different idea of what that means than your repo does.
Constraints are specific and checkable:
- "Do not add dependencies."
- "Match the error-handling pattern in
src/api/." - "Change only files under
src/billing/." - "If the change requires touching the schema, stop and explain why."
That last one is disproportionately valuable. Giving the agent an explicit escape hatch prevents it barreling into a large change when the right answer was a conversation.
Show, do not describe
A single example of your existing pattern outperforms three paragraphs describing it. Models are far better at pattern-matching than at following abstract style rules.
If you want a service written a particular way, point at an existing service and say "follow this shape". The agent will match structure, naming and error handling more faithfully than any written convention.
Order matters more than wording
Attention is strongest at the start and end of the context. Practically:
- Stable rules go in the system prompt — conventions, constraints, tool guidance.
- Reference material goes in the middle — files, docs, schemas.
- The actual request goes last, immediately before generation.
The most common failure I see is burying the request above thousands of tokens of pasted context. The model reads it, then reads a great deal more, and by generation time the instruction is competing with everything that followed.
A system prompt worth having
Long system prompts are not automatically better. What earns its place:
- Project shape — language, framework, where things live. Two or three sentences.
- Hard rules — things that must never happen. Keep this short; a list of thirty rules gets partially ignored.
- Tool guidance — when to search versus read, when to ask rather than guess.
- Definition of done — tests pass, no unrelated diffs, explain what changed.
Repository convention files that tools read automatically are the right home for this. Written once, applied to every session, and version-controlled alongside the code they describe.
Ask for a plan on anything non-trivial
"Before changing anything, list the files you intend to touch and why." This costs one cheap turn and routinely prevents an expensive wrong direction. It also gives you a natural interruption point — you can correct the plan instead of reviewing a large wrong diff.
Say what to do when stuck
Models continue because continuing is more plausible than stopping. Left unaddressed, a stuck agent invents an approach rather than admitting the problem.
Tell it explicitly: "If you cannot find the definition after two searches, stop and ask." Naming the failure behaviour is the only reliable way to get it.
What does not help
- Threats and bribes. "You will be penalised" or "I will tip you" have no consistent effect on current models.
- Role-play preambles. "You are a 10x engineer" adds tokens and little else. Describe the task, not a persona.
- Extreme politeness or rudeness. Neither moves quality measurably.
- Enormous rule lists. Past a point, adherence drops across all of them rather than degrading gracefully.
Repetition is a legitimate technique
If one constraint matters more than the rest, state it twice — once in the system prompt and once immediately before the request. This feels inelegant and it works, because both positions sit where attention is strongest.
Reserve it for the one or two rules whose violation would be expensive. Repeating everything just rebuilds the wall of rules that adherence degrades against.
Give feedback as data, not adjectives
When correcting an agent mid-session, "that is wrong, try again" gives it almost nothing to work with. It will produce a variation, not a correction.
Paste the failing test output. Quote the specific line that broke the convention. Name the file it should have followed. Agents correct well from concrete evidence and poorly from disapproval — which is, in fairness, also true of people.
Iterate on the harness, not the wording
When output is poor, the instinct is to rewrite the prompt. Usually the better fix is upstream: better file selection, tighter tool descriptions, truncated tool output, a smaller task.
Change one thing at a time and re-run the same task. Prompt work without a repeatable test is superstition — you will convince yourself a rewrite helped when the difference was sampling noise.
Common questions
Do longer system prompts produce better results?
Only up to a point. Past roughly a page, adherence tends to drop across all instructions rather than degrading gracefully. Prioritise hard rules and project shape over exhaustive style guides.
Should I tell the model it is an expert?
It has little measurable effect on current models and costs tokens. Describing the task, the constraints and the definition of done is far more productive.
Why does the agent keep touching unrelated files?
It is optimising for a good outcome rather than a minimal diff. State the boundary explicitly and add a definition of done that includes no unrelated changes.