Prompt Engineering for Coding Agents Is Mostly Context Engineering
AI Agents

Prompt Engineering for Coding Agents Is Mostly Context Engineering

Clever phrasing barely moves the needle. What actually improves agent output is what you put in front of it, in what order, with what constraints.

Most prompt engineering advice is written for chat. Coding agents are a different problem: the model runs many turns, accumulates its own output, and acts on a codebase. Phrasing tricks matter far less than what you put in front of it.

Constraints beat instructions

"Write good code" does nothing. Models already try to write good code — they just have a different idea of what that means than your repo does.

Constraints are specific and checkable:

  • "Do not add dependencies."
  • "Match the error-handling pattern in src/api/."
  • "Change only files under src/billing/."
  • "If the change requires touching the schema, stop and explain why."

That last one is disproportionately valuable. Giving the agent an explicit escape hatch prevents it barreling into a large change when the right answer was a conversation.

Show, do not describe

A single example of your existing pattern outperforms three paragraphs describing it. Models are far better at pattern-matching than at following abstract style rules.

If you want a service written a particular way, point at an existing service and say "follow this shape". The agent will match structure, naming and error handling more faithfully than any written convention.

Order matters more than wording

Attention is strongest at the start and end of the context. Practically:

  1. Stable rules go in the system prompt — conventions, constraints, tool guidance.
  2. Reference material goes in the middle — files, docs, schemas.
  3. The actual request goes last, immediately before generation.

The most common failure I see is burying the request above thousands of tokens of pasted context. The model reads it, then reads a great deal more, and by generation time the instruction is competing with everything that followed.

A system prompt worth having

Long system prompts are not automatically better. What earns its place:

  • Project shape — language, framework, where things live. Two or three sentences.
  • Hard rules — things that must never happen. Keep this short; a list of thirty rules gets partially ignored.
  • Tool guidance — when to search versus read, when to ask rather than guess.
  • Definition of done — tests pass, no unrelated diffs, explain what changed.

Repository convention files that tools read automatically are the right home for this. Written once, applied to every session, and version-controlled alongside the code they describe.

Ask for a plan on anything non-trivial

"Before changing anything, list the files you intend to touch and why." This costs one cheap turn and routinely prevents an expensive wrong direction. It also gives you a natural interruption point — you can correct the plan instead of reviewing a large wrong diff.

Say what to do when stuck

Models continue because continuing is more plausible than stopping. Left unaddressed, a stuck agent invents an approach rather than admitting the problem.

Tell it explicitly: "If you cannot find the definition after two searches, stop and ask." Naming the failure behaviour is the only reliable way to get it.

What does not help

  • Threats and bribes. "You will be penalised" or "I will tip you" have no consistent effect on current models.
  • Role-play preambles. "You are a 10x engineer" adds tokens and little else. Describe the task, not a persona.
  • Extreme politeness or rudeness. Neither moves quality measurably.
  • Enormous rule lists. Past a point, adherence drops across all of them rather than degrading gracefully.

Repetition is a legitimate technique

If one constraint matters more than the rest, state it twice — once in the system prompt and once immediately before the request. This feels inelegant and it works, because both positions sit where attention is strongest.

Reserve it for the one or two rules whose violation would be expensive. Repeating everything just rebuilds the wall of rules that adherence degrades against.

Give feedback as data, not adjectives

When correcting an agent mid-session, "that is wrong, try again" gives it almost nothing to work with. It will produce a variation, not a correction.

Paste the failing test output. Quote the specific line that broke the convention. Name the file it should have followed. Agents correct well from concrete evidence and poorly from disapproval — which is, in fairness, also true of people.

Iterate on the harness, not the wording

When output is poor, the instinct is to rewrite the prompt. Usually the better fix is upstream: better file selection, tighter tool descriptions, truncated tool output, a smaller task.

Change one thing at a time and re-run the same task. Prompt work without a repeatable test is superstition — you will convince yourself a rewrite helped when the difference was sampling noise.

Common questions

Do longer system prompts produce better results?

Only up to a point. Past roughly a page, adherence tends to drop across all instructions rather than degrading gracefully. Prioritise hard rules and project shape over exhaustive style guides.

Should I tell the model it is an expert?

It has little measurable effect on current models and costs tokens. Describing the task, the constraints and the definition of done is far more productive.

Why does the agent keep touching unrelated files?

It is optimising for a good outcome rather than a minimal diff. State the boundary explicitly and add a definition of done that includes no unrelated changes.

Similar articles

Agent Cost Control: Patterns That Bound the Worst Case
AI Agents
AI Agents·9 min read

Agent Cost Control: Patterns That Bound the Worst Case

Agent spend is not a per-call problem, it is a per-loop problem. Budgets, caching, tiering and circuit breakers that cap what a bad run can cost you.

Read
Agent Failure Modes: A Taxonomy Worth Memorising
AI Agents
AI Agents·9 min read

Agent Failure Modes: A Taxonomy Worth Memorising

Agents fail in about eight recognisable ways, and each one needs a different fix. A field guide to spotting them from a trace and knowing what to change.

Read
Agent Memory: Managing Context Before It Manages You
AI Agents
AI Agents·8 min read

Agent Memory: Managing Context Before It Manages You

Long agent sessions fail because context fills with noise. Here are the practical strategies for deciding what an agent should remember and what to drop.

Read