Building Your First Agent: A Loop, Three Tools, One Guard
Skip the framework. An agent is a while loop over a tool-calling model. Here is the minimum that works, and the four things that break it first.
The fastest way to understand agents is to write one, and it is much smaller than the ecosystem around it suggests. Anthropic makes the same point in its guidance: the most successful implementations tend to use simple composable patterns rather than frameworks, and many of them are a few lines against the API directly.
Start there. You can adopt a framework later, once you know which of its abstractions you actually need.
The loop is fifteen lines
Everything else is detail on top of this:
messages = [system_prompt, user_task]
for step in range(MAX_STEPS):
reply = model(messages, tools=TOOLS)
messages.append(reply)
if not reply.tool_calls:
return reply.text # model is done
for call in reply.tool_calls:
result = dispatch(call.name, call.args)
messages.append(tool_result(call.id, result))
raise StepLimitExceeded()
Three things are doing all the work. The model chooses actions. Your dispatcher executes them. Results go back into the transcript so the next call sees what happened. The step cap is there because without it the loop has no upper bound.
If you understand that, you understand agents. What follows is how to keep it from falling over.
Three tools, chosen deliberately
Resist adding tools. Each one enlarges the space of wrong choices, and every definition sits in the context on every call. A useful first agent needs about three:
- A read tool. Fetch a file, a record, a page. Support ranges or limits so it cannot return a megabyte.
- A search tool. Find candidates by pattern or query. Without this the model guesses at names, which is the single most common source of wasted steps.
- An act tool. The one thing your agent exists to do — run a command, write a file, call an API. Keep it narrow.
Write the descriptions for a competent new colleague who cannot ask questions. Say what the tool does, when to use it instead of the other one, what the arguments mean, and what happens when there is no result. Vague descriptions are the reason agents pick the wrong tool, and the fix is almost always in the description rather than the system prompt.
Errors are observations, not exceptions
The instinct is to raise when a tool fails. Do not. A failed tool call is information the model can act on, and handing it back is what makes the loop self-correcting.
try:
result = dispatch(call.name, call.args)
except ToolError as e:
result = f"Error: {e}. Try search_files to find the correct path."
Two properties make an error message good: it says what went wrong specifically, and it suggests what to do instead. File not found: src/uti1s.py. Similar files: src/utils.py gets recovered from in one step. Error: ENOENT gets recovered from in four, if at all.
Reserve real exceptions for things the model cannot fix — expired credentials, a service that is down, a policy refusal. Those should terminate the run and surface to you.
Stopping is harder than starting
A first agent usually fails in one of four ways, all of them about termination.
It never stops. The step cap catches this. Set it from observation rather than intuition: run twenty real tasks, look at the step distribution, cap at roughly twice the highest normal value.
It repeats itself. Same tool, same arguments, over and over. Hash each call and detect repeats. On a repeat, append an observation saying the call was already made, with the previous result. This unsticks it more often than you would expect.
It stops too early. The model declares success without verifying. Fix this in the loop, not the prompt: before accepting a final answer, run a check — tests pass, schema validates, the file exists — and if it fails, feed that back as another observation.
It claims success falsely. The most dangerous one, because the trace looks clean. The only reliable defence is an automatic verification step your code controls. Never take the model word for whether the job is done.
One guard before you run it on anything real
If your act tool can modify anything you care about, gate it.
The minimum viable version is a confirmation prompt that displays the concrete action — the actual command, the actual diff, the actual request body — and waits. Not "the agent wants to run a command" but the command itself. People approve summaries reflexively; they read specifics.
Better, if you can arrange it: give the agent a workspace it can destroy. A scratch copy of the repository, a throwaway database, a container that gets deleted afterwards. Then the guard is structural rather than behavioural, and you can let the loop run unattended.
Log enough to debug it
Agent failures are hard to reproduce, because sampling is stochastic and the world moves. Record, per step: the tool name, the full arguments, the full result, the token counts and the finish reason.
Do this on day one. It costs almost nothing and it is the difference between reading what happened and guessing at it. Nearly every hour I have seen wasted on a misbehaving agent was spent reconstructing state that a print statement would have captured.
What to add, in order
- Verification before accepting completion. Biggest single reliability gain.
- Context management once runs get long — prune superseded tool results, then summarise older turns.
- Prompt caching, once the prompt prefix is stable. Agent loops are close to the ideal caching workload.
- Structured output for the final answer, if something downstream consumes it.
- Retries with variation — on a second failure, change the approach rather than repeating the call.
- A framework, last, and only when you can name the specific thing it gives you.
Build the loop, watch it fail on real tasks, and fix what actually broke. That order produces a system you understand. Starting from a framework produces one where the failures happen inside code you did not write.
Common questions
Do I need a framework to build an AI agent?
No. The core loop is roughly fifteen lines against a tool-calling API. Start direct, learn where it breaks on your own tasks, and adopt a framework only when you can name the specific abstraction you want from it.
How many tools should a first agent have?
About three: read, search and one action tool. Every extra tool enlarges the space of wrong choices and occupies context on every call. Tool descriptions matter more than count — most wrong-tool errors trace back to a vague description.
How should an agent handle tool failures?
Return the error to the model as an observation rather than raising. Include what went wrong and what to try instead, so the loop can self-correct. Reserve real exceptions for problems the model cannot fix, such as expired credentials.