How to Write a Goal for an AI Agent
How to write a goal for an AI agent: four fields that stop it redefining done, a before and after rewrite, and how to test the goal before you trust it.
A goal for an AI agent is not a longer prompt. A prompt asks for output you are about to read. A goal is an instruction that has to survive without you: the agent will interpret it while you sleep, hit something unexpected, and decide on its own whether it is finished. Four fields make that survivable. State the objective, state the evidence that proves it is done, state the boundaries it must not cross, and state what should make it stop and ask. Miss the second field and the agent picks its own finish line, which it will do generously.
Why a good prompt makes a bad goal
A prompt lives inside a conversation where you are the error handler. You read the output, notice the gap, and follow up. That loop is doing enormous work, and it disappears the moment the agent runs unattended.
What breaks, specifically:
Nobody checks whether done means done. The agent decides, and it tends to decide early.
Nobody notices scope drift. A task that touches an adjacent problem looks like initiative in a summary and like an incident in a diff.
Nobody sees the moment the agent was confused, which is exactly the moment a human would have asked a question.
That first failure has a name and a public example. OpenAI's safety lead described a model that improved on model laziness while falling short on staying within scope, a trade we covered in what model laziness is. Vague goals sit right on that fault line: an agent that cannot tell what finished looks like will either stop too early or reach too far, and often both in one run.
The four fields
1. Objective, in one sentence with a subject
Write what will be true when the work is done, not the activity you want performed. Improve the onboarding flow is an activity. New users reach the dashboard without hitting the email verification dead end is an outcome, and it can be checked.
2. Evidence of done
This is the field that does the work, and the one almost everybody omits. Name the artefact you will look at to decide the task succeeded. A passing test named in advance. A file that exists. A row in a table. A screenshot of a specific state.
The test for this field: could a stranger with no context look at your evidence and tell you yes or no? If deciding requires judgement, you have written a preference, and the agent will exercise that judgement instead of you.
3. Hard boundaries
A short list of things that are off limits regardless of how helpful they would be. Files it must not touch, systems it must not call, actions it must not take without approval. Keep it to five items. A boundary list long enough to be exhaustive is long enough to be skimmed.
4. Escalation trigger
Say what should make it stop and come back rather than proceed. This is the field that converts a silent wrong answer into a question, and it is worth more than any amount of prompt polish. Good triggers are concrete: a missing credential, an ambiguous requirement, a test that was already failing before it started, more than two attempts at the same fix.
A before and after
Here is a goal of the kind people actually write:
Clean up the API error handling in the payments service and make it consistent.
Add tests where they're missing. Let me know when it's done.Everything is wrong with this in a way that is invisible until it runs. Consistent with what? Which errors? Does clean up include changing response codes, and if so, who is relying on them? Where are tests missing, and is more coverage the goal or is passing the goal? The agent will answer all of these and tell you it is done.
The same task, written as a goal:
OBJECTIVE
Every handler in services/payments/ returns errors in the shape defined in
docs/errors.md. No handler returns a raw exception message to the caller.
EVIDENCE OF DONE
- tests/payments/test_error_shape.py passes (it exists and currently fails)
- grep for 'str(e)' in services/payments/ returns nothing
- the full existing test suite still passes, with no test modified
HARD BOUNDARIES
- do not change any HTTP status code
- do not touch services/billing/ or any file outside services/payments/
- do not add dependencies
- do not modify or delete an existing test
ESCALATE AND STOP IF
- a handler's correct error shape is genuinely ambiguous from docs/errors.md
- a test was already failing before you started
- fixing one handler appears to require changing a status codeThe second version is longer to write and vastly cheaper to review, because every claim in the agent's summary is now checkable in seconds. Note especially the do not modify an existing test boundary. Without it, the cheapest route to a green suite is editing the suite, and agents find cheap routes.
The evidence-of-done field is close kin to acceptance criteria, and the same discipline applies. How to write acceptance criteria for an AI coding agent goes further on the coding-specific version.
What to leave out
Goals rot by accumulation. Each disappointing run tempts you to add a paragraph, and after a month you have two pages the agent reads as ambient noise. Three things to keep out:
Method, unless the method is the requirement. Telling the agent how to do it converts a checkable outcome into an unfalsifiable process.
Encouragement. Be thorough and think carefully cost tokens on every wake and change nothing.
Anything already enforced elsewhere. If a linter or CI blocks it, the linter is more reliable than a sentence.
The exception is context about your situation that the agent cannot infer, which is a genuinely different thing from instruction. How to give AI context about your business covers where that belongs.
Test the goal before you trust it
A goal is a piece of software. Run it against reality before it runs unsupervised:
Ask the agent to restate the goal and list what it will check. If the restatement differs from your intent, the goal is wrong, not the agent.
Ask it for its plan and read it. Cheap, and it catches most misreadings before any work happens.
Run it once while watching. The point is to see where it hesitated.
Read the diff against your evidence list, not against the summary. The summary is the agent's opinion of its work.
Then let it run unattended, and check the first unattended run properly.
Step two has its own method, in how to review an AI agent plan before it runs, and the reporting problem in step four is the subject of what to do when an AI coding agent says it is done.
Why this is becoming the core skill
The goal is now the product surface. When OpenAI announced its always-on Dots agents, the described workflow is that you give a dot a goal and it works toward that goal continuously, bringing completed work back for approval. In that arrangement the goal is the entire specification, written once, executed for days. Prompt craft becomes specification craft.
The broader technique stack sits in our prompt engineering guide.
FAQ
How long should a goal for an AI agent be?
Long enough to include evidence of done and boundaries, which is usually 10 to 25 lines. Past about a page you are adding instructions that get skimmed. If it keeps growing, the task is too big and should be split.
What is the difference between a goal and a prompt?
A prompt is evaluated by you immediately. A goal is evaluated by the agent, later, without you. That is why a goal needs a checkable definition of done and an escalation trigger, while a prompt can rely on you noticing.
Should I tell the agent how to do the task?
Only when the method is part of the requirement, such as a library you must use. Otherwise specify the outcome and the evidence. Prescribing method makes failures harder to detect, because the agent can follow your steps and still miss the point.
What if the agent keeps stopping to ask questions?
That usually means the goal is genuinely ambiguous, and the questions are telling you where. Answer them by editing the goal rather than by replying in chat, so the next run inherits the fix.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


