A team writes the spec for an account agent. It opens with identity language: “You are a helpful account assistant.” Then it lists tasks: summarize activity, identify risk, recommend next steps, draft outreach.
The agent will stay busy while the manager inherits the cleanup. What struggling moment is it there to relieve? Which evidence should it trust, and which decisions stay human? Where should it stop? A list of tasks leaves every one of those open.
A task list makes an agent busy, not useful
An AI agent is software that takes on work instead of waiting for each instruction. The user expresses an intended outcome, the system performs some portion of the work, and the user stays responsible for evaluating, steering, approving, or correcting the result. A coding agent writes the change. A research agent synthesizes the sources. A workflow agent routes the request. With a tool, you manipulate the interface to produce the outcome. With an agent, you hand off the producing part and keep the deciding parts.
A weak agent has a task. A good agent has a job. “Summarize this meeting” is a task. “Help the team leave the meeting with clear decisions, accountable next steps, and unresolved risks visible enough to act on” is a job.
A task list tells the agent what it may produce. A job spec tells it what burden it’s supposed to remove, what evidence it can trust, which decisions remain human, and how the user will know the work is good. It keeps the agent from confusing output with progress.
Without that layer, the agent optimizes for completion. It invents certainty when the evidence is thin. It automates the visible step and leaves the user with the rest.
The job starts with the struggling moment
Write the job around a situation, not a feature. The situation is the job beneath the prompt, and it holds more than any feature name.
“When a renewal is approaching and recent activity suggests the account may be at risk, help the customer-success manager understand what changed and choose a defensible next move, so they can intervene before the relationship deteriorates without wasting time rebuilding the account history.”
That one sentence gives the agent more than “summarize account.” It identifies the trigger, the progress, the user, the time pressure, and the burden to remove. Without the trigger scene, the agent responds generically to a task label. With it, the behavior stays attached to the user’s actual moment of struggle.
A good job spec then expands the sentence into operating guidance.
A job spec has eight parts
Each part answers a question the agent would otherwise answer for itself.
Permission mode decides how far the agent goes on its own. Capability isn’t permission, so the spec names the mode and the conditions for stepping down. Handoff and recovery decide what the user gets back when the agent stops. That return is the pass-back, and it needs the same care as the output. “I can’t help” isn’t a boundary design. The user needs the next route, what context will travel with them, and what to expect while they wait.
Fill-in card
The eight parts of an agent job spec
Answer each field before the agent runs.
- Trigger
- What event activates the job? A date, signal, request, threshold, failure, or change in state.
- Desired progress
- What should become easier, clearer, safer, or finished? Include functional, emotional, and social progress.
- Objects and scope
- Which customer, document, workflow, project, or decision is in bounds? What related objects may the agent inspect or change?
- Evidence standard
- Which sources are authoritative? How current must they be? What happens when sources conflict or are missing? When does the agent say “I don’t have enough evidence”?
- Permission mode
- Does the agent observe, suggest, prepare, act after approval, or act autonomously? Under which conditions should it step down to a safer mode?
- Boundaries
- What can’t the agent infer, promise, publish, delete, or change? What can be drafted but not sent? Which actions are too consequential to automate?
- Definition of done
- What must be true before the job can close?
- Handoff and recovery
- When should the agent ask, escalate, stop, preserve partial work, or return control? What can the user undo?
“Produced a response” is weaker than “produced a response the user can defend, with unsupported claims flagged and next steps clear.” The second is a definition of done the agent can be held to.
Progress has three layers, and each one changes the output
Functional success alone can hide a bad experience. An agent can be safe to trust with one layer and reckless to trust with another inside the very same task.
For the renewal job, functional progress is understanding the signals and preparing an intervention. Emotional progress is feeling prepared rather than blindsided. Social progress is being able to explain the recommendation to a manager and contact the customer without sounding generic or alarmist. Enterprise users are often evaluated on the outcomes of their decisions, and they can’t delegate authority to an agent unless they can defend the agent’s actions to their manager, their compliance team, or their client.
Those layers change what the agent returns. It needs receipts, not only a score. It should distinguish observation from inference. It should preserve the manager’s authorship. It should avoid language that exposes an internal risk classification to the customer.
That’s why the job spec is a design artifact. It affects interface states, controls, copy, and permissions—not only model instructions.
Conditional behavior makes decisions explicit
Adjectives describe a character. Conditions describe a decision. “Be helpful” and “be accurate” aren’t wrong. They’re insufficient: helpful for which job, accurate against which source? Replace vague trait instructions with conditions that change the agent’s behavior.
Fill-in table
Conditional behavior
Write one row for each situation that should change the agent’s decision.
| Situation that changes the job | What the agent must check | What it should do |
|---|---|---|
| [Describe the event or condition.] | [Name the evidence, uncertainty, risk, or permission boundary.] | [State the action, question, limit, or handoff.] |
Example
The renewal job, specified
A renewal is approaching, and recent activity suggests the account may be at risk. The agent must help the customer-success manager choose a defensible next move without taking that choice over.
- Trigger: a renewal is approaching and recent activity suggests the account may be at risk.
- Desired progress: the customer-success manager understands what changed and chooses a defensible next move; feels prepared rather than blindsided; can explain the recommendation to a manager.
- Objects and scope: the account at renewal and its recent activity.
- Evidence standard: attach each risk claim to a source; if the only evidence is more than ninety days old, label the recommendation low-confidence.
- Permission mode: prepare the intervention; the manager chooses the next move. Show conflicting evidence before approval.
- Boundaries: don’t expose the internal risk classification to the customer; don’t draft external outreach on stale evidence alone.
- Definition of done: a response the manager can defend, with unsupported claims flagged and next steps clear.
- Handoff and recovery: escalate when financial, legal, safety, or public commitments are involved, or when the user can’t reverse the action.
A review checklist
If the answers to the checklist below live only in someone’s head, the agent doesn’t have a job spec. It has a prompt and permission to improvise.
Before shipping an agent:
0 of 8 done