An AI agent can finish the job before you have decided what finished means.

Consider a hypothetical case at an industrial equipment supplier. A technician returns from a customer visit with handwritten notes, photographs and a fault log. The field-service manager asks an agent to prepare the service report.

Minutes later, the report looks excellent. Clear structure. Confident diagnosis. Sensible recommendations.

Then the manager reads it closely. The technician recorded a vibration; the report names a cause. The customer mentioned an earlier shutdown; the report treats it as a measured event. A recommended replacement appears without a reference to the correct equipment revision.

The agent has completed the document. The manager has inherited an investigation.

My film How to Brief AI Agents Like a Pro starts with the difference context makes to an AI answer. This essay takes that argument one step further. Once a system can retrieve files, use tools and revise its own work, the brief must describe how someone will accept the result.

My recommendation is simple: write the review before the agent starts. Decide which evidence you will need, which checks must pass and which decisions remain yours. Put those decisions in the brief.

A capable colleague still needs your standards

The colleague analogy in the film is useful. Anthropic's prompting guidance asks users to imagine a capable new employee who lacks their context. Its advice includes clear instructions, reasons for constraints and diverse examples.

That explains why “make this professional” is a weak request. Professional for whom? A maintenance engineer reading on a phone needs something different from a procurement manager comparing contractual commitments.

For the service report, the first sentence could be: “Prepare a draft the service manager can approve for the customer's maintenance lead, separating observed findings, possible causes and recommended next actions.”

Now the agent knows what the report is for. It also has a distinction to preserve. A possible cause belongs in a different category from a recorded measurement, however plausible it sounds.

The next sentence explains the stakes: “The customer may use this report to decide what to repair, so every technical recommendation needs evidence applicable to this equipment.” That reason helps with cases the brief does not spell out.

An example makes the standard visible. Show a previously approved report with permission to use it. Explain which qualities to borrow: the separation of evidence from diagnosis, the restrained language, the way open questions remain visible. A second example with an unresolved fault can show that a useful report need not pretend the investigation is complete.

There is little reason to decorate this with motivational theatre. In a 2025 Wharton study, tips and threats produced no reliable overall improvement on the academic benchmarks tested. A companion study on chain-of-thought prompting found that the benefit depended on the model and task; explicit reasoning models often gained little while using more time and tokens. These are bounded experiments, not verdicts on every prompting technique.

For practical work, I would spend the next sentence on the acceptance criteria. “Every recommended action must point to an applicable source” gives the reviewer something useful. A demand to try harder does not.

Give it a map of the evidence

In Building effective agents, Anthropic describes agents using tools and feedback from their environment in a loop, with the option to pause for human judgment. This matters: an agent can investigate and revise without acquiring authority to approve its own conclusions.

But the loop only helps if it can reach the right evidence.

For our hypothetical report, give it the visit folder, the fault-log file, the equipment identifier and the approved manual revision. State the role of each. Technician notes record what was observed. Customer statements record what the customer reported. The manual supports applicable procedures. An old report demonstrates format.

Those sources are useful in different ways. The old report cannot establish the cause of today's fault. The customer's description cannot silently become a measurement. The manual for a similar machine cannot establish compatibility.

Anthropic's context-engineering article describes a practical approach: provide references such as paths and links, then let the agent retrieve relevant material as it works. It also warns that more context has diminishing returns. My inference is that a well-labelled evidence map can be more useful than an indiscriminate document dump.

That requires actual access. A filename in a prompt does not connect a folder. Before delegating, check that the agent can open the named files and read the relevant formats. If it cannot inspect the photographs, say so in the handover. Let missing evidence remain missing.

Then define how to handle conflicts. If the notes and equipment record show different serial numbers, complete the supported sections and raise that specific discrepancy. Avoid replacing a missing answer with a fluent guess, or stopping all useful work because one section needs a person.

A brief someone can approve

Here is the brief I would start with for that service report. It is a proposed working template, to adapt to the tools and permissions available.

Outcome and reader

Prepare a customer-facing service-report draft for the service manager to review. The reader is the customer's maintenance lead. Keep observed findings, hypotheses and recommended actions separate. Save it as [output format] in [output path] by [deadline].

Sources and their roles

Use [visit folder], including [technician notes], [photographs] and [fault log], plus [equipment record]. Use [approved equipment manual and revision] for this equipment. Follow [approved report sample]'s structure. Treat customer statements as reported information and the sample as a format reference.

Acceptance checks

Verify the equipment identity against the record. Trace each technical finding and recommendation to a file and precise location where available. Check that dates, units and component identifiers agree across the draft. Flag missing or conflicting evidence. Include the next action and proposed owner for each open question.

Permission boundaries

Read the supplied sources and create a new draft in [output path]. Do not change source records, update the service system, send the report or place an order. The service manager approves customer delivery and technical commitments.

Escalation and stopping

Complete supported work. Ask a specific question when a conflict prevents an acceptance check. If a required tool is unavailable, report the blocker without repeated attempts. Work within [time or spend limit] and finish by [deadline]. If either limit is reached, save a partial draft and list what remains.

Handover

Return the draft, a short evidence appendix, the checks performed and any failures. Label it “draft for review”. Separate work completed from decisions awaiting approval.

The template gives the agent room to work and gives the manager a bounded review. The appendix need only make consequential claims easy to trace; reproducing the whole input would create another document to inspect.

If the equipment identity or applicable manual revision cannot be verified, the draft can still describe supported observations. Equipment-specific recommendations remain unresolved. The service manager must close that gap before approving those recommendations for customer delivery.

Acceptance checks also need something outside the prose. Dates can be compared with a record. A total can be recalculated. A recommendation can be checked against the relevant manual passage. Asking the same model whether its report looks convincing is a weaker test.

Permissions need the same treatment. Where the system supports it, enforce the boundary through access controls and approval gates. A sentence saying “do not send” describes intent; a tool that cannot send without approval makes that boundary harder to cross.

The objection that deserves an answer

“If I have to specify all this, I might as well do the work myself.”

For some tasks, that is correct. A brief can become a task of its own. If you need one short paragraph and can check it immediately, an ordinary conversation may be enough. An agent earns its place when retrieving evidence, carrying out several steps and repairing errors adds value you can verify.

The more serious objection is that expertise resists checklists. A service manager knows which symptom is unusual, which explanation is weak and which customer needs a phone call. A template cannot transfer all that judgment.

I agree. The point of the brief is to expose the boundary of delegated work. It should leave the expert with a visible decision, rather than an invisible reconstruction of what the agent assumed.

More agents do not remove that obligation. Anthropic's multi-agent research report describes benefits for searches that can run in independent directions, alongside roughly fifteen times the token use of chat in its own data. The company also identifies tasks with shared context and many dependencies as a poor fit. That is a design tradeoff, not a general instruction to assemble a team.

For this report, separate investigations might help with a large evidence set. They would still need one consistent equipment identity, source standard and handover. Several confident drafts can leave the manager with several sets of assumptions to reconcile.

My practical test is to reuse the brief on a few completed visits, including one with missing evidence. Compare the result with the approved record. Could the reviewer find unsupported claims? Did the agent preserve uncertainty? Did it stop at the agreed boundary? Revise the brief around the failures that matter.

Keep useful corrections. If a reviewer repeatedly changes a diagnosis into a hypothesis, that distinction belongs in the next brief and its example. The standard gets clearer with use.

The next time you delegate a task, pause before describing the work. Imagine the result arriving on your desk. What would you inspect before accepting it? Which missing fact would send it back? What must happen before it leaves your control?

Write those answers first.

Then give the agent the job.