Microsoft's Copilot Studio documentation describes its agents as "autonomous AI assistants" that determine the actions to take based on conversation context. The phrase joins two words that point in opposite directions, and it comes from the company that defined the category for a generation of executives.

When the vendor that named the thing blurs the distinction in its own documentation, every executive downstream inherits the blur. It shows up in roadmaps, steering committees and budget lines, where a configured chatbot and a system that completes work on its own end up under the same heading.

My argument is simple. Calling an assistant an agent is a budgeting error before it is a vocabulary error. It sends transformation money to productivity tools and leaves the work that would change the operating model unfunded. The fix is a test anyone can run in a meeting in under ten seconds.

The gap that should bother you

In McKinsey's late-2025 survey, roughly 88% of organizations reported regular AI use, while only about 6% attributed more than five points of EBIT to it. Adoption is close to universal. Measurable value is rare.

Gaps that wide rarely come from the technology alone. In the transformations I have worked on, they more often come from a definition problem: an organization deploys something, calls it an agent, declares the transformation underway and then gets the returns of a slightly faster chatbot. You cannot get agent-class returns from assistant-class deployments, so the first step is being able to tell the two apart.

The Button Test

Here is the test I use.

Does a human still have to press a button for the core loop to close? If yes, it is an assistant, however much knowledge you added to it.

It is deliberately crude, because the arguments it settles are usually vague. "It has access to all our documents." Does a human still press the button? "We wrote forty pages of custom instructions." Does a human still press the button? "It uses our private knowledge base." The answer tends to come back the same way.

Custom instructions, system prompts and document retrieval are real and useful. They improve the output, but they do not let the system act. They produce a better-briefed assistant. Agency needs three things at once: an autonomous decision, an action taken across systems, and a loop that closes without a human in the middle. Remove any one of them and a person is still pressing the button.

The Agency Ladder

Teams talk past each other because "agent" is being used to name six different things. It helps to name them separately and to be honest about where the work sits.

Adding custom instructions and document access moves a deployment from rung 2 to rung 3. It feels like a leap, but it is one step on a six-step ladder, and the cheapest one. In the roadmaps I see, most systems described as agents sit at rung 3: configured assistants that wait, retrieve and hand the work back to a person.

The work that changes an operating model starts at rung 5, where the system takes a goal, decides the steps, acts across systems and closes the loop, escalating only the exceptions. That is also the work that tends to go unfunded, because the organization has already declared victory at rung 3 and rung 5 looks risky next to a box that is already checked.

How the word got stretched

Two forces did most of the stretching. The first is ordinary vendor incentive: "agent" sells better than "assistant", so products get relabeled upward. The second is more specific, and it is worth being precise about it so the criticism stays fair.

Microsoft did not blur the line because it lacks real agents. Its own preview showed an agent that reads a purchase request from email, checks the requester's department budget in Dynamics 365, verifies supplier inventory through an API, drafts a purchase order in SharePoint and schedules an approval flow, with no human typing after the initial configuration. That is a genuine rung-5 agent.

The same word is then applied to a SharePoint question-and-answer bot, because Copilot Studio is the renamed Power Virtual Agents, a chatbot builder. Three management layers down, "we have Copilot" is heard as "we have agents". Microsoft's best agent demo is the clearest evidence that a document bot is something else.

What the blur costs

This is the part a CFO should care about, because misclassification moves money.

When a rung-3 assistant is filed under transformation, it gets transformation budget and transformation patience: prompt libraries, more single-purpose bots, another knowledge integration. It looks like maturity on a slide. The rung-5 work, a loop-closing agent with a named owner and a real consequence if it goes wrong, looks risky next to that checked box and gets deferred. The cheap rung is overfunded because it is safe and visible, and the valuable rung waits.

That is one plausible reading of the gap between 88 and 6: a budgeting decision more than a technology limit.

An assistant makes a person faster. The gain is linear and usually measured in minutes saved per task, with the person still in the loop and still pressing the button, only sooner. That is real productivity and worth paying for. An agent removes a class of work from someone's queue. That kind of change can alter an operating model, and it is the kind that shows up in a P&L. If you are paying transformation prices and seeing efficiency results, check which rung you actually bought.

The strongest objection

A pragmatic CIO will say the label does not matter. Assistants deliver measurable productivity today, rung-5 agents carry real risk, and many processes should keep a person approving the outcome.

Most of that is right. Assistants are worth buying, and some loops should never close without a person. When an output commits a price, a safety claim or a customer promise, a named human check is a control rather than a failure, as I argue in The Verification Ceiling.

The objection still misses what the test is for. The Button Test asks you to know which kind of system you are funding and to measure it accordingly, and it leaves room for loops that should keep a person in them. Keep humans where the consequence requires them, and stop describing those systems as agents in the budget.

What changes in the budget

Four changes follow, in order.

  1. Fund assistants on the productivity line. They are useful and fundable, and they should be measured as time saved.
  2. Ringfence budget for loop-closing agents with named P&L ownership. If no executive will put their number behind the system, treat it as an experiment.
  3. Make any system that wants the title pass the Button Test before it gets the name and the budget that comes with it.
  4. Measure tasks removed from human queues alongside deployment and adoption figures, because loop closure is what changes the operating model.

Monday move

Take the most important "agent" on your roadmap and draw its loop on one page. Mark every point where a person must copy, approve, send or verify. Keep the points where the consequence justifies a human, remove the ones that exist only because nobody designed them away, and rename the system honestly until the core loop closes.