A CIO in Zurich opens a deck with forty agent tiles: an HR policy bot, a meeting summarizer, a procurement email drafter, and forty more in flight. The pitch is one sentence: "We are scaling agentic AI across the enterprise."

None of it is agentic. These are forty chatbots attached to forty inboxes. None has a goal, none can decide what to do next, and none does anything until a person types first. The deck looks like progress, and the P&L will not move.

I have seen this configuration in industrials, financial services, pharma and consumer goods, in companies with very different cultures and budgets. Through the Pipesignal autonomous pipeline and the first agentic AI deployment in construction at Holcim, I have also seen the minority that works differently. In my experience, the gap between the two groups comes less from budget, talent or model quality than from three repeated mistakes, and one mental model fixes them. They help explain why MIT NANDA's 2025 study found that about 95% of generative AI pilots return no measurable P&L impact.

Reading the 95% correctly

When the MIT NANDA figure landed, many boards concluded they should slow down. I think that reading is backwards. The figure says more about how the pilots were framed than about the technology.

Look at what failed pilots tend to share: a chatbot added to an FAQ page, a copilot rolled out across a department with no defined outcome, a centre of excellence producing presentations about centres of excellence. None answers a business question that matters. They answer the question of how to look as if the company is doing AI.

The deployments that work start with a problem the organization has been losing money on for years and ask whether abundant, autonomous intelligence would change the shape of that problem. If it would, they redesign the workflow around an agent. If it would not, they do not deploy AI. The difference is the question asked at the start.

Mistake 1: calling chatbots agentic

This is the most common error, and it is dangerous because it makes leaders believe they are further along than they are.

A chatbot answers when spoken to. It has no goal of its own, does not know whether the conversation succeeded, and cannot decide to do something different tomorrow because today did not work. Without the marketing language, it is a fluent search interface.

An agent has a goal. It plans, acts on the world, observes the result, updates and tries again. That closed loop of perceiving, deciding, acting and learning is what agency means, and without it you have a conversation.

I wrote about this in The Orchestration Era, and Y Combinator's thesis around AI-fulfilled agencies makes the contrast sharper. It is funding companies where the AI delivers the outcome end to end, such as brand identity, ad creative, media buying and client outreach, while people design the workflow.

The distinction matters in a boardroom because a chatbot called "agentic" gets measured against the wrong baseline. Leaders expect it to compound and budget for it to scale, and it will not. The more useful statement is plain: we have deployed assistants, they reduce friction in specific tasks, and renaming them will not make them agents. I set out a simple test for this in It's Not an Agent If You Still Press the Button.

Mistake 2: supervising AI like a junior intern

The second mistake is reflexive supervision: every output reviewed, every action approved and every decision double-checked by a person, even when that person has less context, recall and time than the model.

The reflex is understandable. Trust has to be earned, liability is real, and the first embarrassing answer on a customer call can cost someone their job. So the default becomes "human in the loop, always".

That default has costs. Harvard Business Review's September 2025 article on "workslop" described AI output that looks finished while pushing cognitive work back onto the reviewer. Research on what has been called the augmentation trap suggests that people who offload thinking to AI and then over-supervise the result can lose the expertise the productivity gain depends on.

Productivity rises steeply with the first layer of oversight, peaks early and then flattens into a long tail. Past the peak, each extra review meeting or approval gate taxes the gain. I go into the mechanics in Human-in-the-Loop Is the New Technical Debt.

The cost is higher in 2026 because capability is rising fast. METR's January 2026 update fitted a doubling time of about three months for the length of tasks frontier models can complete. Frontier systems reached gold-medal level at the International Mathematical Olympiad, AlphaFold predicted structures for around 200 million proteins, and there are documented cases of language models catching errors in scientific papers that survived peer review. Meanwhile, many organizations spend their AI budget having people re-read AI-generated emails.

People remain essential where judgment, accountability and values matter. The shift is to stop assuming the person is always more right than the model, learn by domain which is which, and design the workflow accordingly.

Mistake 3: starting with AI instead of the problem

This is the deepest mistake and the hardest to see, because so much advice reinforces it. The CEO reads about AI, a taskforce forms, capabilities are inventoried and mapped to functions, someone asks where AI can be used, a heatmap appears and the highest-scoring cell is piloted. Eighteen months later the pilot joins the 95%.

The error is in the first question. "Where can we use AI?" puts the technology in the driver's seat and treats the business as a list of slots to fill.

A harder and better question is: which of our hardest problems would change shape if intelligent attention were abundant, autonomous and available on demand? It asks what becomes possible when a constraint the organization has lived with for decades, the cost of intelligent attention, falls sharply. Much of any strategy was built on the assumption that intelligence is scarce, expensive and slow.

When I run this exercise with executive teams, the room often goes quiet before the real answers arrive. Someone says the company has avoided solving customer onboarding for years because it needed a person to read every contract. Someone else says it never built a real account-based sales motion because it could not afford that much research per account. A third says the spare-parts catalogue project was deprioritized three times because the manual clean-up cost too much. Those were the use cases all along, invisible because the team had stopped seeing them as solvable.

I covered this in detail in How to Find AI Use Cases That Make Money, and the Use Case Prioritization tool makes it operational.

The reframe: intelligence as a service

The mental model I use with executive teams assumes three properties.

  1. Abundant: intelligence is no longer rationed by headcount, so you no longer choose which contracts to read carefully, which leads to research or which incidents to analyse. The question becomes what you are still pretending you cannot afford to examine.
  2. Autonomous: it can start, run and finish without a person initiating each instance, triggered by events, schedules and goals. This is what separates an agent from a chatbot.
  3. Capable: it can read a contract, interpret a regulation, generalize from one customer to a thousand and explain its reasoning. The output is less predictable in form than a script's, and it can still be supervised in substance.

With those assumptions, the strategy conversation moves from "which process should we automate?" to "what would we attempt if this kind of intelligence were available throughout the business?", and the two questions rarely produce the same backlog. This is why the AI-native operating model work matters more than the tooling: the tools change every few months, and the operating model decides whether they pay off.

The strongest objection

A CIO could fairly push back that their assistants and copilots are delivering real value, that employees use them, and that asking about the hardest problems invites moonshots that never ship.

The first point is right. Assistants remove real friction and are worth having, provided they are measured as assistants. The second is a risk worth managing. Problem-first does not mean picking the largest problem in the company. It means picking a painful, recurring problem with a workflow around it and an outcome that can be measured within a quarter or two, and then redesigning that workflow properly.

Four steps to start

  1. Inventory your hardest problems rather than your processes. Ask the top team which five problems have cost the most over the last decade and remain unsolved.
  2. Score each by what changes if intelligent attention is abundant. Keep the problems whose shape would change.
  3. Pick a workflow rather than a task. A workflow is a sequence with a business outcome at the end. If you cannot describe the outcome in one sentence, you are looking at a task.
  4. Design for delegation rather than supervision. Decide by domain where people add judgment and where they only add friction. Put people where they add judgment, and put guardrails, observability and rollback everywhere else. The Change Resistance Review helps show where the organization may resist before you find out the expensive way.

Monday move

Put one question on the agenda of your next leadership meeting: which problem have we avoided for years because solving it required too much skilled attention? Take the first answer that has a clear owner and a measurable outcome, and ask that owner to describe, in one page, what the workflow would look like if an agent did the reading, research or clean-up.