Why do two companies with the same budget, the same data, the same talent pool and the same vendors get completely different results from AI? After 17 years inside transformations, I no longer think the answer is technical.
The companies that get value from AI are not the ones with the most sophisticated algorithms or the biggest data science teams. They refuse to deploy AI until it is tied to a business outcome a CFO would defend in front of the board, and they do the unglamorous work of redesigning the workflow around it. This article sets out what that looks like in practice, including the mistakes I have made and the things vendors leave out of their pitch decks.
Most projects fail for non-technical reasons
Analyst reports put AI project failure rates at 70% to 85%, and that matches what I have seen. The technology usually works. Projects fail because there is no clear value thesis, the data infrastructure is not ready, or the organization is not prepared to use what the AI produces.
In 2019 I worked with a financial services firm that spent $3.4 million on an AI customer service platform, with predictive routing, sentiment analysis and the rest. Eighteen months later it was shut down. The technology worked fine. The service representatives did not trust its recommendations and ignored them, and nobody had invested in change management.
The implementations I have led that succeeded all started from a specific, measurable business problem and worked backwards to the technology. The difference between success and failure was rarely technical sophistication. It was clarity about the problem and willingness to change how the work gets done.
Four applications that delivered
Across financial services, manufacturing, retail and healthcare, four kinds of application have consistently produced measurable returns in my work.
Customer service, with the messy numbers
I helped a mid-market financial services firm put in an AI customer service solution. After six months the results were real but messier than a summary suggests. Response times for routine inquiries fell by 60% to 70%, depending on how transfers are counted. Escalations to human agents fell by 43%, although that took four months to settle. Annual savings came to about $2.1 million, of which $800,000 was hiring we had already planned to avoid. Customer satisfaction rose by 12 points, mostly because of speed, and some customers still asked for "a real person".
The technology was the easier part. The harder part was getting the service representatives to trust it, which took longer than the implementation. What made it work was careful integration with the existing CRM and knowledge base: we spent three months mapping the knowledge architecture before writing any AI code. We also positioned the AI as a tool for the representatives rather than a replacement. They became supervisors of the system, handling complex cases and improving it over time.
Demand forecasting
A manufacturer I worked with raised its forecast accuracy from 72% to 91% with AI-driven demand prediction. Inventory carrying costs fell by 28%, on-time delivery improved by 15%, stockouts on critical SKUs fell by 34%, and $4.7 million of working capital was freed.
The old forecasting model used 8 variables. The AI model used 127, including weather, social media sentiment and economic indicators. No human forecaster could combine that many signals, and that kind of pattern recognition across many variables is where AI is strongest.
Process automation that needs judgment
Basic robotic process automation follows rules: if this field contains X, do Y. Intelligent automation copes with context and ambiguity, which makes it suitable for invoice processing, compliance checks and contract analysis.
One recent project automated preliminary contract review and recovered 340 hours a month of professional time. The AI reads each contract, identifies non-standard clauses, flags risks and writes a summary for a lawyer to review. It does not replace the lawyers. It spares them from reading boilerplate so they can spend the time on judgment.
Sales intelligence and lead scoring
When I put AI-driven lead scoring into a B2B technology company, the sales team's close rate rose by 23%, and the team stopped spending time on leads that were unlikely to buy. The model found that leads who engaged with three specific pieces of content within 48 hours of a webinar converted at eight times the normal rate, a pattern no one on the team would have spotted.
What vendors leave out
After hundreds of vendor presentations, I keep a list of what they consistently leave out.
The first is data engineering. The demo in which the AI answers questions about your business assumes your data is clean, integrated and accessible, and it almost never is. I recommend budgeting 60% to 80% of project effort for data preparation, where vendors usually quote 20%. A retailer I worked with wanted AI inventory optimization, a good use case with clear potential return. Its inventory data sat in 14 systems with inconsistent product codes and no reliable view of real-time stock. We spent eight months on data integration before any AI work could start.
The second is integration. AI has to connect to the CRM, the ERP, the data warehouse and the workflow tools, and every connection can fail. I have seen "12-week implementations" take 18 months because nobody scoped the integration properly.
The third is the gap between pilot and production. A team builds a successful pilot on a subset of data, with hand-curated inputs and close attention. Leadership gets excited. Then the project meets real data, real scale and real edge cases, and it breaks. The pilot worked because smart people were looking after it by hand, and production needs automation, monitoring and governance that nobody built.
The fourth is the cost of keeping it running. Models degrade as the world changes, so they need MLOps infrastructure, drift monitoring and regular retraining. I have watched organizations get strong results in the first year and see performance collapse in the second because nobody budgeted for maintenance.
What changes by industry
In healthcare, privacy rules such as HIPAA add complexity, but the real barrier is physician buy-in. I watched a hospital in Ohio spend 18 months building a clinical decision support tool with excellent accuracy and almost no adoption, because physicians saw it as a threat to their judgment and had not been involved in designing it. A different health system asked three senior physicians to choose the problem. They picked something unglamorous, prior authorization for imaging studies, where the AI handled paperwork rather than clinical decisions, and the physicians loved it. In healthcare, AI that removes administrative work wins, and AI that touches clinical judgment loses until physicians help design it.
In manufacturing, the convergence of operational and information technology is underestimated. Factory systems often run on decades-old protocols that do not talk to modern AI infrastructure, latency-critical applications often need edge computing, and change on the shop floor needs different methods from change in an office.
In financial services, regulators want to understand how a model reaches its decisions, and "the AI said so" is not an answer. A bank I worked with built a credit scoring model that beat its legacy system by 23% on default prediction. When the regulator asked for model documentation, the data science team could not explain why the model worked and spent eight months reverse-engineering its own model. I now recommend budgeting 30% to 40% of model development time for explainability and documentation, partly because regulators require it and partly because you cannot improve a model you cannot explain.
In retail, the tension between personalization and privacy is growing. The applications that drive value, such as personalized recommendations, dynamic pricing and targeted marketing, need exactly the kind of data collection consumers are increasingly uncomfortable with, and GDPR and US state privacy laws are narrowing what is possible.
The red flags I watch for
A vendor promising production AI in eight weeks is either overpromising or building something too simple to matter; real implementations take four to twelve months depending on complexity. A goal such as "use AI to improve customer experience" is not a goal, while "reduce average handling time by 30% while keeping CSAT above 4.2" is one. An organization with no data catalog, no data quality process and no master data management should build that foundation before anything else. And a company that asks "we bought this AI platform, what should we use it for?" has the question backwards.
Build, buy or partner
Where the capability comes from is one of the most consequential decisions in any AI initiative. Build when the capability is central to how you compete, when you have strong in-house data science talent, when the use case needs deep integration with proprietary systems or when you need full control of the roadmap. Buy when the problem is well understood and mature products exist, when speed matters, when you lack specialized expertise and when the vendor will keep improving the product. Partner when you need to build capability while delivering results, when you lack the domain expertise, when you want to keep both options open, or when the harder problem is organizational change.
Most organizations end up with a mix: they buy platforms for commodity capabilities, build differentiated applications and use partners for early implementations where they lack experience.
Four questions before any AI project
Before I recommend an AI project, it has to pass four questions.
- Can we state the business problem in one sentence? If it takes paragraphs, we are not ready.
- Is the data ready? I use a 20-point assessment covering everything from completeness to access latency.
- Can we estimate the financial impact to within 25% either way, from a bottom-up business case rather than a claim that it will transform the business?
- Is the organization ready to work differently, with an executive sponsor, capacity for change management and willingness to iterate?
If any answer is no, the project waits. I have disappointed executives by recommending against AI projects they were excited about, and it is better to disappoint early than to fail expensively.
I have turned this into an interactive tool that takes about ten minutes and produces a prioritized list of use cases for your context. It uses the same method I have applied in more than 40 enterprise engagements. It will not replace strategic thinking, but it will give the conversation a structure.
The strongest objection
A reasonable critic would say two things. First, my examples come from my own engagements, which are selected and hard to verify, and the famous failure rates are themselves contested. Second, AI is now cheap enough to test that a four-question gate slows things down; teams should run many quick experiments and let the results decide.
The first point is fair, which is why I give the messy numbers and the failures along with the successes, and why every business case should be measured against its own baseline rather than against anyone's anecdotes, including mine. On the second, I agree that cheap experiments are the right way to learn. The four questions apply at the point where an experiment asks for production budget, integration work and people's time. That is where the expensive failures I have described happened, and none of them failed because the experiment was too slow.
Why this holds as models improve
In July 2024, OpenAI shared an internal framework for measuring progress toward more general AI in five levels. Level 1 is chatbots, which converse and generate text but only react to what they are asked. Level 2 is reasoners. OpenAI's o1, released in September 2024, can work through multi-step problems the way an expert would. A chatbot might tell you a contract clause is unusual, while a reasoner can explain why it is a problem, what follows from it and how it interacts with three other clauses you did not ask about. Level 3 is agents, which use tools and carry out multi-step work with limited supervision. Levels 4 (innovators) and 5 (organizations) remain theoretical.
I misjudged Level 1. In early 2023 I told a client that chatbots changed everything. They changed some things and made certain tasks faster, but they did not change how organizations operate, because they could not plan or act. Level 2 made a bigger impression on the people I work with. I watched a legal team in Zurich go from sceptical to convinced in about 90 minutes as o1 worked through a complex regulatory question. "It thinks like we think," the senior partner said. "Just faster."
Most enterprises are still using AI at Level 1, while the technology has reached Level 2 and early agents are appearing. OpenAI's o3, announced in December 2024, continues that direction, and the gap between an assistant that answers questions and a colleague that handles projects is narrowing. But each level raises the stakes of the same fundamentals. An agent that acts on bad data, without a clear owner or a defined outcome, fails faster and more expensively than a chatbot. The organizations that will benefit from later levels are the ones building the data infrastructure, governance and working habits now.
Monday move
Take the AI initiative with the largest budget in your organization and put the four questions to its owner: the problem in one sentence, the state of the data, the financial impact within 25%, and who will change how they work. If any answer is weak, fix that before the next funding decision rather than after the next status report.