If you have read anything about AI in business this year, you have met the statistic: most pilots fail. Founders and operations heads in Bangalore, Pune and Gurugram quote it in board meetings, usually just before asking for a bigger budget or a smaller one. Understanding why AI projects fail is more useful than quoting the number, because the reasons are specific, and mostly fixable.
Here is what the two best-known sources actually said, what they measured, where the criticism is fair, and what the pilots that do pay back tend to share.
What the MIT report actually says
In July 2025, a team at MIT’s Project NANDA published The GenAI Divide: State of AI in Business 2025. Its executive summary says that despite $30 to 40 billion of enterprise investment, “95% of organizations are getting zero return”, and that “just 5% of integrated AI pilots are extracting millions in value”.
Two details are worth knowing before you repeat that line.
The evidence is smaller than the headline. The report’s own methodology note lists a review of over 300 publicly disclosed AI initiatives, interviews with 52 organisations, and survey responses from 153 senior leaders. Some news coverage at the time quoted larger figures (150 interviews and 350 employees), but those numbers are not in the report itself.
Success was defined narrowly. The authors defined it as deployment beyond the pilot phase with measurable KPIs, with ROI measured six months after the pilot. In the exhibit on custom and task-specific tools, “successfully implemented” meant users or executives remarked on a “marked and sustained productivity and/or P&L impact”. A pilot that saved a team ten hours a week but never showed up in the P&L would not have counted.
The authors are also candid about limits. They write that the figures are “directionally accurate based on individual interviews rather than official company reporting”, that sample sizes vary and that success definitions differ between organisations. The report is labelled “Preliminary Findings” and was not peer reviewed.
Where the criticism is fair
Critics have made three reasonable points. The first is that the 95% figure is hard to trace to a clear dataset in the report. The second is that a six-month window can understate the value of slower enterprise rollouts; the report itself does not dispute this. The third is that a small, self-selected group of interviewees may differ from the wider market. The authors list selection bias as a possible problem, too.
The sensible reading is this. Do not treat 95% as a measured failure rate for Indian businesses. Do treat it as a credible pattern: plenty of pilots impress in a demo and never change how a team works. That pattern matches what most operations leaders will tell you over a cup of chai.
What Gartner measured, and what it didn’t
In June 2025, Gartner issued a press release predicting that over 40% of agentic AI projects will be cancelled by the end of 2027, “due to escalating costs, unclear business value or inadequate risk controls”.
This is a forecast, not a count of failures. The only data in the release is a January 2025 poll of 3,412 webinar attendees about how much they had invested in agentic AI, not how it turned out. Gartner’s analyst Anushree Verma said that most agentic AI projects right now are “early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied”.
The release also warns about “agent washing”, where vendors rebrand chatbots, assistants and RPA as agents. Gartner estimates only about 130 of the thousands of agentic AI vendors are real. If you are comparing vendors this quarter, that is a more actionable sentence than any percentage. We cover the vocabulary in our guide to chatbots, agents and agentic applications.
Why AI projects fail: five patterns that repeat
Put the two sources together with what we see when teams come to us after a stalled pilot, and the causes are rarely the model.
- Nobody owns the result. The pilot sits with IT or an enthusiastic intern. No one in the business is accountable for hours saved or errors avoided.
- The workflow was bolted on, not redesigned. An AI step is added to a process that was already clumsy. If your team still re-keys the output into Tally or Zoho by hand, the saving evaporates.
- The tool doesn’t learn. MIT calls this the learning gap: tools that “don’t learn, integrate poorly, or match workflows”. A tool that makes the same mistake every Monday gets abandoned by Wednesday.
- The first use case was chosen for visibility, not value. The report notes that about half of generative AI budgets go to sales and marketing, while some of the clearest savings it documented came from back-office work.
- There is no definition of success. Without a baseline, “it feels faster” is the only metric, and it doesn’t survive a budget review.
Human in the loop: choosing the process worth automating, and the number that will prove it worked, is a senior person’s decision. A model can draft the workflow. It cannot know which step your finance head quietly distrusts.
What the pilots that pay back have in common
MIT’s own findings on organisations that “cross the divide” are more human than technical. A few that stand out:
- They start from the front line. The report describes successful organisations sourcing ideas from frontline managers rather than central labs, and keeping executive accountability for outcomes.
- They expect the tool to be customised. In the report’s interview sample, external partnerships with customised tools reached deployment about 67% of the time, against about 33% for internal builds. Treat that with care: it is self-reported, it comes from the 52 organisations interviewed, and the authors say the gap may reflect differences in the organisations rather than the approach alone.
- They begin with the back office. Where the report documents clear savings, they come mostly from reduced outsourcing and agency spend, not from firing people.
Add what we see in practice, and a pattern emerges. Pilots that pay back are narrow, owned, measured and supervised.
- Narrow. One task, such as matching UPI settlements to invoices or drafting follow-ups for WhatsApp enquiries, not “AI for the whole company”.
- Owned. One named person who is judged on the outcome and can say stop.
- Measured. A baseline taken before the pilot starts: hours per week, days to respond, errors per hundred invoices.
- Supervised. A person approves outputs until the numbers show the agent has earned more freedom.
Human in the loop: in every build we run, the person who signs off the agent’s work is named before a line of the workflow is written. If no one will put their name to it, that is the answer to whether it is ready.
The Indian context: smaller details that sink pilots
Some failure points are particular to how Indian teams work.
- Messy inputs. Invoices arrive as photos on WhatsApp, GST data is spread across Tally exports, and customers write in Hinglish. An agent trained on tidy English PDFs will struggle. Test with last month’s real files, not the vendor’s sample set.
- Mobile-first reality. Your approvers are often on mid-range Android phones and patchy 4G. If the approval screen needs a laptop, approvals will queue up and the pilot will look slow.
- Data protection. Under the Digital Personal Data Protection Act, 2023, personal data in customer chats and CRMs needs a lawful basis and care in handling. Decide what the agent may see before the pilot, and have your counsel confirm the position for your case.
- Cost honesty. The rupee cost of a pilot includes the review time of your own people, not just the licence. Budget for it, or the pilot will quietly consume your best operations person’s evenings.
How we approach it at Cannyworx
We are a design and AI studio, and the pattern above shapes how we work. AI does the heavy lifting on drafting, matching and first passes. A senior person on our team scopes the workflow with yours, designs the approval screens, and reviews the output before it goes anywhere that matters. We would rather ship one narrow agentic application your team trusts than a broad one it works around.
Where to start
Before talking to any vendor or studio, write down three things on one page: the one repetitive task you would most like off your team’s plate, the person who will own it, and the number that would tell you it worked after six weeks. If you can’t fill in all three, you have found your first piece of work, and it needs no software.
If you would like a structured way to check whether your business is ready for an AI agent, the free readiness scorecard below walks through it in about ten minutes.
Questions people ask
Do 95% of AI pilots really fail?
Not in the way the headline suggests. The MIT NANDA report's 95% figure comes from 52 interviews, 153 survey responses and a review of public projects, and the authors call it directionally accurate. It is a useful warning about stalled pilots, not a precise failure rate for every business.
What is the most common reason AI projects fail?
Usually the work around the model rather than the model itself: no named owner, a workflow that was never redesigned, and tools that don't fit how the team really works. MIT's authors describe this as a 'learning gap' between the tool and the organisation.
How long should an AI pilot run before we decide?
Long enough to see real work flowing through it, typically a few weeks of daily use with a person reviewing every output. Decide in advance what number would count as success, such as hours saved or turnaround time, so the call isn't made on enthusiasm.





