Why 95% of AI projects fail
A 2025 MIT study put a number on it: the vast majority of enterprise generative-AI pilots produce no measurable impact. The problem is almost never the model — it is everything around it.

The number that disturbs
In 2025, MIT researchers studied hundreds of enterprise generative-AI deployments. Their finding: roughly 95% of pilots produce no measurable impact on the income statement. Not "disappointing impact" — no measurable impact at all. Meanwhile, a minority of organizations extract real, recurring value.
This is not a technology problem: the same models are available to everyone. The difference lies elsewhere.
What the study actually says — and what it does not
Before drawing conclusions, the number needs to be read correctly. The MIT study measures the impact of official generative-AI pilots on the income statement — not the usefulness of AI in general. It does not say AI does not work: it says the way companies deploy it produces no measurable value.
One detail in the study is telling: the researchers note that individual use of consumer tools — employees using a chatbot on their own initiative — often produces more value than large official programs. The "shadow AI" outperforms the steering-committee project. Why? Because the employee starts from their real task, iterates fast, and abandons without ceremony what does not work. That is exactly the mindset a good AI project must institutionalize — with governance added.
Another limit worth bearing in mind: "no measurable impact" does not always mean "no impact." Many organizations simply measure nothing — neither before nor after. Their project may be useful, may not be: nobody knows, and that itself is already a management failure.
The four real causes
1. The demo is not the product
A proof of concept impresses in a meeting because it runs on hand-picked examples. In production it meets your real data: incomplete, ambiguous, changing. Without architecture designed for that — reliable retrieval, edge-case handling — it collapses.
2. No memory, no context
The MIT study flags the "learning gap": most tools learn nothing about your processes and forget everything from one session to the next. An assistant that starts from scratch every time never integrates into real work.
3. Zero governance
Who checks the answers? Where does the data go? What happens when the AI is wrong? Without guardrails, traceability, and clear rules, teams — rightly — do not use the tool. And a tool nobody uses has, by definition, a zero return on investment.
4. Starting from the technology, not the business
"We need to do AI" is a starting point that leads to the 95%. Successful projects start from a specific, quantifiable business pain point — then choose the tool.
Useful AI is built — it is not purchased as a kit.
Warning signs of a project heading off the rails
A future member of the 95% is recognizable well before go-live. If you check several of these boxes, it is time to stop and reset:
- The project is called "the AI project" — not "reduce quote-processing time by 30%." When the tool is the subject, the business problem is not.
- Nobody can say what will be measured — no baseline number, no target, no one assigned to watch the dashboard.
- The demo always runs on the same examples — if the tool has never faced an edge case, it is because it cannot handle one.
- End users were never consulted — the project belongs to the executive team or IT, not to the people who will use it every day.
- No clear answers on data — where it goes, who accesses it, what happens when the AI gets it wrong.
- No path to production — the POC lives in a toy environment and nobody has budgeted the integration into the real system.
None of these signals is fatal in isolation. Their accumulation is — and the good news is that they can all be corrected before the budget is spent.
For a mid-market company, concretely ?
The MIT study focuses mainly on large organizations, but the failure mechanism is identical at smaller scale — with a much smaller error budget. A mid-market company that subscribes to three off-the-shelf AI tools without connecting them to its data pays three subscriptions for near-zero gain. Conversely, it is often in mid-size structures that the return is fastest: processes are short, pain points are known to everyone, and decisions are made quickly.
The best first candidates are almost always the same: processing incoming documents (quotes, invoices, files), answers backed by internal documentation, preparation of repetitive deliverables. High-volume, low-ambiguity tasks — exactly where an AI Assessment can quantify the opportunity before a single euro of development is committed.
Anatomy of a project that succeeds
What does a project on the right side of the statistic look like in practice? The details vary, but the pattern is remarkably consistent:
At the start: a pain point, not a technology. The project begins with a sentence like "we spend a huge amount of time re-entering quotes" — not "we need a chatbot." The pain point is quantified (volume, time, error rate), and that number becomes the baseline against which everything is measured.
Then: a deliberately narrow scope. One process, one team, data whose quality has been verified. The temptation to "let everyone benefit" is the number-one enemy: every scope extension before proof multiplies the chances of the project being shelved.
During: real users in the loop from week two. Not an end-of-quarter demo — real cases handled every week, feedback integrated, and the right to say "this doesn't work on this type of file." That is what turns skepticism into adoption.
At the end: an explicit decision. Compare the measurement to the baseline, then decide — scale, adjust, or stop. The 5% do not succeed at every pilot: they know how to stop quickly what does not pay off, and that is precisely what funds the next ones.
How to join the 5%
- Assess before you invest: map the use cases, rank them by impact × feasibility, and scope the first one — that is exactly what an AI Assessment is for.
- Connect AI to your data: answers grounded in your documents (RAG), sources cited, context memory — a foundation, not a gadget.
- Measure before you deploy: an evaluation harness (accuracy, cost, latency) and validated guardrails before going live.
- Keep humans in control: human validation at key checkpoints, full traceability of every action, phased rollout.
It is less spectacular than a demo. It is also what separates an AI project that lives in production from a pilot that gets buried at the end of the quarter. This approach — start from the business, connect the data, measure — is the core principle of the Enterprise AI Core we build.
Frequently asked
Does the 95% figure mean AI does not work in business?
No. The MIT study measures the failure of pilots as they are typically run — generic tools bolted onto processes with no integration and no measurement. Organizations that connect AI to their data and workflows get real, recurring value. It is a methodology problem, not a technology problem.
How long does it take to get measurable value?
It depends on the use case, but the realistic order of magnitude is a few weeks for a well-scoped first perimeter — not eighteen months of platform-building. The key is to start small on a quantifiable business pain point, with a success criterion defined before development begins.
Where to start to avoid joining the 95%?
With an AI Assessment: map the use cases, rank them by impact and feasibility, check the state of your data, and frame a first perimeter with a measurable objective. Investing in a tool before doing that work means choosing the solution before you have defined the problem.
Over 25 years of software craftsmanship. These deep dives come straight from the workshop — the one that builds our SaaS, apps, and AI foundations — not from a marketing department. When a number is quoted, so is its source.
Meet the team