Why Do AI Rollouts Stall? What the Research Actually Says
AI rollouts stall because the difficult part of an AI programme is organisational rather than technical, while the budget and the attention usually go the other way. In the largest available studies, the factors that separate the organisations getting value from the ones that are not are strategy, workflow redesign, leadership behaviour and how much training people actually received, not model quality or tool choice. The most cited statistic on the subject, that 95 percent of AI pilots fail, comes from a non-peer-reviewed preliminary report whose own pages report a roughly 83 percent success rate for a different class of tool.
Published 5 September 2026. Last updated 5 September 2026, the date on which every source cited below was fetched and read at source.
Why this page is shaped as an audit
The market for AI adoption advice runs on a small number of statistics that get copied from article to article without anyone opening the underlying report. Several of them do not survive contact with the source document. This page quotes each figure from the report it came from, names the sample size and the method, and says plainly where a number is weaker than its reputation. Where a widely repeated figure could not be substantiated, it is not printed here at all.
Why do AI rollouts stall?
The short answer is that the licence is the cheap part. Buying an assistant for everyone is a procurement decision that can be executed in an afternoon. Changing how a team writes a quote, checks a claim, hands off a case or closes a month is a management problem that takes quarters, and it is the part that produces the number the board asked for. Programmes stall in the gap between those two things.
That is not an opinion drawn from experience, and it does not need to be. Four independent bodies of evidence point at the same place: a consultancy survey of over ten thousand employees, a vendor survey of twenty thousand knowledge workers, a nationally representative academic study of forty-eight thousand people, and two field experiments with random assignment. They disagree about plenty. They agree that the binding constraint sits with people, workflows and management, and that it is not the model.
It is worth separating two failures that get filed under the same heading. One is that nobody uses the thing. The other is that everybody uses the thing and nothing measurable changes. They have different causes and different fixes, and the second is far more common than the first.
Adoption is not the same as value
Usage figures have risen sharply and keep rising. BCG’s 2026 wave, covering close to 12,000 frontline employees, managers and leaders across more than a dozen markets, found 74 percent of frontline employees describing themselves as regular AI users, while only 36 percent reported receiving adequate upskilling. On the other side of the ledger, Deloitte’s survey of 2,770 director-to-C-suite respondents across 14 countries found that 68 percent said their organisation had moved 30 percent or fewer of its generative AI experiments fully into production, that only 38 percent were tracking changes in employee productivity, and that just 16 percent had produced regular reports for the CFO about value created.
Read together, those tell a specific story. Individual use is close to universal. Organisational change is rare, and organisational measurement is rarer still. A programme where nobody can say what changed is not usually a programme that failed to launch. It is a programme that launched into a workflow nobody redesigned and a measurement system nobody built.
BCG’s 2024 study of 1,000 senior executives across 59 countries makes the same point from the capability side: of the 98 percent of companies at least experimenting with AI, only 26 percent had, in its words, “developed the necessary capabilities to move beyond proofs of concept and begin extracting value”. That 26 percent is a derived figure, the sum of a 4 percent group at the leading edge and a 22 percent group starting to generate value, so treat it as a segmentation rather than a survey answer. The shape of the finding is what matters, and it recurs everywhere.
Is it true that 95 percent of AI pilots fail?
No, and the way that number is used is worth taking apart, because almost every page in this category repeats it and almost none of them have read the document.
The source is The GenAI Divide: State of AI in Business 2025, by Aditya Challapally, Chris Pease, Ramesh Raskar and Pradyumna Chari, dated July 2025. Its own notes page describes it as “Preliminary Findings from AI Implementation Research from Project NANDA” and carries the disclaimer that “the views expressed in this report are solely those of the authors and reviewers and do not reflect the positions of any affiliated employers”. It is not peer reviewed, it is not an institutional position, and it is not published on MIT’s own site. Every accessible copy is a third-party mirror, including the one linked here.
Its methodology, quoted from the same page, is “a systematic review of over 300 publicly disclosed AI initiatives, structured interviews with representatives from 52 organizations, and survey responses from 153 senior leaders collected across four major industry conferences”. Fifty-two organisations is a qualitative sample. It can generate hypotheses. It cannot support a percentage quoted to the whole economy.
The methodology numbers in circulation are wrong
The coverage that made this report famous restated its method as “150 interviews with leaders, a survey of 350 employees, and an analysis of 300 public AI deployments”. The report says 52 organisations and 153 senior leaders. Those numbers were then inherited by hundreds of downstream articles without correction. If you are going to cite the methodology of this report, cite the report.
Then there is what the figure actually measures. The 95 percent is the residual of a funnel drawn for one category only, described in the report as embedded or task-specific generative AI tools, of which roughly 5 percent reach production. And the report defines its own success criterion in a research note: “We define successfully implemented for task-specific GenAI tools as ones users or executives have remarked as causing a marked and sustained productivity and/or P&L impact.” So a pilot counts as failed when no executive in a 52-organisation interview sample remarked on a marked and sustained impact. That is a perception measure, not an audited return.
The decisive detail is in the same section. Two paragraphs after the 95 percent, the report states: “Generic LLM chatbots appear to show high pilot-to-implementation rates (~83%).” One report, two tool classes, two very different outcomes. The headline took the worse one and dropped the qualifier. The assistant sitting on an ordinary desk is a general-purpose chatbot, which is the class the same report scores at 83 percent.
None of that makes the report worthless. Its actual conclusion is the useful part, and it agrees with everything else on this page: “The core barrier to scaling is not infrastructure, regulation, or talent. It is learning.” That sentence deserved the attention. The percentage got it instead.
Where does the effort actually need to go?
The most durable answer in the literature is BCG’s allocation rule, and it is worth quoting in full rather than in the inverted form that circulates: “Our experience, corroborated by our new research, indicates that about 70% of the challenges relate to people and process, about 20% are technology issues, and only 10% involve AI algorithms (which often occupy a lot more organizational time and resources).” The same report turns that into a prescription: focus 70 percent of effort and resources on people-related capabilities, 20 percent on technology, and 10 percent on algorithms.
Two years later BCG’s 2026 wave, covering close to 12,000 frontline employees, managers and leaders across more than a dozen markets, puts the comparison in one line: employees with clear strategy but limited access to tools outperform those with strong access but no direction. BCG frames the whole 2026 analysis around why strategy matters more than tools. Note what that statement is and is not. It is a directional finding about which of two constraints binds harder. We have deliberately not attached a percentage to it, because no quantified strategy-versus-tools figure appears on BCG’s own page for this study, and several figures in circulation that purport to quantify it could not be found at source.
Microsoft’s 2026 Work Trend Index, a survey of 20,000 knowledge workers across ten markets fielded between 18 February and 7 April 2026, decomposes the same effect: “Organizational factors like culture, manager support, and talent practices account for more than 2x the reported AI impact of individual factors like mindset and behavior (67% vs. 32%).” In the same study only 26 percent of AI users say leadership is clearly and consistently aligned on AI.
Note what that ratio implies for a budget. If two thirds of the effect is organisational, a plan whose only line items are licences and a lunchtime demo has bought the smaller third and skipped the larger one. This is the same reasoning we apply when scoping a deployment rather than a course, which is why the training we run is attached to a system going live rather than sold as a standalone curriculum. The shape of that is set out on our AI training for companies page.
Does training change anything measurable?
Yes, and the dose-response is the cleanest thing in the whole evidence base. BCG’s 2025 survey of 10,635 employees across 11 markets plotted regular AI usage against hours of training received. The bars run from 18 percent regular users among those who received no training at all, to 63 percent at one to five hours, 82 percent at five to ten hours, and 89 percent above ten hours. The report’s own summary of the chart is that 79 percent of respondents who received more than five hours of training are regular AI users, compared with 67 percent of those who received less.
Format matters too, and by less than people assume. The same survey shows in-person training associated with a 12 percentage point lift over remote, and access to a coach with a 14 percentage point lift over none. Both are real and both are small next to the gap between no training and any training. The first hour is worth more than the venue.
Against that, only 36 percent of employees in the same survey say their training was enough. In BCG’s 2026 wave the figure was unchanged at 36 percent, a year later, with usage 20 points higher. The gap is not closing on its own.
Read the direction of causation carefully
These are survey correlations, not experiments. Organisations that invest ten hours per person in training are different from organisations that invest none, in ways that also affect usage. The honest reading is that training hours are the clearest available marker of a serious programme, and that the marker moves usage across an unusually wide range. For causal evidence about what AI does to output quality, the two field experiments in the sections below are the better source.
What has the single largest effect?
In BCG’s 2025 data the largest single differential is not training and not tooling. Among 3,537 frontline employees, 82 percent of those with clear leadership support on generative AI use were regular users, against 41 percent of those without it, a gap of 41 percentage points. The same section reports that only 25 percent of frontline employees say they experience that support. Positive sentiment about the impact on job enjoyment moves from 15 percent to 55 percent on the same split.
Leadership support in this data is not a memo. It is whether people believe they are allowed, whether their manager visibly uses the tools, and whether anyone has told them which tasks are appropriate. That last one is cheap and is skipped constantly. Its absence has a second cost, because staff who lack a sanctioned route do not stop, they route around: BCG’s 2025 survey found 37 percent say their company is not supplying the right tools, and that when corporate solutions fall short, 54 percent say they would use unauthorised AI tools. What that produces is covered in our companion piece on what employees actually paste into AI tools.
Who gains most from the tools?
Here the evidence gets stronger, because it stops being survey data. Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied a staggered rollout of a generative AI assistant across 5,179 customer support agents, published as NBER Working Paper 31161 and subsequently in the Quarterly Journal of Economics. Their finding, verbatim: “Access to the tool increases productivity, as measured by issues resolved per hour, by 14% on average, including a 34% improvement for novice and low-skilled workers but with minimal impact on experienced and highly skilled workers.”
The distribution matters more than the average. The gain concentrates on the least experienced staff, and the mechanism the authors describe is that the system disseminates the working practices of the more able workers to everyone else. This is a training effect delivered through a tool. It is also the reason a rollout that begins with the most senior people, who need it least, tends to produce a disappointing pilot and an unfairly negative verdict.
Why usage-only training makes things worse
The most important result for anyone designing a programme is the one that cuts against the sales pitch. Fabrizio Dell’Acqua and colleagues ran a pre-registered field experiment with 758 BCG consultants randomly assigned to no AI, AI, or AI plus a short prompt-engineering overview. On 18 realistic consulting tasks inside the capability frontier, the AI groups “completed 12.2% more tasks on average, and completed tasks 25.1% more quickly”, and produced “more than 40% higher quality” results, with the largest gains going to those below the average performance threshold, who improved by 43 percent against 17 percent for those above it.
Then the same paper reports the other half: “For a task selected to be outside the frontier, however, consultants using AI were 19 percentage points less likely to produce correct solutions compared to those without AI.”
That is percentage points, not percent, and the distinction is widely botched. It means the same tool, in the same hands, on a task that merely looks similar, converted a large gain into a measurable loss. The tasks inside and outside the frontier were chosen to be comparable in apparent difficulty. Nothing on the screen tells you which one you are holding.
Follow that through and the conclusion for training is uncomfortable. A programme whose success metric is usage, seats activated, prompts per week, licences consumed, is optimising for the exact behaviour that produces the loss on the second class of task. Driving usage without teaching judgement moves work from the 12 percent column to the minus 19 point column and reports it as adoption. The competence that has to be taught is the boring one: how to tell whether this task suits the tool, and how to check the answer before it leaves the building. That competence is also what the EU expects, and the regulator has said as much in plain language, which we cover in what EU AI Act Article 4 actually requires.
The nationally representative evidence says this is where the real exposure sits. In the University of Melbourne and KPMG global study of 48,340 people across 47 countries (University of Melbourne & KPMG, 2025 (n=48,340)), 66 percent of employees report having relied on AI output at work without critically evaluating the information it provides, 72 percent report putting less effort into their work because of AI, and 56 percent report having made mistakes in their work from AI use. The report’s own diagnosis is that this complacent use “may be fueled by inadequate training, guidance, and governance of responsible AI use at work”.
Five claims, and what the sources say
| The claim as usually repeated | What the source actually says |
|---|---|
| 95 percent of AI pilots failSource: MIT NANDA report, pp. 2 and 6 | A preliminary, non-peer-reviewed report, 52 organisations interviewed. The figure is a funnel residual for task-specific tools only, and the same report gives roughly 83 percent for general-purpose chatbots |
| Most value comes from picking the right modelSource: BCG, Where’s the Value in AI? | About 10 percent of the challenge is algorithms, 20 percent technology and data, and 70 percent people and process |
| Training is a nice-to-haveSource: BCG, AI at Work 2025 | Regular usage runs from 18 percent with no training to 89 percent above ten hours, and only 36 percent say their training was enough |
| AI makes everyone more accurateSource: Dell’Acqua et al., HBS WP 24-013 | Inside the capability frontier, yes: 12.2 percent more tasks and 25.1 percent faster. Outside it, 19 percentage points less likely to be correct |
| Adoption is an individual mindset problemSource: Microsoft, 2026 Work Trend Index | Organisational factors account for 67 percent of reported AI impact against 32 percent for individual mindset and behaviour |
What the evidence says to do instead
Nothing here argues against deploying AI. It argues that the deployment decision is a small part of the work, and that the order matters. Read as a sequence, the studies above point at five steps that are unglamorous and cheap relative to the licences.
Pick the processes before the tools
BCG’s 2026 analysis of close to 12,000 employees finds that those with a clear strategy but limited access to tools outperform those with strong access but no direction. Name two or three processes where an improvement would actually show up in a number somebody already reports, and start there rather than issuing licences broadly and waiting to see what happens.
Source: BCG, AI at Work 2026Budget in the 70/20/10 proportion, not the reverse
If about 70 percent of the difficulty is people and process, a plan whose entire cost line is software has funded the easy 10 to 20 percent. Put the majority of the effort into redesigning how the work runs and into the people doing it.
Source: BCG, Where’s the Value in AI?, 2024Give more than five hours, and start with the least experienced
The usage curve steepens sharply above five hours and again above ten. The productivity gain in the strongest causal study concentrates on novice and low-skilled workers, at 34 percent against minimal impact for the highly skilled. Starting a pilot with your most senior people inverts the effect you are trying to measure.
Source: BCG, AI at Work 2025 (n=10,635)Teach the frontier, not the prompt
The teachable skill is telling apart a task the tool handles well from one that merely resembles it, and checking the output before it goes anywhere. In the consultant experiment, a short prompt-engineering overview did not protect the group working outside the frontier.
Source: Dell’Acqua et al., HBS WP 24-013Make managers visible users, and say what is allowed
The 41 percentage point gap between employees with and without clear leadership support is the largest single differential in the data, and only a quarter of frontline employees report experiencing it. Naming the sanctioned tools and the permitted tasks costs nothing and removes the main reason people improvise.
Source: BCG, AI at Work 2025 (n=10,635)A last note on measurement, since it is the step that gets skipped. If only 16 percent of the organisations Deloitte surveyed produce a regular report to the CFO on value created, the other 84 percent have no mechanism that could tell them whether the thing worked. Decide the number before the rollout, take the baseline before the tool arrives, and accept that a genuine before-and-after on two processes is more informative than a dashboard covering twenty. The same discipline applies to any system going live, which is why our page on what an implementation actually involves spends more time on the weeks after launch than on the launch.
If you are earlier than that and still deciding where to point the first project, how to start using AI in a small business and how to automate recurring business processes cover the selection question, and ten practical business uses of ChatGPT is the concrete version for a team that has licences and no plan. If the tools are already in use and nobody has decided which ones are sanctioned, start with what each vendor does with your data instead, because that decision constrains everything downstream. Lithuanian readers will find the same argument, with the Lithuanian adoption data, in kodėl DI diegimas įmonėje strigo.
The summary is short. Rollouts stall where nobody redesigned the work, nobody said out loud what was allowed, and nobody was taught to tell a suitable task from an unsuitable one. All three are management problems, all three are cheap next to the software, and all three are what the research keeps finding when it goes looking for the difference between the organisations getting a return and the ones that are not.
Frequently Asked Questions
The figure comes from a preliminary, non-peer-reviewed report from a research group, based on a systematic review of over 300 publicly disclosed initiatives, structured interviews with 52 organisations and survey responses from 153 senior leaders. The 95 percent is a funnel residual for task-specific generative AI tools only, and the success criterion is whether users or executives remarked on a marked and sustained impact. The same report separately states that generic LLM chatbots appear to show high pilot-to-implementation rates of around 83 percent. The widely quoted methodology, 150 interviews and a survey of 350 employees, is not what the report says.Source: MIT NANDA, State of AI in Business 2025 (mirror)
BCG puts it at about 30 percent in total. Its research finds that about 70 percent of the challenges relate to people and process, about 20 percent are technology issues, and only 10 percent involve AI algorithms, and it recommends allocating effort and resources in the same proportion. Its 2026 analysis of close to 12,000 employees adds a directional finding in the same direction: those with a clear strategy but limited access to tools outperform those with strong access but no direction.Source: BCG, Where’s the Value in AI?, 2024
The clearest available dose-response comes from BCG’s 2025 survey of 10,635 employees. Regular AI usage was 18 percent among those who received no training, 63 percent at one to five hours, 82 percent at five to ten hours and 89 percent above ten hours. In-person delivery was associated with a 12 percentage point lift and access to a coach with a 14 percentage point lift. These are correlations from a survey rather than experimental results, so read them as a marker of a serious programme rather than a guaranteed return per hour.Source: BCG, AI at Work 2025 (n=10,635)
It depends entirely on whether the task falls inside the tool’s capability frontier. In a pre-registered experiment with 758 consultants, those using AI on 18 tasks inside the frontier completed 12.2 percent more tasks, worked 25.1 percent more quickly and produced more than 40 percent higher quality output. On a task selected to be outside the frontier, the same participants were 19 percentage points less likely to produce a correct solution than those working without AI. The figure is percentage points, not percent.Source: Dell’Acqua et al., HBS WP 24-013
The least experienced staff. A study of 5,179 customer support agents on a staggered rollout found that access to the tool increased issues resolved per hour by 14 percent on average, including a 34 percent improvement for novice and low-skilled workers, with minimal impact on experienced and highly skilled workers. The authors describe the mechanism as the system disseminating the practices of more able workers.Source: Brynjolfsson, Li & Raymond, NBER WP 31161
Predominantly organisational. Microsoft’s 2026 Work Trend Index, based on 20,000 knowledge workers across ten markets surveyed between 18 February and 7 April 2026, reports that organisational factors such as culture, manager support and talent practices account for more than twice the reported AI impact of individual factors such as mindset and behaviour, at 67 percent against 32 percent. Only 26 percent of AI users in the same study say leadership is clearly and consistently aligned on AI.Source: Microsoft, 2026 Work Trend Index
Because the sanctioned route is missing rather than because the rules are strict. BCG found that 37 percent of employees say their company is not supplying the right tools, and that 54 percent would use unauthorised AI tools when corporate solutions fall short. Only 25 percent of frontline employees say they receive sufficient leadership support on how and when to use AI at work, while regular usage runs at 82 percent among those who do have that support against 41 percent among those who do not.Source: BCG, AI at Work 2025 (n=10,635)
Against a baseline taken before the tool arrives, on a small number of processes, using a number somebody already reports. Deloitte found that only 38 percent of organisations were tracking changes in employee productivity and only 16 percent had produced regular reports for the CFO on value created, which leaves the large remainder with no mechanism capable of answering the question. Usage counts, seats activated and prompts per week are activity measures, and optimising for them is what produces adoption without impact.Source: Deloitte, State of Generative AI, Wave 3
Founder & CEO, AInora
Building AI digital administrators that replace front-desk overhead for service businesses across Europe. Previously built voice AI systems for dental clinics, hotels, and restaurants.
View all articlesReady to try AI for your business?
Hear how AInora sounds handling a real business call. Try the live voice demo or book a consultation.
Related Articles
Shadow AI: What Employees Actually Paste Into AI Tools
What staff put into unmanaged assistants, why bans push usage further underground, and what an approved-tool list changes.
EU AI Act Article 4: What the AI Literacy Rule Actually Requires
The 27 July 2026 rewrite, why Article 4 sits outside the fine schedule, and what a programme actually contains.
Does Your AI Vendor Train on Your Data?
Training defaults, EU residency and retention across the major assistants, each cell sourced to the vendor doc it came from.
How to Start Using AI in Your Small Business
Choosing the first process, and what to do before buying anything.