What AI actually does to productivity
In short: The evidence does not say AI works or does not work. It says AI helps most where the work is already written down and the person doing it is inexperienced, and helps least — or hurts — where the work is complex and the person is already expert. In the one randomised trial of experienced developers, participants were 19% slower and believed they had been 20% faster. That 39-point gap between perception and measurement is the finding that should change how you run a pilot.
There are two ways to write about AI and productivity, and both are useless. One collects the enthusiastic numbers and calls it a transformation. The other collects the disappointing ones and calls it a bubble.
The actual literature is more interesting than either, because the studies disagree in a way that turns out to be systematic. Once you line them up, the disagreement resolves into a single pattern — and that pattern has direct consequences for which project you should start first.
Every figure below is linked to its source. Where a number is weaker than it looks, it says so.
The trial where the experts got slower
In early 2025, METR ran a randomised controlled trial on 16 experienced open-source developers across 246 real tasks in their own repositories — mature projects averaging over a million lines of code. Tasks were randomly assigned to allow or forbid AI assistance.
| What they predicted before starting | 24% faster |
| What they estimated afterwards | 20% faster |
| What was measured | 19% slower |
The interesting number is not the slowdown. It is the 39-point gap between what the participants experienced and what actually happened. These were not naive users; they were expert maintainers of code they knew intimately, and they were confidently wrong about their own week.
Be careful with this result. Sixteen people, one very specific population, and early-2025 tooling. It is not evidence that AI makes developers slower in general. It is strong evidence that self-reported productivity is not evidence of anything, which is a different and more useful claim.
The study where the novices gained a third
Run the same technology past a different population and the sign flips. Brynjolfsson, Li and Raymond studied the staggered rollout of a generative-AI assistant across 5,179 customer support agents. Issues resolved per hour rose 14% on average — but that average hides the finding:
- +34% for novice and low-skilled workers
- close to nothing for the experienced ones
The authors describe the mechanism as the model disseminating the practices of the better workers and helping newer ones move down the experience curve. In other words the tool was not adding capability so much as distributing capability that already existed inside the company — and there is nothing to distribute to someone who is already at the top of that distribution.
Put the two studies next to each other and the apparent contradiction disappears. Novices gained a third. Experts gained nothing, or lost.
The boundary you cannot see
The third study explains why even the gains are unreliable. In a preregistered field experiment, 758 BCG consultants completed 18 realistic tasks with and without AI.
Inside the range of things the model does well, the results were strong: 12.2% more tasks completed, 25.1% faster, 40% higher quality.
Then the researchers included tasks designed to sit just outside that range:
- consultants working alone were right 84% of the time
- consultants using AI dropped to 60–70%
The authors named this the jagged technological frontier: the boundary of competence is uneven and does not follow human intuitions about difficulty. Two tasks that look equally hard to a person can sit on opposite sides of it, and the model gives no signal about which side it is on. It is most convincing precisely where it is wrong.
Where the money went, and where the returns were
The fourth data point is about organisations rather than individuals. MIT Media Lab's Project NANDA surveyed the state of enterprise adoption in The GenAI Divide — 52 executive interviews, 153 leader surveys, 300 public deployments — and reported that 95% of pilots showed no measurable P&L impact.
That number has travelled much further than its methodology. It is a survey of executives, not a measurement of outcomes, and it is not peer-reviewed. Treat it as a description of how people feel about their own pilots.
Read that way, the headline is not the interesting part. This is: more than half of generative-AI budgets went to sales and marketing, while the highest returns were found in back-office automation. The money went where the work is visible. The returns were where the work is repetitive.
The report attributes the failures to a learning gap rather than to model quality — pilots stall because the tools cannot retain feedback or adapt to context. Which is another way of saying the context was never written down.
The control group is disappearing
The last finding is the one almost nobody is writing about, and it is the reason to act now rather than next year.
METR ran a larger follow-up with newer tools from August 2025. In February 2026 they published that the data cannot answer the question. Not because the effect was small — because they could no longer build a control group:
"An increased share of developers say they would not want to do 50% of their work without AI, even though our study pays them $50/hour."
"30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI."
"This implies we are systematically missing tasks which have high expected uplift from AI."
METR believes developers are probably faster now than their 2025 estimate suggested, but refuses to put a number on it: their own data is, in their words, "only very weak evidence for the size of this increase."
Sit with that. The best-designed measurement of this question in the world has become harder to run because people will not go back. Whatever is true about your own organisation's baseline, it is easier to measure today than it will be in six months, and at some point the honest answer becomes "we can no longer tell what this changed."
What this means for your first project
Four consequences follow directly from the evidence above, and none of them are about which model to use.
Do not trust how it feels. METR's participants were expert, motivated, and wrong by 39 points. If your evidence that a pilot is working is that the team likes it, you have measured enthusiasm.
Expect the gain on your juniors. The largest measured effect anywhere in this literature was +34% for inexperienced workers. If you are choosing between giving the tool to your best person or to your newest, the evidence points at the newest — which is the opposite of how most licences get allocated.
Assume the boundary is invisible. Build the review step in from the start, scaled to what a wrong answer costs. The BCG consultants who trusted the convincing answer did worse than those who had no help at all.
Measure the before, this month. Not because measurement is virtuous, but because the counterfactual is genuinely evaporating. A rough number now beats a precise argument later.
The honest summary
The evidence does not support "AI transforms knowledge work." It does not support "AI is overhyped" either. What it supports is narrower and more actionable:
The returns are not in the model. They are in whether the work has been written down, who is doing it, and whether anyone recorded what it cost before you started. Every study above is a different route to that same conclusion — and it is a conclusion about your organisation, not about the technology.
Frequently asked questions
Does AI actually make people more productive?
It depends almost entirely on who is using it and whether the work is specified. A study of 5,179 customer support agents found a 14% average gain — but 34% for novices and close to nothing for experienced staff. A randomised trial of 16 expert open-source developers found the opposite sign entirely: they were 19% slower with AI on their own mature codebases.
Why did experienced developers get slower with AI?
METR's participants worked in repositories they knew intimately — averaging over a million lines. The time spent prompting, reviewing and correcting output exceeded the time saved, because their own recall was already faster than the review loop. The striking part is that they could not tell: they estimated afterwards that AI had made them 20% faster.
What is the jagged frontier?
The boundary of what AI does well is uneven and does not match human intuitions about difficulty. In a study of 758 BCG consultants, tasks inside that boundary showed 40% higher quality and 25% faster completion; on tasks placed just outside it, consultants working alone were right 84% of the time while those using AI dropped to 60–70%.
Is it true that 95% of AI pilots fail?
That figure comes from an MIT Media Lab survey of executives, not from measurement of outcomes, and it is repeated far more often than its methodology is. Treated as a survey it is still informative: it reports that most pilots showed no measurable P&L impact, and that more than half of budgets go to sales and marketing while the highest returns were found in back-office work.
How should we measure whether AI is working for us?
Record the baseline before you start — hours, error rate, or what the work delays downstream. Perception is unreliable in both directions, and it is now becoming difficult to run a control group at all, because people refuse to work without the tools. The window for measuring your own before is closing.