The automation numbers people quote are not real
Two controlled trials, opposite results, one shared conclusion. And a measured gap between what workers believed AI did for them and what it actually did.
In the one controlled trial that measured it properly, experienced developers using AI tools were 19 percent slower, and afterwards estimated they had been 20 percent faster. That gap, 39 percentage points in the flattering direction, on work they had just finished, is the most useful fact in this entire subject. Every percentage you have been quoted about automation is self-reported, and self-report is what that study measured being wrong.
The trial that found a loss
In July 2025 METR published a randomised controlled trial on experienced open-source developers working in their own repositories. Sixteen developers, all with moderate AI experience, completed 246 real tasks in mature projects they had worked on for around five years on average. Each task was randomly assigned to AI allowed or AI not allowed, and when allowed they used what they liked, mostly Cursor with the then-current Claude models.
Before starting, the developers forecast that AI would cut their completion time by 24 percent. Afterwards, having done the work, they estimated it had cut it by 20 percent. The measured result was that tasks took 19 percent longer with AI available.
METR were careful about the scope and so should anyone citing it. Sixteen developers is a small sample, it is a snapshot of early-2025 tool capability, and it covers one specific setting: expert practitioners on codebases they know intimately. They have since revised their experiment design, which is what a serious research group does. It is not proof that AI makes people slower in general.
What it is proof of is narrower and more damaging: that people cannot accurately report their own productivity change from using these tools, even immediately afterwards. Which means every vendor case study, every survey of perceived time savings, and every "we automated 80 percent of our operations" claim is built on a measurement instrument that has been tested and failed.
The trial that found a large gain
If the story stopped there it would be a hit piece, and the evidence does not support that either. The most rigorous study pointing the other way is larger and cleaner.
Brynjolfsson, Li and Raymond studied the staggered rollout of a generative AI assistant across 5,179 customer support agents, published in the Quarterly Journal of Economics. Productivity, measured as issues resolved per hour, rose 14 percent on average. Customer sentiment improved and staff retention improved with it.
But the average conceals the finding. The gain was 34 percent for novice and low-skilled workers, and close to nothing for experienced, highly skilled ones. The mechanism the authors identify is that the tool spreads the practices of the best agents to everybody else, moving new people down the experience curve faster.
Both studies say the same thing. AI substantially helps people who do not yet know how to do the work, and does approximately nothing for people who already do.
Read them together and the apparent contradiction dissolves. METR studied experts on work they knew deeply and found a loss. Brynjolfsson studied a mixed workforce and found the entire gain concentrated among the least experienced. Same pattern, two directions, and the pattern is about who is using it rather than about the technology.
Laid out as one range, the measurements stop looking like disagreeing studies and start looking like a single gradient with the skill level on one axis:
| Who was actually measured | Measured effect |
|---|---|
| Experts, on code they had worked in for years | 19% slower |
| Experienced support agents | close to nothing |
| The whole workforce, averaged | 14% faster |
| Novice and low-skilled agents | 34% faster |
That is a 53 point spread, from the same category of tool, in the same era, measured properly at both ends. Which is the practical reason to distrust any single figure you are quoted: it is a point somewhere on that range with the context stripped off, and the context is the only thing that determines which point you would land on.
Two numbers are worth carrying out of this. 39 points is how wrong people were about their own just-completed work, which tells you what self-report is worth. 53 points is how much the real effect varies by who is holding the tool, which tells you what somebody else's number is worth. Neither of those is an argument against the tools. They are an argument against believing a percentage without knowing who it was measured on.
What that implies for a real business
This is where the evidence becomes operationally useful, because it tells you where to point the tools and, more importantly, where not to.
| Expect little or nothing | Expect the real gain | |
|---|---|---|
| Who | Your most experienced technician, doing work they have done for twenty years | New hires, the admin layer, and anybody currently learning the job |
| Why | They are already at the practice frontier the model was trained toward | The tool carries the best existing practice to somebody who has not built it yet |
| The risk | Time lost reviewing output that is worse than what they would have produced | Over-reliance without the judgement to catch a wrong answer |
| What to do | Leave them alone, or use it only for the paperwork around their work | Deploy it here first, with review, and measure the throughput |
The practical version for a fifteen-person service company: do not aim any of this at the master craftsman diagnosing a fault. Aim it at the person answering the phone, the person producing the quote, and the person who joined four months ago and is still asking how things are done. That is where the 34 percent lives, and it is also the part of the business with the most turnover, which compounds the benefit.
It also reframes onboarding as the highest-value application. If the effect is largest for people early on the curve, then the return is not a one-off efficiency gain, it is a permanently shorter ramp for every future hire. In a trade where hiring is the binding constraint, that is worth more than the hours saved.
How to measure your own, since you cannot trust the report
The whole point of the perception gap is that asking people whether something helped will give you a wrong answer, delivered confidently. So the measurement has to be structural.
- Pick one countable output that existed before you changed anything: quotes issued per week, jobs invoiced per week, enquiries handled per person
- Get eight weeks of baseline before the tool arrives, because after it arrives you will never be able to reconstruct it
- Change one thing at a time, since two simultaneous changes produce an unattributable result and an argument
- Never ask anyone how much time it saved them, and specifically do not put that question in a survey
- Watch quality alongside throughput, because faster output that gets reworked is not faster
- Give it eight weeks before judging, since the first fortnight measures the learning curve rather than the tool
Refusing to ask people how much time it saved will feel rude, and it is the most important rule here. The people using it are not lying, and their estimate is still unusable, because that is precisely what the METR study demonstrated on a group of skilled professionals reporting on work they had just done. If you want to know whether something worked, count the output.
What an honest claim looks like
Given all of the above, here is the standard I would apply to any figure quoted about automation, mine included.
A credible claim names the specific task, not the business. "We reduced the time from enquiry to quote issued from four days to one" is checkable. "We automated our operations" is not a claim, it is a mood. Anything expressed as a percentage of a whole company is almost certainly the second thing wearing the clothes of the first.
It also states what was measured and against what. A number with no baseline is not a result. And it is bounded: administrative and coordination work in a service business is a real slice and a minority slice, with skilled labour, parts and vehicles making up most of the cost base. A claim that a large share of a whole business was automated is arithmetically describing something that would have to include the technicians, and it does not. That bound is the honest part of the operating thesis and the part most often left out.
The reason to be this strict is not scepticism for its own sake. It is that the underlying opportunity is genuinely large, and inflated claims are what make an owner who checked one against their own business dismiss the whole subject. A measured 14 percent on the right population, with a permanently shorter onboarding ramp attached, is a very good outcome. It just does not sell as well as a number somebody made up. Where the metric that resists this kind of flattery lives is in the two ways to raise revenue per person, and the practical deployment map is in AI-native operations, without the hype.
The short version
- Two numbers to carry: 39 points is how wrong people were about work they had just finished, which prices self-report. 53 points is how much the real effect varies by who holds the tool, which prices somebody else's number.
- METR's 2025 trial: experienced developers were 19 percent slower with AI on repositories they knew well, having forecast 24 percent faster and reported 20 percent faster afterwards.
- That 39-point perception gap is the finding. Self-reported productivity change is a measurement instrument that has been tested and failed.
- Brynjolfsson, Li and Raymond across 5,179 support agents: 14 percent average gain, 34 percent for novices, near zero for experienced workers.
- Both point the same way. Deploy this at the admin layer and at new hires, not at the technician who has done the work for twenty years.
- Measure output you counted before the change, one variable at a time, over eight weeks, and never ask anyone how much time it saved them.
Questions I get on this
Does AI actually make people more productive?
Why should you not trust reported time savings from AI?
Where should a small business apply AI first?
Sources: METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (July 2025, 16 developers, 246 tasks; METR has since revised the experiment design). Erik Brynjolfsson, Danielle Li and Lindsey Raymond, "Generative AI at Work", Quarterly Journal of Economics, covering 5,179 customer support agents.
