The AI Productivity Paradox: Why More Technology Doesn't Always Mean Better Work

Last quarter I did something I had been putting off for about a year. I opened my company card statement, filtered it down to software vendors, and counted.

Eleven AI subscriptions. Four of them I had not opened in sixty days. Two were doing the same job. One I could not remember signing up for at all, and the timestamp on the welcome email suggested I had done it at 11pm during a free trial I fully intended to cancel.

The total was not going to bankrupt anyone. That was the part that bothered me.

Every line item looked reasonable on its own, which is exactly why none of them ever got questioned. A bloated stack never arrives as one bad decision. It arrives as forty small ones that each made sense on the day they were made.

So I went looking for evidence that all this software was actually working. I expected a mixed picture, some wins against some waste. What I found was stranger. The most careful measurements we have of AI at work keep landing on the same uncomfortable finding: people are worse at judging their own productivity than they are at doing the job.

That gap has a name, and it is older than the internet.

Below I will walk through the study that broke my confidence in my own stack, the company that spent eighteen months proving the point in public, the costs that never make it into a business case, and the five questions I now ask before anything gets a card number.

The paradox already had a name before AI showed up

In 1987, the economist Robert Solow reviewed a book for the New York Times and dropped a line that outlived nearly everything else written about computers that decade. You could see the computer age everywhere, he wrote, except in the productivity statistics.

American firms had spent the 1970s and 1980s pouring capital into information technology. Productivity growth slowed anyway. Erik Brynjolfsson gave the phenomenon its name in a 1993 paper, and the term stuck: the productivity paradox.

It resolved itself in the 1990s when growth finally picked up. Then it came back in the 2000s and never fully left.

Brynjolfsson revisited the question with two co-authors in 2018 and sorted the explanations into four buckets. Maybe the technology falls short of the hype. Maybe it works and we are measuring the output wrong. Maybe the gains are real but get burned up in zero-sum competition. Or maybe new technology takes years to diffuse, get implemented, and reach the point where it pays off. The authors found the last one most convincing.

Read that fourth explanation again, because it reframes the whole argument. The bottleneck was never the machines. It was the reorganisation of work around the machines, which is slow, unglamorous, politically awkward, and impossible to buy.

We are running the experiment again. The tools are better. The blind spot is identical.

The number that made me stop buying tools

Zylo's 2026 SaaS Management Index tracks license and spend data across a large sample of organisations. Two figures in it explain the entire shape of the current moment.

What it measuresThe number
Apps at the average organisation305
Apps at large organisations696
Change in app count, year over yearDown 0.07%
Change in total SaaS spend, year over yearUp 8%
Change in AI-native SaaS spend, year over yearUp 108%
Share of SaaS that IT actually manages13%
Renewals the average organisation handles per year211

Portfolios stopped growing. Spending did not. AI-native software spend more than doubled while the app count sat flat, which means the money is going into a smaller number of increasingly expensive things.

The 13% figure is the one that should worry any operations lead reading this. IT manages roughly one app in eight. The other 87% was bought by a team lead with a card, which is how the same organisation ends up with the duplicates I will come back to shortly.

Zylo also found that organisations underestimate their own app count by 1.7x and their own spend by 3x. And 61% reported cutting projects or initiatives because of SaaS cost increases they had not planned for.

That is 211 renewal conversations a year. Nearly one per business day, before anyone has done any work.

The study that broke my confidence

Between February and June 2025, the research group METR ran a randomised controlled trial that almost nobody in the AI industry wanted to look at directly.

They took 16 experienced open-source developers and gave them 246 real tasks on repositories they already knew well, mature projects averaging a million lines of code with more than 22,000 GitHub stars each. Tasks were randomly assigned to an AI-allowed arm, using Cursor Pro with Claude 3.5 and 3.7 Sonnet, or an AI-forbidden arm.

Going in, the developers predicted the tools would make them 24% faster.

After finishing, they reported they had been about 20% faster.

They were 19% slower.

The screen recordings showed where the time went. Developers accepted fewer than 44% of AI suggestions, then spent 9% of total task time reviewing and rewriting the code they did accept. Less time was spent typing. More time was spent prompting, waiting, reading, and repairing. The work felt lighter while taking longer.

Sit with the size of that error. A roughly 39-point gap between what these people believed about their own output and what a stopwatch recorded.

The researchers were careful about scope, and their caveats deserve repeating: the results do not mean AI is useless in other settings, the models were early-2025 vintage, and experienced engineers working in a codebase they have memorised is close to the worst case for an assistant. Fair points, all of them.

The finding that survives every one of those caveats is the perception gap. Almost every corporate claim about AI productivity traces back to self-report or to proxy counts like pull requests merged. METR showed self-report can be wrong by 40 percentage points in the confident direction.

Upwork's Research Institute surveyed 2,500 people across four countries and found the same fracture from a different angle. 96% of C-suite leaders expected AI to raise productivity. 77% of employees using AI said it had added to their workload. Another 47% said they did not know how they were supposed to get the gains their leadership was expecting.

Reviewing AI output ate the largest share of the extra time at 39%. Learning the tools took another 23%, and 21% reported absorbing more work on the theory that AI had made them more capable.

Where the hours actually go

What METR caught was overhead, the thing no business case has a row for. It comes from three well-documented sources, and none of them arrived with AI.

The toggle tax

In 2022, Rohan Narayana Murty and his co-authors published a study in Harvard Business Review that observed around 140 people at 20 Fortune 500 companies over five days.

The average worker toggled between applications about 1,200 times a day. Each switch cost a little over two seconds. That adds up to roughly four hours a week spent reorienting, or 9% of the working year, which is about five working weeks.

One finding stuck with me more than the headline number. To push a single supply-chain transaction through, each person involved switched roughly 350 times across 22 different applications.

Gloria Mark's attention research at UC Irvine puts the recovery cost of a significant interruption at 23 minutes and 15 seconds. Not every toggle triggers a full reset. Enough of them do.

This study was published before ChatGPT launched. Every AI assistant added since then is another window.

AI makes a broken process fast

Here is the mechanism nobody wants to name.

A tool applies force to whatever workflow it finds. If the workflow underneath is sound, the tool compounds it. If the workflow is a mess, the tool industrialises the mess.

Take a marketing team with no approval process. Drafting used to be the constraint, so the chaos downstream stayed hidden behind the bottleneck. Add an AI writer and the team can now produce forty pieces a week instead of four. The constraint moves to review, where two people are now drowning in drafts that a machine can generate faster than a human can read. Output went up ten times. Published work did not move. The team is busier, more tired, and measurably less effective than it was in February.

That is Brynjolfsson's implementation lag, arriving on schedule, dressed as a subscription.

You already own this tool twice

Zylo's data on redundancy reads like a joke until you check your own stack. The average organisation runs fifteen duplicate online training apps. Eleven project management tools. Ten separate team collaboration platforms.

More than half of workers surveyed by Cornell and Lokalise said the platforms they use daily overlap in what they do. 79% said their employer had taken no steps to consolidate anything.

The waste is measurable. BetterCloud's 2026 data puts unused licenses at 51%, the highest recorded, with only 49% of provisioned users logging in within the past 30 days. Large enterprises average roughly $18 million a year in spend on software nobody opens.

Klarna spent eighteen months proving this in public

Most companies bury this experiment. Klarna ran it on stage.

On 27 February 2024, Klarna and OpenAI published a joint announcement about Klarna's AI customer service agent. The first-month numbers were hard to argue with:

  • 2.3 million conversations handled, roughly two-thirds of Klarna's total volume
  • The workload of about 700 human agents
  • 75% of all customer chats across 23 markets and 35 languages
  • Average resolution time down from 11 minutes to under 2
  • Projected savings of $40 million across 2024

CEO Sebastian Siemiatkowski described himself as Sam Altman's favourite guinea pig. Klarna froze hiring for more than twelve months. Headcount fell 22%, down to around 3,500, mostly through attrition, and staff were told to lean on AI to cover the gaps their departing colleagues left behind.

Then the numbers underneath the numbers started moving.

Customer satisfaction declined on complex interactions. Complaints rose. Repeat contact rate climbed, which is the metric that matters, because it means people were coming back a second and third time about a problem that had been marked resolved.

In May 2025, Siemiatkowski told Bloomberg the company had "gone too far". He said the focus on cost had produced lower quality, and that customers needed to know a human was reachable.

Klarna is now rehiring. The replacement model is hybrid: remote agents on flexible hours, recruited from students and rural populations, working with AI tools rather than being replaced by them. AI absorbs the volume. People take the judgment calls.

The diagnosis worth stealing from this story is about measurement. Klarna's dashboard reported averages, and the average was excellent. The failure was hiding in the distribution, in the small percentage of conversations that were emotionally loaded or legally sensitive or simply weird, which is precisely the set of interactions that decides whether a customer stays.

Klarna is not an isolated case. IBM automated large parts of its HR function and rehired when the system could not handle anything requiring judgment. Commonwealth Bank of Australia reversed 45 layoffs after concluding the roles had never been redundant. Forrester now predicts that half of all AI-attributed layoffs will be quietly reversed, a pattern analysts have started calling the layoff boomerang.

The costs that never reach the business case

Klarna's reversal exposed something the original spreadsheet had no row for: unwinding the decision cost more than the decision had saved. Recruiting, onboarding and retraining a support function is expensive, and no AI replacement business case I have ever seen models the price of being wrong.

Here is what the case usually counts against what actually lands.

Cost typeWhat the business case countsWhat you actually pay
FinancialThe license feeLicense, integration build, the premium tier you upgrade to in month four, the annual commit at renewal
TimeOnboardingOnboarding, migration, the review pass on every generated output, 211 renewal conversations a year
HumanNothingDecision fatigue, learning fatigue, the 71% burnout rate Upwork measured among full-time employees
OperationalNothingSecurity review, compliance mapping, IT support, one more data flow to document under GDPR

That last row is getting more expensive. Every AI assistant bolted onto a tool is a new path corporate data travels down. Running a hundred tools with AI inside them means governing a hundred AI contracts.

What it looks like when the technology works

None of the above is an argument against buying software. I write reviews of AI tools for a living. The tools work, in the specific conditions where they work.

The pattern in every deployment I have seen pay off has four features, and Klarna's rebuilt support desk has all of them.

A narrow problem with a countable frequency. Not "improve productivity." Something like "we answer the same 40 shipping questions 900 times a week."

A baseline measured before the tool arrives. METR could see the 19% only because they had a control arm. Most teams start measuring after adoption, which makes comparison impossible and guarantees the perception gap goes undetected.

A named owner. Somebody whose job it is to check whether the thing is still earning its place.

Something retired to make room. One in, one out. This single rule would have caught most of my eleven subscriptions.

Automating 65% of a support queue at high quality while keeping human capacity for the other 35% delivers a smaller number on the slide than full replacement. It also survives contact with reality, which full replacement did not.

Five questions I ask before anything gets a card number

I stopped using a framework with a clever acronym. These are the questions, in order.

  1. What breaks today, how often, and what does each occurrence cost? If I cannot put a frequency and a number on it, I am buying a feeling.
  2. What is the baseline? Measured before purchase, in writing. Without it I will be relying on the same self-report that was off by 40 points at METR.
  3. What am I retiring? If nothing comes out, the toggle tax goes up and I have made the day worse to make a task easier.
  4. Who reviews the output, and does that review cost more than the generation saved? 39% of Upwork's respondents were spending their new time here.
  5. What does month 13 look like? The annual commit, the price rise, the migration cost of leaving. Renewal is where the waste gets locked in for another year.

Stop measuring activity

The perception gap survives because most teams track things that feel like productivity. Activity is easy to count. Outcomes are what you were buying.

Activity metric (feels productive)Outcome metric (is productive)
Messages sentTime from request to resolution
AI outputs generatedOutputs shipped without a rework pass
Seats provisionedSeats active in the last 30 days
Tools adoptedTools retired
Tickets closedRepeat contact rate
Hours loggedRevenue per employee

Look again at the right-hand column of that last pair. Repeat contact rate is the number that would have caught Klarna a year early, and it costs nothing to track.

Where this goes next

The direction of travel is fewer tools with more autonomy inside each one.

BetterCloud's 2026 research found that 42% of organisations with at least 1,000 employees are scaling agentic AI across multiple business functions. Those organisations have identified an average of 88 agentic use cases, with 32% already in production. 10% are running fully autonomous workflows today, and 60% expect AI agents to completely own key workflows within two years.

Gartner projects that more than 70% of organisations will centralise SaaS management on a dedicated platform by 2028, up from under 30% in 2025.

Set those against the Zylo figures from earlier in this article. Portfolios flat, spend up 8%, AI-native spend up 108%. Consolidation is happening. Cost is still climbing, because the surviving platforms are absorbing the AI capability and charging for it.

Which means the decision I got wrong at 11pm on a free trial is about to get harder. The tools are moving from things you use to things that act while you are not looking, and the review burden that made METR’s developers slower does not disappear when the agent runs unattended. It moves. Somebody still has to check the work.

Solow's line took about a decade to stop being true. The firms that got there first were not the ones that bought the most computers. They were the ones that rebuilt the work around them, and that part has not been automated yet.

Comments

Join the discussion and share your perspective.