
Task-level AI gains are well documented. Organization-level gains are largely absent. An NBER study of roughly 6,000 executives found that 89-95% of firms saw no measurable productivity or employment impact over three years, while the same period produced consistent evidence that individual workers do save real time.
Both findings are correct, and the gap between them is a measurement problem before it is a technology problem. Time saved on a task only becomes productivity if it flows into work that matters rather than dissolving into tool-switching, review overhead, and administrative residue - and most organisations have no instrumentation capable of telling the difference. Teams that deploy employee monitoring software after an AI rollout usually discover their baseline data predates the rollout, which makes before-and-after comparison impossible. The measurement has to be designed before the tools arrive, and almost nobody does that.
What the published numbers actually say
Cited figures for AI time savings vary by a factor of six. They are not contradicting each other - they measure different populations.
Source
Reported saving
Population measured
St. Louis Fed
~2.2 hours per week
US workers using generative AI
Goldman Sachs
40-60 minutes per day
Firms past the experimentation stage
McKinsey
6.4 hours per week
Organisations running production AI agents
McKinsey, adjusted
~1.4% of total hours
All workers, including non-users
NBER/Microsoft
3.6 hours per week
Email management specifically, a 31% reduction
The last two rows explain most of the confusion. The 6.4-hour figure describes organisations that built AI into workflows, not organisations that bought licences. Average the same data across an entire workforce, including people who never open the tools, and it drops to roughly 1.4% of hours.
If you are quoting a headline number to justify spending, know which population it describes. Most do not match yours.
Why savings disappear before reaching the P&L
Four mechanisms account for most of the leak, and each is measurable if you decide to.
Saved time fragments. Twenty minutes recovered from drafting does not become twenty minutes of deep work. It becomes three interruptions, a Slack thread, and a coffee. The saving is real, and the value is not captured, which is the point both McKinsey and the St. Louis Fed make when they distinguish time saved from value captured.
Review overhead eats the gain. Upwork research found 77% of freelancers using generative AI reported it added to their workload, primarily through review and validation. This is the most underweighted cost in the category: output arrives faster, and someone has to check it, and checking unfamiliar output is slower than checking your own.
Trust gaps create duplicate work. Stack Overflow found 84% of developers using or planning to use AI tools, but only 3% highly trusting the accuracy of outputs. Low trust means verification, and verification means the task is partly done twice.
Adoption concentrates. Gartner found 37% of employees skip AI tools because colleagues are not using them, and only 13% of workers report their company offered training. A tool used intensively by a fifth of the team produces a fifth of the possible organisational gain, however impressive the individual results.
The pattern: gains are real at the task level, and organisations lose them between the task and the quarter.
Measure the input first: is anyone actually using it
Before measuring savings, establish usage. Most organisations cannot answer this, which makes everything downstream guesswork.
Licence utilization. How many assigned seats were used in the last 30 days? Compare to seats purchased. Gaps here are the cheapest thing to fix in the entire exercise.
Depth of use. Daily users behave differently from monthly ones - one analysis found a third of daily users save four or more hours weekly, a figure that occasional users never approach. Segment by frequency before averaging anything.
Distribution by team and role. Concentration matters more than the mean. Engineering at 80% adoption and operations at 10% produces a misleading organisational average and points directly at where the training gap sits.
Which tools, in what mix? Writing and coding assistants have penetrated fastest; scheduling and analytics move more slowly. Different categories carry different return profiles, and averaging across them hides both the wins and the waste.
Application-level tracking answers all four. Tools that log which applications ran and for how long - including AI assistants - turn "we bought Copilot licences" into "34% of assigned seats were used last month, concentrated in two teams." That is the difference between a procurement record and a measurement.
Then measure the output: did anything change
Usage data tells you the tools are open. These four signals tell you whether it mattered.
Cycle time on comparable work. How long from start to shipped, for a category of work that existed before and after? This is the strongest available signal because it captures the whole path, including the review overhead that erodes the gain.
Volume at constant quality. Tickets closed, documents produced, code merged - paired with a quality measure. Volume alone is worthless: AI makes it trivially easy to produce more output of lower value, an effect that one analysis costs at roughly $186 per employee monthly in wasted downstream handling.
Rework rate. How often does work come back? If it rises after AI adoption, the tools are shifting effort downstream rather than removing it. This is the single clearest indicator of the review-overhead problem, and almost nobody tracks it.
Where do recovered hours go? The question that separates real gains from illusory ones. If drafting time dropped 30% and nothing else changed, either the hours went somewhere invisible or the estimate was wrong. Application-level time data can show whether recovered hours moved into higher-value categories or were scattered.
How to actually run the measurement
Establish the baseline before deployment. This is the step that gets skipped, and skipping it is unrecoverable - you cannot reconstruct last quarter's cycle time after the fact. Capture at minimum: time distribution by application category, cycle time for two or three recurring work types, and volume with a quality measure. Four weeks is enough.
Pick one work type, not the whole organisation. Organisation-wide averages are exactly where the signal disappears. Choose something repetitive and measurable - support responses, standard reports, routine code review - and measure that properly.
Run a control if you can. Two comparable teams, one adopting, one not, over the same period. Rarely politically possible, and worth attempting because it is the only design that survives the objection that something else changed.
Measure for at least a quarter. Early weeks show a learning dip, then an enthusiasm spike, then something closer to reality. Anything shorter measures novelty.
Segment by depth of use. Report daily users, weekly users, and non-users separately. The organisational average is the least informative number you can produce.
Ask people what changed. Quantitative data shows that drafting time fell. It does not show that the person now spends the recovered time verifying output they do not trust. Structured interviews with a handful of heavy users surface the review-overhead problem faster than any dashboard.
What not to do
Do not measure keystrokes or activity levels as a productivity proxy. AI-assisted work involves less typing by design. An employee producing more with fewer keystrokes will look less productive on any input-based metric, which is precisely backwards.
Do not attribute every change to AI. Team composition, seasonality, process changes, and client mix all move the same numbers. Without a baseline or a control, attribution is an assertion.
Do not treat vendor benchmarks as your forecast. Figures like 6.4 hours weekly describe organisations that redesigned workflows around the tools. A company that bought licences and sent an announcement email is not in that population.
Do not measure individuals. AI adoption data at the individual level invites exactly the wrong conversation and produces defensive usage patterns. Team and role-level aggregation answers the business question without creating a performance-review problem.
What a realistic result looks like
Organisations that measure carefully generally find something like this.
Task-level savings are genuine and smaller than the headlines - closer to the St. Louis Fed's 2.2 hours weekly than McKinsey's 6.4, unless workflows were actually redesigned.
Savings concentrate in a minority of heavy users, so organisational impact tracks adoption depth far more than tool capability.
A meaningful share of the gain returns as review overhead, and this is normal rather than a failure.
Recovered hours default to fragmentation unless someone deliberately redirects them.
That is a defensible result and a useful one. It is also considerably more honest than the alternative, which is quoting a vendor benchmark and hoping nobody asks how it was verified.
Frequently asked questions
How much time does AI actually save employees?
Published figures range from about 2.2 hours weekly for general users to 6.4 hours for organisations running production AI agents, with Goldman Sachs reporting 40-60 minutes daily at firms past experimentation. Averaged across a whole workforce including non-users, the figure falls to around 1.4% of total hours.
Why do most companies see no productivity gain from AI?
An NBER study of roughly 6,000 executives found 89-95% of firms reported no measurable impact over three years. The common causes are shallow adoption, review overhead consuming the saving, and recovered time fragmenting instead of flowing into higher-value work.
How do I measure AI ROI in my own team?
Baseline before deployment, pick one repeatable work type, measure cycle time and volume with a quality check, segment results by depth of use, and run for at least a quarter. Without a pre-deployment baseline, the comparison cannot be made afterwards.
What metrics show whether AI is working?
Cycle time on comparable work, output volume paired with quality, rework rate, and where recovered hours went. Rework rate is the most diagnostic and least tracked - rising rework means effort moved downstream rather than disappearing.
Can activity monitoring measure AI productivity?
Application-level data answers usage questions well: which tools ran, how often, by which teams. Input-based metrics like keystroke counts or activity percentages do not, because AI-assisted work involves less typing by design and would register as reduced productivity.
Why does AI add to some people's workload?
Review and validation. Upwork research found 77% of freelancers using generative AI reported increased workload, and Stack Overflow found only 3% of developers highly trust output accuracy. Low trust means verification, and verification returns part of the saving.
How long before AI adoption shows measurable results?
Plan for at least a quarter. Early weeks show a learning dip, followed by an enthusiasm spike, before settling. Measurement periods shorter than that capture novelty rather than a durable change in how work gets done.
Should we measure AI usage per employee?
Aggregate by team and role rather than individual. Individual-level adoption data invites performance conversations that produce defensive usage, and the business question - where the tools are working and where training is missing - is answered at team level anyway.
