SaaS
Nobody Actually Knows if AI Improved Output: Measure the Work, Not the Usage
Published: 2026-07-13
The Problem
Companies mandate AI use and count seats and prompt volume like a KPI, but nobody honestly measures whether that usage actually lifted an individual's or a team's output.
Why Now
Knowledge work outside engineering still has no equivalent measurement tool, and a causal, quality-adjusted measure of before versus after can catch the real gains that usage dashboards miss.
Recommended Talent
Someone who has worked in product analytics and causal inference (A/B and holdout design) with a feel for building performance metrics that resist gaming
The Problem
Companies push people to use AI, then measure the result by usage: how many seats, how many prompts, what percentage of code suggestions were accepted. The trouble is that usage is not output. Heavy use does not guarantee better work, and under a mandate, usage metrics are the easiest thing to inflate. A manager watches the blue line on a dashboard climb and believes it is working, while the actual question, whether the team’s output got better, goes unanswered.
The evidence points straight at that illusion. In an NBER survey of about 6,000 executives, over 80% detected no discernible productivity impact from AI. On the engineering side the telemetry is harsher. Faros AI tracked 22,000 developers over two years and found that although 75% use AI tools, most organizations saw no measurable performance gain, and while throughput rose 66%, incidents and bugs rose faster: the probability of a production incident per merged PR more than tripled year over year, and bugs per developer were up 54%. Usage up, quality down, in the same window. Managers are judging blind in the middle of it.
Why Now
Two currents overlap right now. One is that AI mandates have gone mainstream. The more companies force usage from the top, the more thirst there is for a measure of real effect rather than raw usage. The other is a blind spot in existing tools. Engineering already has measurement: Faros AI and DX with its DX Core 4 read throughput and quality off git and PR telemetry. But two gaps remain. First, even these tools still report a lot of usage and throughput, and rarely isolate causally whether a given person truly improved because of AI. Second, knowledge work with no git log, like sales, support, operations, marketing, and legal, has almost no equivalent. The mandate lands on all of those functions, yet measurement stops at the engineering team.
flowchart LR
A[Pre-AI baseline<br/>output and quality per person] --> C[Causal measurement<br/>holdout, staggered rollout]
B[Output after AI<br/>speed, rework, error rate] --> C
C --> D[Real per-person gain<br/>quality-adjusted delta]
D --> E[Honest signal for managers<br/>output, not usage]
How to Build It
Start narrow, in one function where output is visible. For customer support, that is time to resolution, reopen rate, and quality-review scores. Three things matter. First, record a per-person baseline before rollout. Second, measure causally: use a staggered rollout, some people turn AI on first and others later, or a holdout group, so you pull the real improvement rather than a correlation. Third, adjust for quality: if work got faster but rework, errors, or incidents rose, subtract that loss from the output. The per-person delta that comes out goes quietly to the manager. Keep access and display narrow so it reads as a coaching signal, not surveillance.
Revenue fits a team or per-seat subscription. To lower the early resistance to adoption, do not rank people from day one. Draw a team-level map first, showing which kinds of work AI actually lands on and where it spins uselessly. Individual rankings come only after trust is built.
Success Criteria
Two conditions have to hold. First, you must be able to define a gaming-resistant, quality-adjusted output metric for each function. If the metric is flimsy, this tool decays into yet another usage dashboard. Second, the organization has to be able to accept the honest signal that a given mandate did not work. Handing leadership that pushed AI a verdict of no effect is politically uncomfortable, and without a culture that tolerates that discomfort, the tool ends up in a drawer. Clear both and you strip out the usage game and, for the first time, price where AI genuinely makes the work better. Then there is one question left: in our team, how many tasks did AI actually improve, right now?
Related Content
Build this together
Find collaborators