Most design organizations reporting on AI right now are reporting activity: tools adopted, people trained, prompts submitted, assets generated, hours saved, usually self-reported.
Every one of those numbers can go up while nothing about the business changes, and finance leaders have started to notice. The question worth asking is whether the work got better or arrived sooner, and almost nobody is instrumented to answer it.
Start with the sentence you are trying to be able to say
Before choosing metrics, write down what you expect AI to change, specifically enough that it could be wrong.
"We expect to reduce elapsed time from research completion to first design decision from an average of eighteen days to under seven." That is a claim. It can be measured, it can fail, and if it succeeds, it is worth money.
"We expect AI to make the team more efficient" cannot fail, which is why it is the most common objective in the category and the least useful.
The specificity matters because different objectives require entirely different measurement. Reducing repetitive work, accelerating research analysis, increasing concept exploration, improving handoff quality, and shortening review cycles are five different programs, each requiring its own measurement approach. Teams that skip this step end up collecting whatever their tools happen to export, which is always activity data.
Measure the chain, not the task
Task-level savings are real and insufficient, and the best studies in the field show why.
Brynjolfsson, Li, and Raymond's study of 5,179 customer support agents, published in the Quarterly Journal of Economics, found average productivity up 14 percent. The average hid the finding that matters: novice workers improved 34 percent while experienced, highly skilled workers saw minimal impact. Noy and Zhang's randomized experiment in Science, covering 453 professionals on writing tasks, found completion time down 40 percent and quality up 18 percent, with the largest gains again going to the initially weakest performers.
If you report a blended average from your own team, you will hide the same thing they found. Segment by seniority. A 25 percent average across your design org probably means your juniors got substantially faster and your principals did not move, which is worth knowing, because your principals are where the queue forms.
Both studies also share a structural limit worth stating plainly: they measured bounded, individually-owned tasks, and they measured the worker during the task.
A research summary that took eight hours now takes ninety minutes. Then it waits nine days for a stakeholder review cadence that did not change. Net effect on the delivery date: zero. The task metric reports a 5x improvement. The project reports the same March date it reported in January.
Measures that reflect the chain rather than the task:
- Elapsed days from research completion to the first design decision citing it
- Elapsed days from approved concept to development readiness
- Revision cycles caused by context arriving late rather than by changed requirements
- Handoff questions raised after development began, segmented by cause
- Percentage of shipped work using approved design system components
- Time required to locate existing research on a known topic
- Issues identified before development versus after
These are harder to collect, and they are the ones that correspond to money.
Quality, or the correction tax
Faster output that requires heavy correction is not faster.
The number worth tracking is the correction tax: what proportion of AI-assisted output is accepted with minor edits, what proportion needs substantial rework, and what proportion gets discarded. A team generating drafts in ninety seconds that take three hours to fix has bought itself a worse process with better optics.
Assess against accuracy, completeness, accessibility, design system compliance, evidence of actual customer need, and brand fit. DORA's finding that 30 percent of developers report little to no trust in AI-generated code is a quality signal expressing itself as a review burden, and review capacity is the constrained resource in most delivery organizations.
The team metric nobody wants to collect
AI adoption changes the experience of doing the work, and not uniformly in the direction the pitch deck promised.
It removes administrative drudgery, which people welcome. It also shifts the composition of the job toward reviewing and correcting machine output, which some practitioners find substantially less satisfying than producing work themselves, and it can create pressure to increase volume without any corresponding clarity about priority.
Track time spent on repetitive versus strategic activity, confidence using approved tools, clarity of governance, time available for actual customer contact, and, importantly, where AI is creating additional work rather than removing it. That last question surfaces things no dashboard will, and the answers are usually specific and fixable.
Pair the numbers with conversation. Quantitative data tells you what moved. Only your team can tell you why.
Connecting to customer outcomes without overclaiming
The strongest evidence is that customers had a better experience: task completion, error rates, conversion, satisfaction, support contact volume, abandonment, feature adoption, accessibility performance, successful self-service, retention.
Attribution will be imperfect and you should say so plainly rather than construct a causal story that will not survive scrutiny. Many things influence experience performance.
What you can defend is a documented chain: we identified this journey gap during design, using this method, which included AI-assisted analysis; we resolved it before launch; and this is what that journey stage did afterward. That is an honest claim about process contribution, and it holds up better in a review than a confident percentage nobody can reproduce.
Two statistics to leave alone
Both are circulating heavily in this conversation, and both will cost you credibility with an informed audience.
The METR "developers are 19 percent slower with AI" finding. The finding is real, from early 2025, and widely quoted. But METR published an update in February 2026 reporting serious problems with their follow-up: developers refused to participate without AI even at $50 an hour, and 30 to 50 percent admitted avoiding tasks they expected AI to handle well. METR now states their newer estimate "is a lower-bound on the true productivity effects of AI" and is abandoning the design. Citing the original figure today without that context is a factual error, and technically literate readers will catch it.
The "95 percent of GenAI pilots fail" report. It is everywhere. It is 52 executive interviews, 153 survey responses, and roughly 300 implementation reviews from July 2025, self-described as preliminary and anonymized, with no vendor case studies and no governance analysis. It has been widely and reasonably criticized.
DORA's 2025 research supports the same argument on far better methodology: 90 percent adoption, over 80 percent believing they are more productive, and higher adoption correlating with increases in both delivery throughput and delivery instability. Use that instead. It is stronger and it will survive a challenge.
Governance belongs in the scorecard
An AI program that improves delivery speed while creating privacy, accessibility, or compliance exposure has not succeeded. It has moved the cost somewhere that shows up later and larger.
Worth reporting alongside efficiency: proportion of work happening in approved tools, incidents involving restricted information, proportion of AI-assisted work receiving required human review, frequency of accessibility validation, and documented sources for consequential claims.
The 2026 WebAIM Million provides the cautionary case: 95.9 percent of top home pages carrying detectable WCAG failures, average errors up 10.1 percent, the first regression in six years, with WebAIM pointing partly at AI-assisted coding practices. Those organizations shipped faster. Whether they would describe the year as a success depends entirely on what they measured.
Baseline first
None of this works retroactively. Before expanding AI use, document current performance on the last two or three projects: how long the work took, where it stalled, how much rework occurred, which activities consumed the most time, what quality issues recur, and how customers currently perform in the experience.
It takes a couple of days and it is the difference between demonstrating improvement and asserting it. Teams that skip it end up arguing from anecdote in exactly the meeting where anecdote does not work.
The framework, briefly
Four questions, asked together:
Efficiency. Is the work taking less time and effort, measured across the chain rather than the task?
Quality. Is the output accurate, complete, accessible, and consistent? How much correction does it need?
Experience. Are customers and the team better off?
Governance. Is the organization using this responsibly and consistently enough to keep doing it?
No single number answers whether AI is working. But a team that can speak to all four is having a fundamentally different conversation with its executives than one presenting a prompt count.
Schedule an AI-Enabled Experience Design Working Session
To read more about where AI creates friction across the delivery process, read our blog on the delivery gap.
For a deeper look at measuring design-to-development handoffs, read our blog on handoff metrics.
Sources
- Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond, "Generative AI at Work," The Quarterly Journal of Economics, vol. 140, no. 2, 2025, pp. 889-942. https://academic.oup.com/qje/article/140/2/889/7990658
- Shakked Noy and Whitney Zhang, "Experimental evidence on the productivity effects of generative artificial intelligence," Science, vol. 381, 2023. https://www.science.org/doi/10.1126/science.adh2586
- DORA, "Balancing AI tensions: Moving from AI adoption to effective SDLC use," 2025. https://dora.dev/insights/balancing-ai-tensions/
- METR, "We are Changing our Developer Productivity Experiment Design," February 24, 2026. https://metr.org/blog/2026-02-24-uplift-update/
- WebAIM, "The WebAIM Million," 2026. https://webaim.org/projects/million/

