Why AI isn't automatically making design teams more effective

Vice President, Digital Experience and Engagement
  • Twitter
  • LinkedIn

In 2023, Jakob Nielsen reviewed three controlled studies of generative AI in business work and put the average productivity gain at roughly 66 percent. Customer support agents handled 13.8 percent more inquiries per hour. Business professionals produced 59 percent more writing per hour. Programmers completed projects 126 percent faster.

Three years later, those task-level numbers have held. The thing every design leader was told would follow has not arrived. Projects are not finishing sooner.

DORA's 2025 research found that 90 percent of technology professionals now use AI at work and more than 80 percent believe it has made them more productive. The same research found that higher AI adoption correlates with an increase in software delivery throughput and an increase in software delivery instability, simultaneously. Thirty percent of developers report little to no trust in AI-generated code.

Faster and less stable, at the same time, in the same organizations.

That is the honest state of AI in delivery work right now, and experience design teams are living inside it. The research summary arrives Tuesday instead of Thursday, and the release still ships in March. There are ten concept directions instead of three, and the review cycle still takes three weeks. Something is absorbing the gains.

 

The gains are real, but they stop at the task boundary.

Read Nielsen's caveats and the whole picture changes. His figures came from GPT-3.5. Most participants were measured during a single use of the tool, often the first time they had ever used it. And the qualifier he wrote in 2023 turned out to be the most important sentence in the piece: the productivity gains "accrue only while workers are performing those tasks that receive AI support."

That is the entire problem, stated three years before most teams noticed it.

Design work is not a sequence of independent tasks. It is a chain of handoffs, each one a place where context gets compressed, translated, or dropped. Research findings become a deck. The deck becomes a set of requirements someone rewrote from memory. Requirements become a journey map in a design file. The journey map becomes screens. The screens become a handoff document. The handoff document becomes twenty Slack questions during the sprint.

AI has gotten extremely good at the work inside each box. It has done almost nothing about the space between them, and the space between them is where design projects spend their time.

So the summary that took eight hours now takes ninety minutes, and it sits unread for nine days because the stakeholder review cadence did not change. Net effect on the timeline: zero. Net effect on the team's confidence in AI: negative, because they can feel the effort they saved evaporating downstream.

 

AI amplifies whatever your workflow already is

DORA's researchers put this more bluntly than most vendors will. Where teams have quality infrastructure, AI acts as a powerful collaborator. Where they have "fragmented tooling, siloed data, or fragile infrastructure," AI "will simply help them generate technical debt faster."

Design organizations have an equivalent, and this year it produced a visible, measurable regression.

The 2026 WebAIM Million found that 95.9 percent of the top one million home pages had detectable WCAG 2 failures, up from 94.8 percent in 2025. Average errors per page rose 10.1 percent to 56.1. That broke six consecutive years of gradual improvement. WebAIM attributes part of the reversal to "increased reliance on 3rd party frameworks and libraries and automated or AI-assisted coding practices," alongside page complexity growing 22.5 percent in a single year and ARIA attributes climbing 27 percent to more than 133 per page.

Six years of slow progress on web accessibility reversed in the first year that AI-assisted production went mainstream. More output, generated faster, against a weaker set of standards, at greater complexity. The tools did what they were asked. The system around them had no way to catch what came out.

That is what amplification looks like when the underlying workflow is not ready. A team with rigorous accessibility gates, a mature component library, and clear review ownership gets a genuine accelerant. A team without them gets the same problems it already had, arriving faster and in higher volume.

 

Three places to look before you look at another tool

Most AI adoption in design starts with a demo and a use-case list. Start with the calendar instead. Pull the last three projects and find where the weeks went.

Look at the wait states. Not the work, the waiting. How many days elapsed between a research readout and the first design decision that used it? Between a finished concept and development readiness? These gaps are usually measured in weeks, and they are almost never on anyone's status report, because nobody owns the space between two owners.

Look at what gets rebuilt. Count how often someone recreated a journey map, a set of requirements, a component, or a research finding that already existed somewhere in the organization. Every instance is a retrieval failure, and retrieval is the one thing AI is unambiguously excellent at. This is usually the highest-return place to start, and it is the least glamorous.

Look at where feedback arrives. If stakeholder input consistently lands after concepts are built rather than before, generating concepts faster makes the problem worse. You will produce more work that gets reversed. Several teams have discovered this the expensive way over the past two years.

None of these three questions are about AI. That is the point. They tell you where AI would compound and where it would just add volume.

 

Measure the chain, not the task

Counting tools, prompts, seats, or assets generated tells you about activity. It tells you nothing about whether the work got better or arrived sooner. Most design organizations reporting AI progress to their executives right now are reporting activity, and the executives are starting to notice.

Measures that reflect the chain:

  • Elapsed days from research completion to the first design decision that cites it

  • Elapsed days from approved concept to development readiness

  • Number of revision cycles caused by late-arriving context

  • Percentage of shipped work using approved design system components

  • Number of handoff questions raised after development begins

  • Time required to locate existing research on a known topic

Set a baseline on the last two projects before expanding AI use anywhere. Without one, you will have opinions instead of evidence, and opinions do not survive a budget review.

 

Where to start when the operating model is not the first move

There is a version of this conversation that ends in a redesigned operating model, and eventually most organizations get there. It is not where anyone should start, because it is expensive, slow, and impossible to justify before you have proven anything.

Start with something bounded and broken.

An accessibility and usability sweep across a large property is a good first candidate. It is a well-defined problem, the current state is measurable, AI can extend coverage across pages and states no human team would get to, and the output is a prioritized remediation list with an owner and a date. The team learns where AI is reliable and where it needs a human check, on work that has a defensible return, before anything structural gets touched.

The same logic applies to a research repository nobody can search, a component library that has drifted from what is in production, or a handoff process generating the same twenty questions every sprint. Each is small, each is measurable, and each teaches the organization something true about where AI belongs.

Then the operating model conversation happens with evidence behind it instead of a slide.

 

The uncomfortable part

The teams getting real leverage from AI right now are not the ones with the best tools. They are the ones whose workflows were already coherent enough to amplify. Clear ownership, a design system that reflects production, research people can find, review cycles that happen before the expensive work.

That is unwelcome news, because it means the work that unlocks AI is the unglamorous work most design organizations have been deferring for years. It is also good news, because that work has a known cost and a known return, and it pays off whether or not the next model release lives up to its launch video.

The technology is not the constraint. The chain is.

Building the AI-enabled experience

Building the AI-Enabled Experience Design Workflow for In-House Teams

Read our perspective