Describe an audience and a goal to any current model and you will receive a journey map in under a minute. Stages, actions, needs, emotional states, pain points, opportunity areas, formatted and ready to present.
The output quality is often reasonable, and that is exactly what makes it a problem. It arrives looking identical to the artifact your organization produces after six weeks of customer research, and nothing about its appearance signals the difference.
A journey map makes a claim about how real people move through an experience. When that claim comes from research, it carries evidence. When it comes from a prompt, it carries the model's aggregate prior about how people like that generally behave. Both render the same way in a slide.
Six months later, nobody remembers which one they were looking at, and a roadmap has been built on it.
Use generation to decide what to investigate
The generated journey is legitimately useful in one specific role: as a hypothesis you are about to attack.
Produce it early, before fieldwork, and treat it as a list of assumptions rendered visually. Which stages does it claim exist? Which pain points does it assert? What does it take for granted about how customers make this decision? Every one of those is now a research question with a shape, which beats starting from a blank discussion guide.
Then label it, in the file name and on the artifact itself, as an assumption map. That label is the only thing standing between a hypothesis and a stakeholder who saw it once and now believes it.
The stronger application runs in the other direction. Give the model your actual research and ask it to organize the evidence you already have: group findings by journey stage, compare patterns across segments, identify which stages have thin evidence, and surface where qualitative feedback and behavioral data disagree.
That disagreement question is the highest-value one you can ask of a research corpus, and the one humans are worst at answering, because it requires holding two datasets in mind at once and noticing the seam. Where customers say one thing and do another is almost always where the real design problem lives.
Generate more than one future
Design teams converge too early. The first plausible journey becomes the target journey, usually within a week, and usually because somebody had to present something on Thursday.
Generation breaks that habit well. Ask for three structurally different approaches to the same customer need: one that solves it through self-service, one that solves it with guided human support, and one that solves it by changing the underlying business process so the need never arises. Those are three different bets with different cost structures and different organizational consequences, rather than three visual treatments of the same idea.
The same technique works on the conditions teams forget about until QA finds them. Run the journey again for a returning customer instead of a new one, for an enterprise account instead of a small one, and for someone using a screen reader. Run it for the customer who is confident and again for the one who is not sure they are even in the right place. Then run the exception case, and the moment where a digital channel hands off to a human.
Those variations surface missing states, which is where projects usually lose their schedule, because the gaps get discovered during development rather than during design.
From journey to prototype, without the map going stale
Journey maps have a well-known failure mode. They get made, presented, admired, and then the screens get designed by someone looking at a different document entirely.
The connection point is requirements. A journey stage that has been validated should generate its downstream artifacts directly: the user stories, the content requirements, the functional requirements, the edge cases, the testing tasks, and the measurement plan for whether the stage actually improved.
This is mechanical translation work of exactly the kind AI handles well, and it is the step most teams still do by hand and inconsistently. Done properly it also creates a traceable line: this screen exists because of this requirement, which came from this journey stage, which came from this research finding. When a developer asks in month four why a step exists, somebody can answer.
The reverse check is equally valuable and almost never run. Compare the prototype against the journey and ask what the journey documented that the prototype does not address. That gap is usually real, and usually invisible until someone looks for it.
Constraints are what make generated prototypes usable
AI design tools produce polished-looking interfaces from thin prompts. Nielsen Norman Group's evaluation against real design scenarios documented a consistent signature in the results: visual style that is flat and interchangeable, generic aesthetics with no brand differentiation, related information separated by excessive spacing, absent visual hierarchy, poor color contrast, and patterns applied from entirely the wrong context. In one case the tools imported social-media profile conventions onto a learning dashboard, which promoted secondary information over primary.
Those are context failures rather than capability failures. The tool was asked to design something without being told anything true about the organization, the audience, or the system the design has to live inside.
How much does supplying constraints actually change the output? Guriță and Vatavu measured it at the 2025 Web for All Conference, comparing UI code generated from accessibility-agnostic prompts against the same tasks with accessibility-oriented prompts.
The agnostic prompts produced a 58 percent violation rate under expert evaluation. The accessibility-oriented prompts produced 19 percent.
The per-criterion numbers are starker. Under agnostic prompting, alternative text and information structure failed 100 percent of the time, keyboard navigation failed 80 percent of the time, and status messages failed 90 percent. Adding accessibility direction to the prompt took keyboard accessibility from 80 percent violations to zero and more than tripled the success rate on focus visibility.
Two conclusions follow, and design leaders should hold both. Telling the tool what you need works, dramatically, and most teams are not doing it. But 19 percent remains a substantial failure rate, so better prompting raises the floor without removing the need for review.
Supply the full constraint set and the effect compounds: approved design system components and their states, brand standards, accessibility requirements, content guidelines, device and technical limitations, validated user needs, business rules, and the interface states that must exist.
Nielsen Norman Group has begun calling this practice UX-context design, though their own framing is that experiments "suggest" curated context improves output while "important questions remain." The W4A measurement is the harder evidence, and it is the one worth citing.
Real content early, not lorem ipsum
Prototypes built on placeholder text fail in a predictable way. The layout works beautifully with a 40-character product name and breaks when the real name runs 90 characters with a trademark symbol. The error message that read cleanly in the mockup turns out to be a three-sentence legal disclosure. The form was elegant until the compliance-required consent copy arrived.
Generating realistic content early catches those structural problems while they are still cheap to fix. That means plausible long and short variants, genuine error scenarios, the actual regulatory language, and localized examples that expand the way German expands.
Content strategists still review the material for accuracy, tone, reading level, and brand fit. The point is that the prototype gets stress-tested against reality before anyone has invested in building the reality.
Review before you spend on participants
Customer testing is the most expensive validation you have. Spending a session watching someone discover an inconsistent label wastes a participant.
A structured pre-test review, whether AI-assisted or not, should catch the obvious problems first:
- Missing interface states
- Inconsistent terminology across screens
- Unclear calls to action
- Gaps between journey stages
- Likely accessibility failures
- Conflicts with design system standards
- Scenarios the prototype simply does not handle
Clear those, and the session can spend its full hour on the questions only a human can answer. Does this actually address the need? Does the customer understand what happens next? Do they trust it enough to continue?
Those are the questions worth paying a participant for.
The judgment that stays with you
Generation expands the number of options a team can consider. It cannot tell you which one is right, because rightness here is a judgment about people, business constraints, technical reality, and organizational appetite, all held at once.
The risk in this category has little to do with AI producing bad journeys and prototypes. The risk is that it produces a great many acceptable ones, very quickly, and a team mistakes that volume for progress. You can now generate forty journey variants and know precisely as much about your customers as you did on Monday.
Understanding still comes from talking to people. The tools just clear more room to do it.
![]() | Building the AI-enabled experience design workflow for in-house teams |
Sources
- Alexandra-Elena Guriță and Radu-Daniel Vatavu, "When LLM-Generated Code Perpetuates User Interface Accessibility Barriers, How Can We Break the Cycle?" Proceedings of the 22nd International Web for All Conference (W4A '25), April 2025. https://dl.acm.org/doi/10.1145/3744257.3744266
- Huei-Hsin Wang and Megan Brown, "Good from Afar, But Far from Good: AI Prototyping in Real Design Contexts," Nielsen Norman Group, October 24, 2025. https://www.nngroup.com/articles/ai-prototyping/
- Tony Alicea, "UX-Context Design: Using UX Knowledge to Inform AI-Generated Design," Nielsen Norman Group, July 24, 2026. https://www.nngroup.com/articles/ux-context-design/


