A team tests its AI product before launch. The demo goes smoothly. In usability sessions, people complete the designed flow and praise the prototype.
Then a busy afternoon arrives, and the old system stays open as insurance. Does the result hold up when the source is stale and the user is late? Does anyone know what needs review? Can the work survive the next step? Demos are built to avoid those conditions.
The demo is a special occasion
Demos remove friction on purpose. The data is clean. The path is known. The presenter understands the product. The story begins at the most flattering moment and ends before maintenance, ambiguity, or recovery appears.
On a busy afternoon, the user is late. The source is stale. The customer record is incomplete. Two priorities conflict. The person who understands the workaround is out. The AI is moderately confident, not certain. Someone needs an answer before the system has every input it would prefer.
A product that works in a calm demo can fail on a normal day of use. That’s where it earns or loses its place.
The situation needs validating, not the artifact
Standard usability testing is designed to catch friction. Can the user find the button? Can they complete the flow without getting stuck? Useful, but insufficient for products making recommendations or acting on a user’s behalf.
Situational validation asks whether the product holds up inside the trigger and constraints of the real job.
Recruit someone who is currently trying to get the job done. Recruit by demographics and hand them a scenario, and they’re performing. Give them a believable account, document, incident, or decision. Include the awkward details. Let the system encounter missing context. Make the user explain the result to another person. Watch what they check, ignore, correct, and refuse to delegate.
The question isn’t only “Can they use it?” It’s “Can they rely on it here?”
The ordinary-day version needs fidelity in the situation and the response
An ordinary-day prototype doesn’t need a complete backend. In a Wizard-of-Oz simulation, a person manually provides the system’s behavior. That can make the behavior feel real enough to test. What needs fidelity is the situation and the response—not every implementation detail.
Fill-in card
Ordinary-day prototype
Build the trigger and constraints of the real job into the prototype.
- Trigger
- the actual trigger that creates urgency.
- Data quality
- realistic data quality, including gaps and contradictions.
- Tools and people
- the tools and people the user normally relies on.
- Pressure
- time pressure and competing work.
- Consequence
- a consequence if the result is wrong.
- Evaluate, defend, or hand off
- the need to evaluate, defend, or hand off the result.
Example
The day-to-day scene, as a prototype
Someone needs an answer before the system has every input it would prefer.
- Trigger: someone needs an answer before the system has every input it would prefer.
- Data quality: the source is stale, and the customer record is incomplete.
- Tools and people: the person who understands the workaround is out.
- Pressure: the user is late, and two priorities conflict.
- Consequence: Not known yet: What happens if the result is wrong?
- Evaluate, defend, or hand off: the AI is moderately confident, not certain, and the user explains the result to another person.
Fast output can hide burden migration
AI products can appear fast while moving labor into hidden places.
The system generates the brief in seconds, but the user spends twenty minutes verifying it. It routes the ticket automatically, but the rep has to reconstruct why. It suggests the next action, but the manager opens four sources before feeling safe enough to approve.
Measure the whole loop. Time to generated output is rarely the same as time to accepted result.
Look for repeated checking, corrections that don’t persist, uncertainty about scope, and moments where the user takes the task back. That isn’t relief. It’s the same labor with a friendlier interface.
The reaction is the evidence
People can sincerely praise a prototype and still avoid relying on it.
Behavior is harder to fake. Do they use the result in the next step? Do they explain it in their own words? Do they keep the old system open as insurance? Do they ask for another case? Do they hesitate before an irreversible action? Do they trust too quickly when they shouldn’t? Does the task technically succeed while they say, “I’d probably double-check this before I sent it”?
Both under-trust and over-trust are design findings.
A successful product helps confidence match capability. The user knows what can be accepted quickly, what needs review, and when a human should take over.
An ordinary-day test plan
Before shipping:
0 of 8 done
If the product only works when the data is perfect, the user is patient, and a founder narrates the value, it doesn’t work yet.