Product teams are good at reviewing screens. The hierarchy is clear. The layout is polished. The empty state has the right tone. The prototype flows.
None of that review shows what the AI does when it’s unsure. What will people learn to check twice? Which correction will they have to repeat? When will they stop trusting the result? The review covered only what the screen reveals.
The screen is the visible part
What does the system do when it isn’t sure? Does it ask or guess? What does it remember? When does it recommend, prepare, or act? How does it expose the evidence behind a result? What happens when the user corrects it? Can someone recover without starting over?
Those decisions determine the relationship more than the color of the prompt box.
A product can look calm and behave chaotically. It can sound friendly and repeatedly ignore the user’s intent. It can have a perfect design system and no consistent definition of approval, confidence, or undo.
If the product doesn’t define the relationship, the user has to. That work may be manageable once. It becomes exhausting when the user has to do it every time the stakes change.
Each interaction teaches the user what kind of partner the system is
People build a mental model of intelligent systems through repeated encounters.
They learn whether the product tells the truth about what it knows. Whether “approve” means once or forever. Whether a correction changes only this output or future behavior. Whether a progress indicator reflects real work. Whether the AI admits when a source is missing.
In deterministic software, consistency often means the same component behaves the same way. In AI products, consistency also means the same promise behaves the same way. “Source-backed” should mean something stable. “Draft” should remain provisional. “Reversible” should come with a real recovery path. “Remember this” should have visible scope.
The relationship has a lifecycle
Onboarding a static product teaches the interface. Onboarding an agentic product teaches the relationship.
At first, the user needs recognition. The product should understand the situation without demanding a long briefing.
Then the user needs relief. The first result should remove a meaningful burden while staying easy to evaluate.
As use repeats, the product earns permission. The AI may move from suggesting to preparing, and eventually to acting within clear boundaries. New levels of consequence, context, or autonomy require new evidence.
A wrong click in normal software is a user’s mistake. A wrong agent action feels like the product failing the relationship. So when something goes wrong, the system needs to preserve the relationship. It should show what happened, keep useful work, provide a path back, and learn from the correction at the right scope.
Over time, memory and autonomy become valuable only if the user can see and change them. Hidden adaptation feels less like intelligence and more like the rules changing behind the user’s back.
The choice not to act is a product behavior
Good AI design is partly restraint.
The system doesn’t need to generate whenever it can. It shouldn’t rewrite settled work, reopen decisions, or turn every completed task into a new suggestion. It shouldn’t perform confidence when the evidence is weak. It shouldn’t take control of the part of the job the user needs to own.
In some moments, the AI should suggest and stay quiet. In others, it should show what it used before it recommends anything. In others, it should prepare the work but stop before the commitment point.
A humane system knows when to ask, when to wait, when to narrow scope, when to hand back, and when the most useful response is a quiet confirmation that the work is done.
A relationship critique reviews behavior, not just artifacts
A screen critique asks whether the page communicates. A relationship critique examines the system’s behavior at each decision point.
Fill-in card
Relationship critique
Ask each question of the system’s behavior, not just the page.
- Belief
- What does the system believe is happening?
- Authority
- What authority does it have right now?
- Uncertainty
- What uncertainty is visible?
- Evidence
- What evidence can the user inspect?
- Change, reject, undo
- What can be changed, rejected, or undone?
- Memory
- What will the system remember from this interaction?
- Next decision
- Who owns the next decision?
- If the AI is wrong
- What happens if the AI is wrong?
Prototype those questions inside realistic situations. Watch whether people relax into appropriate trust or start babysitting the system. If people start babysitting the system, the product has added work.
Example
Relationship critique: manual or autopilot
A finance operations manager wants help processing a month of company expenses. As soon as the manager connects the account, the product asks them to turn on autopilot. So they leave autopilot off. The product asked for the final behavior before giving the manager a way to try the first one.
- Belief: rules it inferred during setup.
- Authority: all or nothing. On autopilot, the AI will categorize transactions, merge duplicates, request missing receipts, and post expenses. Otherwise the manager keeps doing everything manually.
- Uncertainty: not visible. The manager hasn’t seen how the system handles split transactions, whether it can tell software from office supplies, which employees it’ll contact, or what it considers approved when a receipt conflicts with the card charge.
- Evidence: nothing the manager has seen before choosing.
- Change, reject, undo: decline, and keep doing everything manually.
- Memory: Not known yet: What will the system remember from this interaction?
- Next decision: the manager’s, forced as autopilot or manual.
- If the AI is wrong: Not known yet: What happens if the AI is wrong?
The design standard
The goal isn’t to make AI seem human. It’s to make its participation considerate, legible, and bounded.
For the relationship you’re designing:
0 of 6 done