Error Recovery Is Product Strategy

Undo, comparison, preservation, escalation, and graceful recovery determine whether one bad result becomes a minor interruption or permanent mistrust.

The Field Guide

Methods and tools to design AI products people trust and keep using.

Read Me Because

Users don’t decide whether to trust a product only when it

Explore this concept

A smooth demo can make almost any product feel trustworthy. Then the user does real work.

One wrong action can turn trust into cleanup work. What broke, and what still holds? Can the person undo the change? Who takes over when the system can’t? Those answers stay hidden until something fails.

The happy path hides the contract

The real contract appears in those moments. Does the product admit what happened? Preserve completed work? Explain the consequence? Offer a safe route forward? Let the user restore the prior state?

Or does it show a generic error, erase the work, and send the user back to the beginning?

Recovery isn’t cleanup around the product. It’s part of the value proposition.

Failure changes the relationship

A user can forgive a system for being imperfect. It’s harder to forgive a system that hides uncertainty, loses work, or makes the user absorb the cost of its mistake. The mistake may be technical. The recovery is emotional.

“You approved this action” protects the product. “This action was approved, and here’s what changed, who saw it, and how to restore the previous version” helps the user.

AI raises the stakes because failure is often ambiguous. The system may return something fluent but unsupported. It may finish most of a task and omit one part. It may act correctly against the wrong interpretation.

Recovery begins with making the failure legible. Say what happened in the language of the work: “I updated the internal note, but didn’t send the customer reply.” Separate what completed from what didn’t. Name the missing source or permission. Mark the weak claims. Show whether external state changed. Honesty shortens the path back to trust.

Users take more reasonable risks when they know how to get back

A mistake with a visible repair path is survivable. A mistake with no obvious repair path becomes evidence that the AI is dangerous. Design recovery before the error.

Fill-in card

Safety net

Let users see how to get back.

Preview
what the user sees before a consequential action.
Reversibility
whether the change can be undone—fully, until published, for a limited time, partially, only by manual recovery, or not at all.
Diff
what changed, kept where the user can compare it.
Undo window
how long the action can be undone, kept visible.
Checkpoints
the versions the user can return to.
Scope
what the change touches, made explicit.
Confirmation
the stronger confirmation required when recovery is limited or impossible.

Reversibility isn’t only a post-failure feature. It’s pre-action confidence.

If recovery requires detective work, users stop experimenting. And undo is uneven: an edited file can be reverted; a sent email can’t. Users stop using the parts they can’t undo.

When users can see the safety net, they’re more willing to delegate.

Recovery should narrow the problem, not restart the job

The worst recovery path throws away valid work because one part failed.

If an internal research library becomes unavailable after twelve findings have been saved, keep the findings. Explain the blocked route. Offer public sources and ask before rerouting. If a workflow fails at step four, show the first three completed steps and what the next attempt will reuse.

Fill-in card

Recovery path

Narrow the problem instead of restarting the job.

Valid work
what stays valid when one part fails, kept rather than thrown away. Mark which inputs, drafts, records, and decisions still hold; which ones are stale; and which ones were never committed.
Failed part
what failed or became blocked, explained.
Alternative
the route to offer instead, asked about before rerouting. A blocked route isn’t permission to loosen the rules that made the work safe.
Next attempt
what the next attempt will reuse.

Example

Recovery path: the unavailable research library

An internal research library becomes unavailable after twelve findings have been saved. The recovery path preserves those findings and asks before it switches to public sources.

  • Valid work: the twelve findings already saved. Keep the findings.
  • Failed part: the internal research library, which became unavailable. Explain the blocked route.
  • Alternative: public sources, offered as an alternative. Ask before rerouting.
  • Next attempt: Not known yet: What will the next attempt reuse beyond the twelve saved findings?

The same principle—narrow the problem—applies to human handoff. When the system escalates, the person receiving the work should get the trigger, context, attempts, evidence, uncertainty, and current state. “Contact support” with no context is abandonment dressed as recovery.

A correction’s scope is a decision

A correction can fix the current output, the current object, the project, the user’s preference, or the system rule. Those are different decisions.

Ask or show the scope. “Change for this incident” shouldn’t become a permanent default. A global policy correction shouldn’t remain trapped in one run.

Let the user set a boundary, adjust a rule, mark an exception, or change the workflow. Recovery creates learning only when the product knows what the correction means.

A recovery review

For every high-value workflow:

0 of 8 done

Products don’t earn durable trust by never failing. They earn it by making failure survivable, honest, and smaller than the job the user came to complete.