Ease of Evaluation Is the New Ease of Use

In delegated systems, the user’s burden shifts from operating the interface to judging what the system hands back.

The Field Guide

Methods and tools to design AI products people trust and keep using.

Read Me Because

When software does more of the execution, the user’s hardest job

Explore this concept

The AI produced the report in thirty seconds instead of three hours. Then the user spends forty minutes checking claims, rebuilding context, comparing sources, correcting tone, and deciding what changed between versions.

Three hours became thirty seconds plus forty minutes of doubt. Which claims are sourced? What did the system assume? What changed since the last version? The hard part is no longer doing the work.

The complexity moves to evaluation

Traditional UX focused heavily on the gulf of execution. Can users find the control, understand the label, complete the steps, and make the software do what they intend?

AI can collapse that gulf. A person describes the outcome and the system handles the operations. No need to master menus, formulas, filters, or workflow builders.

But the complexity doesn’t disappear. It moves.

The user now has to evaluate a result they didn’t produce step by step. Was the intent understood? Are the sources sound? What assumptions slipped in? What changed? Is the output safe to use, send, publish, or automate?

The user is deciding whether the work is true enough, fitted enough, safe enough, and ownable enough to act on.

The new usability failure isn’t “I can’t make the computer do it.” It’s “I can’t tell whether what the computer did is right.”

Fast output can create slow work

Generation time is a seductive metric. Faster generation doesn’t help if the user’s burden is verification.

Review burden is often invisible because it happens after the celebrated moment of generation. It’s distributed across hesitation, side searches, repeated prompts, and quiet manual reconstruction.

If every automated step creates an equal inspection tax, the product hasn’t removed labor. It has relocated it.

The result should be inspectable

Ease of evaluation starts by structuring the output around the decisions the user needs to make. “Here’s the summary” versus “here’s the summary, with the three sources used to produce it.” “I rewrote the code” versus “I changed these three files and left the billing function alone.” Only the second version treats judgment as the user’s actual job.

Fill-in card

Make the result inspectable

Structure the output so the user can inspect it.

Sources
Show them at the claim they support, not in a pile at the bottom.
Observation and inference
Separate them.
Assumptions
Mark them.
What changed
Highlight it.
Previous version
Preserve it.
Confidence
Make it specific—confidence in the source, the claim, the plan, or the action.

Example

Make the result inspectable: a vendor agreement review

A procurement lead uploads a twelve-page vendor agreement and asks an AI assistant to identify clauses that differ from the company’s standard terms. The answer appears almost immediately: Three clauses require attention. They don’t know whether the assistant compared the agreement with the current policy, which passages produced the findings, or whether “require attention” means legally unusual, commercially unfavorable, or merely different from a template.

  • Sources: The passages that produced each finding, shown at the clause.
  • Observation and inference: Not known yet: Which findings are observations, and which are inferences?
  • Assumptions: Whether the assistant compared the agreement with the current policy.
  • What changed: Not known yet: Which clauses differ from the company’s standard terms?
  • Previous version: Not known yet: Is there an earlier version of the vendor agreement to compare?
  • Confidence: Not known yet: How confident is the assistant in each finding?

The instant answer hasn’t created relief. It has created another object to verify.

Don’t expose every internal step. Transparency asks whether the user can see more. Evaluability asks whether the user can judge better. Show the smallest amount of evidence that makes the next decision safe.

A reviewer scanning a routine draft may need a compact “source-backed” signal. Someone approving a consequential external action may need the full diff, rationale, risk, and rollback plan.

Evaluation should scale with stakes.

Judgment needs somewhere to land

“Human in the loop” is a comforting phrase, but it’s not a design until the loop gives the human something real to judge with.

Generated work should become an object the user can inspect and change.

A plan needs owners, dates, dependencies, assumptions, and unresolved questions. A recommendation needs evidence, alternatives, and a next action. An edited image needs before-and-after comparison and restore. A workflow change needs scope, preview, and undo.

A block of fluent text makes the user infer all of that structure. A designed object makes it visible.

The product should also preserve settled decisions. If the user approved the tone, selected the direction, or rejected an option, the next pass shouldn’t casually reopen it. Convergence reduces evaluation load by shrinking the number of live possibilities.

Accepted results are the better measure

Better metrics include:

  • Time from request to accepted result.
  • Percentage of output changed before use.
  • Number of sources opened for verification.
  • Repeated correction rate.
  • Undo and recovery success.
  • Escalation caused by missing context.
  • Willingness to delegate the same job again.
  • Difference between system confidence and user trust.

A product can produce more while performing worse on every one of these.

An evaluation audit

Choose one important AI result:

0 of 8 done