Generative Outcome-Oriented Design

Use generative tools and agents to push toward better outcomes faster, not more options, with every round anchored to the job, the intended behavioral shift, and the specific relief the product needs to create.

The Field Guide

Methods and tools to design AI products people trust and keep using.

Read Me Because

Generative outcome-oriented design is iterative design

Explore this concept

The question used to be “How do we add AI?” Now it’s “How much can our agents produce?” One prompt becomes a fleet—agents generating features, flows, and copy in parallel, faster than any team could review.

That’s the promise of scaling agentically. It’s also the danger. The risk was never technical failure. It’s relevance failure—shipping things that work but answer the wrong question—and agents reproduce that failure at volume.

When you scale agentically, slop scales too

Slop is the output everyone instantly recognizes as AI: generic, overconfident, vaguely polished, and empty at the center. A single piece of slop earns a click and gets forgotten. A scaled fleet of it teaches your users, at scale, that you don’t understand them.

When every team has the same models and the same agents, understanding the question is the only durable advantage.

Scaling amplifies whatever you point it at. Aim agents at assumptions and you mass-produce slop. Scaling agentically by design means aiming them at the job—on purpose, every time.

GOOD Iteration chases better outcomes, not more options

GOOD Iteration is our way of using generative tools without letting the work become generic, random, or feature-led. We don’t use AI to produce more screens just to have more options. We use it to push toward a better outcome faster. That means every round of exploration is anchored to the job, the intended shift in user behavior, and the specific relief or progress the product needs to create.

It validates the JTBD, scalability, and agentic feasibility through research, not assumptions—before agents build at scale—so what ships is meaningfully better and delivers value from the first use.

Fill-in card

The three questions

Every iteration must answer these before it scales.

JTBD
Is this the real job? Does it create the relief and progress the user is actually trying to make—not the feature we assumed they wanted?
Scalability
Does the value hold when it scales? The experience has to stay sharp when a fleet of agents multiplies output across users and volume.
Agentic AI
Can agents actually build, run, and sustain it? The work can be delegated to agents and still deliver on real intent.

If a round can’t pass all three with evidence, it isn’t ready to scale.

Scaling by design starts from the struggling moment

The old loop—build, ship, measure, hope—breaks when a fleet of agents can build almost anything overnight.

Fill-in table

Scaling by accident vs. scaling by design

Check which way a round is being scaled.

What changesScaling by accidentScaling by design
Agents start fromInternal brainstorming, opinions, and the loudest person in the roomThe user’s real job-to-be-done and the struggling moment that triggered the search
SequenceGenerate and ship at volume, then measure for success after launchRapidly prototype and validate with users before agents build at scale
Agents are used toMass-produce screens and options—motion mistaken for progressGenerate embodiments of relief and accelerate toward a clearer outcome
Governed byVibes, prompts, and whatever’s in the agent’s training dataA shared constitution—the job spec, trigger, frame, and language guide every agent reads
Validated againstNothing—assumptions stand in for proof until launchThe job, scalability, and agentic feasibility, each proven with evidence before scale
ResultSlop, scaled: fast, forgettable, trust-eroding—at volumeStickiness, scaled: relief users recognize, return to, and rely on

Encode the struggling moment first. Before you point a single agent at the work, define what made someone look for a better way, what they’re trying to escape, what relief looks like, and what would make a new path feel safe enough to try. Agents inherit your aim—give them the job, not just the task.

Go wide with agents before you go deep, then generate relief, not option-sprawl. Put prototypes in front of real users in the situation that triggers the job. If the intent doesn’t land in a prototype, it won’t land in production, and you’ve just stopped agents from scaling the wrong thing.

Then govern with a constitution. Capture the job spec, the trigger scene, the positioning frame, and the language guide as structured files every agent reads. “Build a collaboration feature” scales slop. The same prompt plus the job spec scales relief.

Example

The three questions for “Build a collaboration feature”

A team is about to point its agents at the prompt “Build a collaboration feature.” Before the fleet runs, the round goes through the three questions.

  • JTBD: Not known yet: What made someone look for a better way, what are they trying to escape, and what would relief look like?
  • Scalability: Not known yet: Does the value hold when agents multiply the output across users and volume?
  • Agentic AI: Not known yet: Can agents build, run, and sustain it and still deliver on real intent?

No question passes with evidence, so the round isn’t ready to scale. The next move is to encode the struggling moment and add the job spec to the prompt.

Chat produces answers. Work changes things.

A blank prompt box is a powerful doorway. It’s also a poor place for work to live. Plain chat buries the artifact in the thread, so every improvement becomes another message and another full response the user has to copy, compare, and mentally merge.

Generative outcome-oriented design closes the gap between the answer and the product. The AI doesn’t merely respond. It creates or changes a meaningful object in the user’s world.

Return the result where the job can continue. Notes belong with the meeting record. A customer summary belongs in the customer record, not marooned in a chatbot transcript. Structured objects give judgment somewhere to land, and settled choices remain settled.

The conversation becomes the steering layer. The object becomes the product.

Prove it, then let the fleet run

The real risk of this moment isn’t that you fail to use agents. Everyone will. The real risk is that you scale a fleet of them pointed at shallow assumptions while someone else understands the job better than you do.

Before a round scales:

0 of 7 done

Prove the JTBD, prove it scales, prove agents can build it—then let the fleet run.