The question used to be “How do we add AI?” Now it’s “How much can our agents produce?” One prompt becomes a fleet—agents generating features, flows, and copy in parallel, faster than any team could review.
That’s the promise of scaling agentically. It’s also the danger. The risk was never technical failure. It’s relevance failure—shipping things that work but answer the wrong question—and agents reproduce that failure at volume.
When you scale agentically, slop scales too
Slop is the output everyone instantly recognizes as AI: generic, overconfident, vaguely polished, and empty at the center. A single piece of slop earns a click and gets forgotten. A scaled fleet of it teaches your users, at scale, that you don’t understand them.
When every team has the same models and the same agents, understanding the question is the only durable advantage.
Scaling amplifies whatever you point it at. Aim agents at assumptions and you mass-produce slop. Scaling agentically by design means aiming them at the job—on purpose, every time.
GOOD Iteration chases better outcomes, not more options
GOOD Iteration is our way of using generative tools without letting the work become generic, random, or feature-led. We don’t use AI to produce more screens just to have more options. We use it to push toward a better outcome faster. That means every round of exploration is anchored to the job, the intended shift in user behavior, and the specific relief or progress the product needs to create.
It validates the JTBD, scalability, and agentic feasibility through research, not assumptions—before agents build at scale—so what ships is meaningfully better and delivers value from the first use.
Fill-in card
The three questions
Every iteration must answer these before it scales.
- JTBD
- Is this the real job? Does it create the relief and progress the user is actually trying to make—not the feature we assumed they wanted?
- Scalability
- Does the value hold when it scales? The experience has to stay sharp when a fleet of agents multiplies output across users and volume.
- Agentic AI
- Can agents actually build, run, and sustain it? The work can be delegated to agents and still deliver on real intent.
If a round can’t pass all three with evidence, it isn’t ready to scale.
Scaling by design starts from the struggling moment
The old loop—build, ship, measure, hope—breaks when a fleet of agents can build almost anything overnight.
Fill-in table
Scaling by accident vs. scaling by design
Check which way a round is being scaled.
| What changes | Scaling by accident | Scaling by design |
|---|---|---|
| Agents start from | Internal brainstorming, opinions, and the loudest person in the room | The user’s real job-to-be-done and the struggling moment that triggered the search |
| Sequence | Generate and ship at volume, then measure for success after launch | Rapidly prototype and validate with users before agents build at scale |
| Agents are used to | Mass-produce screens and options—motion mistaken for progress | Generate embodiments of relief and accelerate toward a clearer outcome |
| Governed by | Vibes, prompts, and whatever’s in the agent’s training data | A shared constitution—the job spec, trigger, frame, and language guide every agent reads |
| Validated against | Nothing—assumptions stand in for proof until launch | The job, scalability, and agentic feasibility, each proven with evidence before scale |
| Result | Slop, scaled: fast, forgettable, trust-eroding—at volume | Stickiness, scaled: relief users recognize, return to, and rely on |
Encode the struggling moment first. Before you point a single agent at the work, define what made someone look for a better way, what they’re trying to escape, what relief looks like, and what would make a new path feel safe enough to try. Agents inherit your aim—give them the job, not just the task.
Go wide with agents before you go deep, then generate relief, not option-sprawl. Put prototypes in front of real users in the situation that triggers the job. If the intent doesn’t land in a prototype, it won’t land in production, and you’ve just stopped agents from scaling the wrong thing.
Then govern with a constitution. Capture the job spec, the trigger scene, the positioning frame, and the language guide as structured files every agent reads. “Build a collaboration feature” scales slop. The same prompt plus the job spec scales relief.
Example
The three questions for “Build a collaboration feature”
A team is about to point its agents at the prompt “Build a collaboration feature.” Before the fleet runs, the round goes through the three questions.
- JTBD: Not known yet: What made someone look for a better way, what are they trying to escape, and what would relief look like?
- Scalability: Not known yet: Does the value hold when agents multiply the output across users and volume?
- Agentic AI: Not known yet: Can agents build, run, and sustain it and still deliver on real intent?
No question passes with evidence, so the round isn’t ready to scale. The next move is to encode the struggling moment and add the job spec to the prompt.
Chat produces answers. Work changes things.
A blank prompt box is a powerful doorway. It’s also a poor place for work to live. Plain chat buries the artifact in the thread, so every improvement becomes another message and another full response the user has to copy, compare, and mentally merge.
Generative outcome-oriented design closes the gap between the answer and the product. The AI doesn’t merely respond. It creates or changes a meaningful object in the user’s world.
Return the result where the job can continue. Notes belong with the meeting record. A customer summary belongs in the customer record, not marooned in a chatbot transcript. Structured objects give judgment somewhere to land, and settled choices remain settled.
The conversation becomes the steering layer. The object becomes the product.
Prove it, then let the fleet run
The real risk of this moment isn’t that you fail to use agents. Everyone will. The real risk is that you scale a fleet of them pointed at shallow assumptions while someone else understands the job better than you do.
Before a round scales:
0 of 7 done
Prove the JTBD, prove it scales, prove agents can build it—then let the fleet run.