A team using AI can generate polished interface concepts, onboarding copy, and workflows before lunch. The output looks ready to ship. Then a user brings a real request, and the AI misunderstands it, acts beyond its scope, or forgets something important. The team has produced more work, but the user now has more work to do.
That’s AI slop: work that’s faster to generate than to check, and it erodes trust. Did the system understand what the person was trying to get done? What did it change? Can they correct it without redoing the work? The fix starts with how the product behaves, not how it looks.
Slop is an interaction design problem
How did it get this bad? How did our tools make us so sloppy? Lobbed slop grenades aren’t crafted by hand. They’re born of low-effort button clicks, poor judgment, and simple ignorance. Production used to be slow, and that friction filtered out a lot of wrong ideas. Agents remove the friction. If the starting assumption is wrong, the agents don’t catch it. They multiply it.
Bad writing, generic copy, ugly AI visuals, and hallucinated facts are symptoms. The deeper failure is a system producing work with no validated direction behind it, so we treat slop as a product design problem, and it’s not a difficult mystery to solve. Slop enablers stem from the flows, form, and function of a product. Slop enablement is a design anti-pattern in the interaction design. But that doesn’t mean it should be ugly.
Visual design still counts. People form an opinion in milliseconds, and a careless interface can kill trust before the product gets a chance to prove itself. But visual polish is the opening handshake, not the whole relationship. A cleaner interface doesn’t help if the problem is uncertainty about scope, or if the user’s judgment is poor or misinformed. And yes, this is the job of product design.
Then there’s function. The deeper design work shows up after the click:
- Does the AI understand what the user is actually trying to accomplish?
- Does it know enough to act, or should it ask a question first?
- Can it show where an answer came from and how confident it is?
- Does it make a useful recommendation or dump ten options back on the user?
- Can someone edit, reject, undo, or recover without starting over?
- Does a correction improve the next result, or does the same mistake return five minutes later?
- When the system gets blocked, does it explain what happened and offer a safe path forward?
Those choices determine how AI works, responds, and feels. They’re design choices even when no one opens Figma to make them. Now read the list from the user’s side. Did it get what I meant? Should I double-check this before I send it? Will it make the same mistake tomorrow? In a market that decides on trust and judgment, those aren’t questions you want your users asking themselves. Each one they have to ask is work the product handed back.
A beautiful product that makes silent changes, invents certainty, and forces people to recheck everything is badly designed. A visually simple product that asks at the right moment, keeps its scope clear, and makes every action reversible can feel remarkably thoughtful.
The interface makes that thoughtfulness visible, but the product decisions behind it come first.
GOOD Iteration connects intent to desired outcomes
Agents amplify whatever you point them at. Point them at a validated job and they scale relief. Point them at assumptions and they mass-produce trust damage.
There’s a discipline behind this: GOOD Iteration, aka Generative Outcome Oriented Design. Build toward outcomes you’ve validated, not output you’ve assumed. Slop is what the absence of that discipline looks like once agents are involved.
The JTBD check: was the job validated before the agents ran? What struggling moment triggered the need, what progress is the user after, what old way are they trying to escape, what would make them feel understood rather than processed? Without that, agents inherit tasks instead of intent—and that’s where slop begins.
The scalability check: does the value hold as output multiplies? Not can the infrastructure handle the volume, but can trust withstand it. Does review burden shrink or grow, does trust deepen or flatten, do corrections become reusable, does the job survive as volume rises? A product can look scalable while scaling mistrust.
The agentic feasibility check: can the agent do the work without detaching from intent? Not just can it produce the output, but can it understand the job, use real evidence, stay inside boundaries, ask when the trigger is unclear, make its result easy to evaluate, and stop before uncertainty turns into confident damage.
And a big thing that keeps a fleet pointed the same way is shared judgment—the job spec, the trigger, the evidence standard, the boundaries, the definition of good, written where every agent inherits it. Slop is what happens when agents inherit tools but not judgment. GOOD Iteration doesn’t slow agents down. It gives their speed something true to attach to.
The agentic era rewards teams that move fast, but speed is about to stop being scarce. Every team will have agents. Every team will be able to generate, automate, and ship more than before. Production is becoming table stakes.
What stays scarce is direction. The differentiator won’t be how much a team can produce. It’ll be whether that production is anchored to validated user progress before the fleet runs.
Knurture’s process runs from human truth to system behavior
GOOD Iteration is how we put that into practice. We don’t start by asking what the AI can generate. We start with the moment a person needs help and the outcome they’re trying to reach. The process has nine steps, and each one answers a question the screens can’t answer on their own.
Sequence map
Knurture’s process
Walk one AI feature through all nine steps before anyone scales it.
- Find the real situation: name the status quo, the trigger, and the progress the person wants; a feature request isn’t a job.
- Define the first relief: name the smallest useful result that shows the product understands what someone is dealing with.
- Identify the objects and transitions: map what the person works on, what can change, what must stay stable, and who owns each move.
- Set the AI’s role and boundaries: for each action, decide whether the AI observes, suggests, prepares, acts with approval, or acts on its own, and what it may remember.
- Prototype the relationship, not just the screens: put a realistic trigger and a real object in front of people and watch what they do when they’re uncertain.
- Design every state where trust can break: cover working, partly done, missing context, low-confidence, blocked, contradicted, denied, and wrong.
- Test for relief, not throughput: track time to an accepted result, corrections, review effort, recovery, escalation, and delegation over time.
- Encode the judgment into the system: carry the reasoning into states, components, decision trees, interaction rules, and code.
- Ship one complete slice, then compound: release one instrumented journey a real user can finish, then fold what you learn back in.
1. Find the real situation
The brief names a requested feature, not necessarily the progress a person needs.
“Build an assistant.” “Summarize the account.” “Recommend the next action.” “Automate onboarding.” Those are feature requests. They don’t tell us what changed in the user’s day, what they’re trying to escape, what they’re afraid of getting wrong, or what progress would feel like.
We look at the status quo people have learned to tolerate, the trigger that makes it unacceptable, and the functional, emotional, and social progress they’re trying to make. We study workarounds, hesitation, repeated explanations, manual checks, and the moments where trust breaks.
That gives the AI a job to serve instead of a task to perform.
2. Define the first relief
AI products often chase a dramatic reveal. We’d rather find the first moment where the weight drops.
What’s the smallest useful result that helps someone think, “Okay, this understands what I’m dealing with”? It might be one correctly prioritized action, a credible first draft, a visible comparison, or a risky choice caught before it becomes a problem.
That first relief gives the product something concrete to design toward. It also keeps the team from confusing surprise with value. A result can be impressive once and still be useless in day-to-day use.
3. Identify the objects and transitions
Once the job is clear, we map the things involved in getting it done. Generative outcome-oriented design keeps the AI’s results attached to the objects they change, not buried in a chat thread.
Fill-in table
Objects and transitions
Answer these for every object the AI touches before you design a screen for it.
| Question | What to write down |
|---|---|
| What is the user looking at? | The object, named the way the user names it |
| What can change? | The fields and parts the AI or the user may alter |
| What needs to remain stable? | What stays fixed so the user can compare and trust |
| Which objects depend on one another? | The links that make one change ripple into another |
| What states can each object enter? | The states the object can enter and the transitions between them |
| What should be generated, suggested, compared, approved, or left alone? | Which parts may be generated, suggested, compared, or approved, and which must be left alone |
| Who owns each transition, and what’s the way back? | The owner, the evidence needed, and the undo when the cost of error is high |
Then we define the transitions. A support case can move from open to investigated to resolved. A recommendation can move from proposed to approved to acted on. A generated draft can move from provisional to edited to accepted. Each transition needs visible ownership, evidence, and a way back when the cost of error is high.
Example
Objects and transitions for an AI-improved project plan
| Question | Project plan |
|---|---|
| What is the user looking at? | A project-plan object, not a paragraph in chat history |
| What can change? | Owners, dates, dependencies, assumptions, and sources, one part at a time |
| What needs to remain stable? | The previous version, so the user can compare against it |
| Which objects depend on one another? | The dependencies the plan lists |
| What states can each object enter? | Provisional, edited, accepted |
| What should be generated, suggested, compared, approved, or left alone? | The AI suggests changes; the user compares them with the previous version and approves them; parts nobody asked to change are left alone |
| Who owns each transition, and what’s the way back? | The user moves the plan to accepted; choices that came from the system stay marked, and the previous version stays available |
The project plan is now a product people can inspect, change, and compare instead of a vague “AI feature.”
4. Set the AI’s role and boundaries
For every meaningful action, we decide whether the AI should observe, suggest, prepare, act with approval, or act on its own. We define what it can infer, what it must ask, what evidence it needs, and when it has to stop.
Autonomy can grow as trust grows. A user may do the task manually the first time, approve an AI-prepared version the next time, and allow the system to handle it later. The product should make that progression visible instead of taking control without showing it.
The boundary work also covers memory. What should the system remember? For how long? Is a one-time correction a preference, an exception, or a signal that the rule is wrong? Helpful memory preserves context. Bad memory turns one choice into a permanent assumption.
5. Prototype the relationship, not just the screens
A static mockup can show what an AI product looks like. It can’t prove how the relationship feels.
We build situational prototypes around realistic triggers and real objects. Sometimes the AI is simulated behind the scenes. Sometimes only one narrow path works. That’s enough to test the hard parts early: whether the system asks the right question, whether the result feels credible, whether the level of initiative feels appropriate, and whether someone can tell what to do next.
We’re not looking for compliments about the interface. We’re watching what people do when they’re uncertain. Do they pause? Re-read? Check another source? Undo the action? Correct the same thing twice? Hand the task back to themselves?
Those reactions tell us more than “I’d use this.”
6. Design every state where trust can break
The happy path is the easy part.
We design what happens when the AI is still working, only partly done, missing context, low-confidence, blocked by a source, contradicted by evidence, denied permission, or simply wrong. Each of those is a state the person has to be able to read. We decide how progress is shown, how partial work is preserved, how uncertainty is worded, and how the user recovers.
Microcopy—the short labels, instructions, and messages in an interface—is part of the behavior. So are timing, motion, hierarchy, defaults, approvals, before-and-after views, and the location of the undo control. Every one of those details can make the AI feel calm and legible or slippery and infuriating.
When evidence is uncertain, the system shows that uncertainty and gives the person a next move.
7. Test for relief, not throughput
Generation count, token volume, automation rate, and time saved can all look good while the experience gets worse. Volume isn’t a trust metric.
We look for signals that the burden actually moved off the user: time to an accepted result, correction rate, repeated corrections, review effort, successful recovery, appropriate escalation, and whether people choose to delegate more of the job over time.
If the AI produces ten times more work and creates ten times more checking, nothing was automated. The labor was relocated.
The question isn’t whether the user still has work to do. It’s whether the new work is better work. Better work is smaller. It has clearer boundaries. It makes the next step easier to see.
8. Encode the judgment into the system
A good prototype isn’t enough if the logic disappears during implementation.
We carry the reasoning into states, components, decision trees, interaction rules, voice guidance, edge cases, accessibility notes, instrumentation, and code. Figma and the frontend should describe the same product behavior. Humans and coding agents should be able to see not only what a component looks like, but when to use it, when not to use it, and what it does under pressure.
That shared judgment keeps the product coherent as more people and more agents contribute to it. Slop scales when correction stays stuck in the last output instead of becoming judgment for the next one. Corrections stop being one-off fixes and become better system behavior.
9. Ship one complete slice, then compound
We’d rather ship one honest, instrumented journey than a fleet of half-designed AI features.
A complete slice lets a real user enter with a real need, work with the AI, understand the result, recover from trouble, and leave with something useful. Once it’s in the world, we compare actual behavior with what we expected. Where did people stall? What did they check? What did they trust too quickly? What did they refuse to delegate?
Then we fold what we learn back into the objects, rules, components, and boundaries. The next slice gets faster without getting sloppier.
Humane AI doesn’t need to imitate a person
A humane AI product doesn’t need to pretend it’s human. It needs to respect the human using it.
A friendly avatar can be delightful. A conversational tone can lower the temperature. Neither one makes an AI humane. Humane AI protects a person’s attention, competence, agency, and time.
It doesn’t make someone repeat context the product already has. It doesn’t hide a guess behind a confident sentence. It doesn’t turn every completed task into a pitch for another task. It doesn’t take an irreversible action because the user clicked something once. It doesn’t force people to become full-time reviewers of work they supposedly delegated.
Instead, it:
- Meets the user inside a real situation, not at an empty prompt.
- Asks when intent is ambiguous and stays quiet when it isn’t.
- Takes the right amount of initiative for the stakes.
- Shows what it changed, what it used, and what still needs judgment.
- Keeps generated work editable, reversible, and attributable.
- Preserves useful context without becoming creepy or presumptuous.
- Makes progress visible without faking precision.
- Hands off cleanly when a person should take over.
- Learns from correction instead of repeating the same failure at scale.
- Finishes the job without manufacturing more work.
In practice, humane AI shows its evidence, leaves work editable, and hands off when someone needs to decide.
A quick slop test
Before scaling an AI experience:
0 of 8 done
If the answers to the checklist questions are vague while the screens are specific, generation has outrun design.
What Knurture is really designing
We design the visible interface, but we don’t stop there.
We design the intent model that helps the system understand what someone is trying to do and what outcome would count as progress. We design the objects and states that keep generated work grounded. We design the permission model that determines when the AI may act. We design the language that makes confidence and uncertainty legible. We design the recovery paths that keep one bad result from destroying trust. We design the system that carries those decisions from prototype into production.
These choices distinguish a product that adds AI from one designed for people to rely on.
That’s GOOD Iteration: build toward outcomes you’ve validated, not output you’ve assumed.
Generation can be fast, but the team still has to set its direction, boundary, and standard of care.
That judgment appears in the evidence, permissions, and recovery paths people use.
It reduces the checking and frustration that make generated output feel like slop.