Skip to content

Jul 3, 2026 · 16 min read

The Mental-Model Reset: Designing Systems That Decide

Moving from deterministic screens to probabilistic behavior is the biggest design shift since mobile. A before/after worksheet to reorient your team’s assumptions.

We’ve hit the wall in the traditional design review.

We’ve been working through what it means to bring agents into Mural . Agents that can create, synthesize, and connect to the outside world. Agents that can churn through buckets of workshop data and hand you back an analysis. Normal review, normal cadence, and then the most normal question in our profession: Can we see the screens?

And the answer is no longer “yes”. The work may be done, but when an agent is composing the output itself, deciding at runtime what an analysis should look like for this user with this data, there is no screen to show. There’s a range of screens. A distribution of screens. The artifact we’ve organized our entire craft around — the designed, reviewed, approved screen…

…has stopped being the unit of work.

Agentic UX design reviews have left us with a question we’re still chewing on: how do we design in this mindset at platform scale?

How do we transition our thinking, not for one deeply thoughtful flow or feature, but for a whole system where agents can build things we’ve never seen?

I want to argue that before your team learns a single tactic for designing AI products; before the pattern libraries, before the chat guidelines, before the trust affordances, they need to break through the same wall we are working through. The mental model has to break first.

Because what’s changing isn’t the toolkit. It’s the entire design process.

Everyone wants to skip this part

Here’s where I see most design orgs right now, (including mine, which is in transition): they’ve treated AI as a feature. Somebody added a sparkle icon, a chat panel slid in from the right, the deck went to the exec team, done. The designers on those teams are still operating with fully deterministic instincts — enumerate the states, spec the flows, polish the artifacts, but apply them to a system that doesn’t hold still.

It mostly works, for a while, because a chat panel is a deterministic frame around probabilistic content. The frame behaves. And so nobody’s mental model gets challenged until the day the product roadmap says the word agent, and suddenly the thing inside the frame is taking actions, composing interfaces, making calls your team never reviewed.

Maggie Appleton called the chat-panel reflex what it is, back in 2023, in her Language Model Sketchbook:

But it’s also the lazy solution.
Maggie Appleton

Her point was that chat is the obvious tip of the iceberg, the interface you reach for when you haven’t yet re-thought what the material can do. Her sketches of background daemons that critique your writing as you work, models used as “epistemic rubber ducks,” are what it looks like when someone has actually made the mental shift and is designing from the new assumptions rather than the old ones.

You can’t get your team to that kind of thinking by teaching tactics. Tactics layer onto whatever mental model is already installed. If the installed model is deterministic screens, every tactic gets bent back into a screen.

So name the reset first. Here’s what it looks like:

No, this is not another “mobile moment”

The comparison every leader reaches for is mobile. We survived responsive design, we survived native apps, we’ll survive this. It’s a comforting analogy, and I think it’s wrong, and it’s worth being precise about why it’s wrong, because the analogy is exactly what lets teams underestimate the reset.

Mobile changed the constraints: screen size, input method, context of use. It was a hard decade. But through all of it, one thing never moved — the button you shipped did the same thing every time anyone tapped it. Your design was a contract, and the software honored it identically for every user, every session. Determinism survived mobile untouched.

Jakob Nielsen’s historical framing is the clearest I’ve found. In AI: First New UI Paradigm in 60 Years, he counts exactly three UI paradigms since 1945: batch processing, then command-based interaction from about 1964. He’s explicit that command lines, GUIs, and every smartphone platform are all the same second paradigm. Mac, Windows, iOS: one continuous era of “user issues a command, machine executes it, repeat.” The third paradigm, what he calls intent-based outcome specification, is the first genuinely new one in sixty years: the user states the outcome they want, and the system decides how to get there. As Nielsen puts it, this “completely reverses the locus of control.”

Pause here for a moment.

In Nielsen’s accounting, the mobile revolution, the one that restructured our teams, our tools, our titles, doesn’t even register as a paradigm change. It was a form factor change within the same paradigm.

Retrospectively, this rings true.

Interface paradigms of computing are batch processing, command-based interaction, and intent-based outcome specification.
GUIs and mobile were the same paradigm and this shift is categorically different Jakob Nielsen

The practical consequence for a design org: when the user stops specifying steps, the designer stops specifying steps too. Our entire artifact chain — the flows, wireframes, redlines, state diagrams, all comprise a technology for specifying steps.

That’s the thing that just went probabilistic.

The material that misbehaves (on purpose)

The deepest assumption in a screen designer’s head is so deep it doesn’t feel like an assumption: same input, same output. Every QA plan, every acceptance criterion, every “can I see the error state?” in every crit you’ve ever run assumes it.

A model-backed system breaks it by design. The output is a draw from a distribution: usually good, sometimes great, occasionally bizarre, and no amount of craft on your side makes it stop being a distribution.

Josh Clark and Veronika Kindred’s Sentient Design is the most complete treatment of this I’ve read, and its central move is the right one: treat AI as a design material, with a grain and a texture and characteristic weaknesses, the way HTML or ink or steel have weaknesses.

In Clark’s talks, the material properties are echoed: the system is probabilistic, not deterministic; it’s better at interpreting and transforming than at knowing. Designing with it means defensive design, accommodating failure and uncertainty as a first-class part of the work, not as an edge case cleanup pass.

I’ve found the fastest way to make this land with a team is not to explain it but to show the spread. Run the same prompt in your favorite UI generation tool twenty times and pin the outputs side by side. The reaction goes discomfort, then laughter, then somebody asking the real question, “So which one of these did we design?”

None of them…
…All of them?

You designed the possibility space — and if that sentence feels uncomfortable, good.

That’s the mental-model reset happening.

Clark and Kindred call the far end of this radically adaptive interfaces, experiences that reshape their content, structure, and interaction in real time against the user’s intent. NN/g’s less-imaginative term is generative UI, and their conclusion matches what we ran into at Mural almost word for word: when the interface is assembled at runtime, designers need to shift to outcome-oriented design, defining the goals, constraints, and rules the generation must operate within, rather than the discrete elements themselves.

Triangle diagram of Sentient Design experiences across three attributes: grounded, interoperable, and radically adaptive
There’s an emerging shared map of AI-mediated experience types beyond chat Josh Clark and Veronika Kindred

That’s the job now, at least for these surfaces. Less playwright, more improv director. A playwright writes every line and the actors deliver it verbatim, eight shows a week, identical. An improv director never scripts a line; she designs the scene: who’s on stage, what they know, what the rules of the game are, what’s out of bounds. Then the performance happens, differently, every single night.

We are all improv directors now. Most of us are still grading ourselves on how well we write scripts.

Your user got a promotion

There’s another thing that breaks, and it’s the one that turns this from a craft problem into a leadership problem: these systems act.

Not render. Act.

They send the email, run the code, reorganize the data. The industry noticed this in 2023 and named it “agentic,” but the design thinking is older: Christopher Noessel mapped this territory in Designing Agentive Technology back in 2017, studying systems that do things for you while your attention is elsewhere, and he’s been explicit recently that today’s agentic AI is the same discipline with a vastly more capable engine.

Noessel’s phrase for the shift is that agentive tech means “giving users a promotion”: the user moves from operator to manager of the work. And a manager needs an entirely different interface than an operator does. Working from his framework, the design surface stops being “the flow” and becomes a set of questions almost none of our existing patterns answer:

  • Setup and tuning: how does a user express preferences and constraints to something that will act unsupervised?
  • Visibility: how does the system expose its plan before, during, and after it acts?
  • Handoff and takeback: how does the user grab the controls mid-task, and how does the agent gracefully return them?
  • Decommissioning: how does someone wind down an agent they no longer trust or need?

Read that list again and notice: every one of those is a relationship problem that just happens to manifest in screens. That’s the promotion your design team is getting, whether they asked for it or not — the same promotion the user got.

But the pendulum can’t swing all the way the other direction. Deciding is not binary, and the mental model reset is not “make everything adaptive.”

Determinism is now a design decision instead of a default, which means somebody has to actively choose where the product stays fixed, predictable, boring. The primary toolbar should probably behave the same way every single time. The billing flow, absolutely. Guardrails, undo, confirmation on consequential actions. Google’s People + AI Guidebook has an entire chapter on errors and graceful failure that reads like it was written for this moment, and its core stance is the right default: assume the system will be wrong, and design the recovery path with the same care you design the happy path.

Choosing where the product does not decide is as much a part of the new mental model as accepting where it does.

Assumptions to unlearn

This is the before/after worksheet: one column of deterministic instinct, one column of what replaces it. Run your team through it against a real surface (instructions in the next section), not as an abstract reading exercise.

# The assumption to unlearn What’s true now
1 The screen I ship is the screen the user sees. The screen may be composed at runtime. You ship the rules of composition — goals, constraints, components the system can use, and what it may never do.
2 Same input, same output. Same input, distribution of outputs. Quality is a property of the distribution, not of any single render.
3 QA enumerates the states and signs off. You can’t enumerate the states. You sample behavior — run it many times, score the spread, set thresholds. (Your eng partners call these evals. Learn the word.)
4 Errors are bugs; the target is zero. Some error rate is a property of the material. The design question moves from “prevent it” to “how does the user detect and recover, and what’s the cost when they don’t?”
5 Nothing happens unless the user does it. The system acts while attention is elsewhere. Setup, plan visibility, handoff, takeback, and decommissioning are now core flows, not settings-page afterthoughts.
6 Edge cases are rare states we document at the end. There is no edge — the happy path drifts continuously into failure. Confidence thresholds and fallbacks are the center of the design, not the margins.
7 A crit reviews the artifact. A true crit happens after it is built, and reviews behavior over many runs. If in a pre-build crit, the work is reviewed as directional.
8 Copy is written once, by us. The system generates language on the fly. You design the voice constraints: personality, tone boundaries, forbidden claims, required disclosures, but not the sentences.
9 Trust comes from polish. Polish inflates expectations that a probabilistic system will sometimes miss. Trust comes from calibrated expectations and visible reasoning. Over-promising via sheen is a design defect.
10 Done means shipped. Behavior drifts — models update, data shifts, prompts age. Done means monitored, with someone owning the watch.

Two notes on using it. First, the left column is a list of correct ideas for the last thirty or forty years of software, held by your best people most firmly, because their careers were built on being right about them. Treat the unlearning with respect or it won’t happen. Second, rows 5 and 10 are organizational, not craft. They will create new questions about ownership and staffing that a worksheet can’t answer. Capture those and take them offline. Don’t let them derail the session.

We’re all resetting ourselves

As a profession, we’re all working through this reset. I’m in the midst of it with my org, and here’s my plan: This is a 60 to 90-minute working session plus about an hour of prep. I’m writing it as instructions you can share.

  1. Pick one real surface. A live AI feature in your product, or the roadmap item closest to shipping. Decision point: if the only AI you’ve shipped is a chat panel, use the chat panel — but review what happens inside and after it, not the frame around it. Do not use a hypothetical. Hypotheticals let everyone stay comfortable.

  2. Prep: run it twenty times. If the feature is live: Before the session, have one designer run the same realistic input through the surface twenty times and screenshot every output. (If you’re pre-ship, use the staging build or even raw model outputs — fidelity matters less than variance.) Pin all twenty on a board, unlabeled. Trade-off: twenty runs of one input beats five runs of four inputs. You’re teaching the team to see a distribution, and a distribution needs depth to be visible.

  3. Open with the spread. First ten minutes: everyone looks at the twenty outputs and individually sorts them into three piles — great, acceptable, not okay (a product like Mural is great for this). No discussion yet. If the twenty outputs are essentially identical, you picked a surface that isn’t actually probabilistic — swap it now, or the whole session becomes theater.

  4. Fight about the piles. Next thirty minutes: compare sorts. The disagreements are the entire point . They expose that your team has no shared quality bar for a distribution, because they’ve never needed one. Write down every disagreement as a draft rubric line (“an analysis that omits the data source is not okay, even if the summary is good”). You will leave this session holding your first informal eval spec, and most teams don’t realize that’s what they built until you tell them.

  5. Run the table above against what they just saw. Next twenty minutes: go through the ten rows above and, for each, ask “where did we watch this assumption break in the last hour?” Rows that connect to a specific pinned output will actually stick. Rows that don’t connect, skip without guilt.

  6. Rewrite three crit questions. Last twenty minutes, and this is the piece that makes the reset durable instead of a nice workshop: pick three questions your team asks in every design review and rewrite them for probabilistic surfaces. Ours: “Can I see the error state?” became “What’s the worst plausible output, and what happens in the next ten seconds after it appears?” “Why this layout?” became “What rules generated this layout, and what else could they generate?” “Is the copy final?” became “What are the boundaries on what the system is allowed to say here?” Trade-off: rewriting the crit changes behavior faster than any training deck, but it only works if you, the leader, actually ask the new questions in the next review. The first time you ask “can I see the screen?” out of habit, the old model reboots itself.

  7. Assign the two follow-ups. One: someone turns the pile-sorting rubric into a real conversation with your engineering partner about evals — the draft you made in step 4 is the input. Two: someone maps where the product should stay deterministic — the surfaces where adapting would be a bug, not a feature. Both have owners and dates before anyone leaves the room, or neither happens.

That’s it. One Mural board, one enlightening (and possibly fun) meeting. What you’re buying is a team that stops bending AI work back into screen-shaped thinking — which is the prerequisite for every tactical skill you’ll want to build next.

Tools & resources

The screens we can’t see (yet)

Back to the design review. We’re in a world where there is nothing to mock up. What teams will eventually bring back instead are a set of rules: What an agent-built analysis must always include, what it may never claim, which pieces of the interface stay fixed no matter what the agent decides, and what the user grabs when they want the controls back.

These will be the least screen-like design deliverables we’ve approved as leaders. And I’m increasingly convinced this output will be the most designed — because every rule in it is a decision about someone’s experience, made deliberately, ahead of time, exactly like our craft has always demanded.

Like jazz, or improv, the performance changes every night now. Embrace the riff. Design the stage.