· 4 min read

The Judgment Layer

AI can produce almost anything that looks finished. It cannot tell you whether it's right. Where a designer's judgment has to sit in that gap.

AI can hand you something that looks like a decision. Usually it isn't one. Design systems taught me the difference years before AI did: a raw colour is a description, a semantic token is a decision, and most of the frustration I see with AI-assisted work comes from people operating entirely in the first layer and calling it the second.

A model can produce a screen, a copy draft, a render, that looks like a decision was made. Often none was. It's a very confident shrug. It is describing what plausible output looks like for a brief like this one, the same way blue-500 describes a colour without saying what it is for. Whether that output is actually right for this product, this user, this moment, is a separate question, and it is one the model structurally cannot answer, because answering it requires knowing things that are not in the prompt.

The gap between plausible and right

I noticed the pattern most clearly not in client work but in something with nothing to do with software: I had a model generate renders of a house I just bought, to see rooms before a contractor touched anything. The renders were consistently plausible. The proportions were confidently wrong until I supplied actual measurements. The materials looked convincing and matched nothing sold in a hardware store. The light was always golden hour, because that photographs well, regardless of which way the room actually faces. I wrote the specifics up separately — it is a cleaner example of this argument than anything from client work, because a house doesn't negotiate. It is exactly the size it is.

That gap between plausible and right is the whole story. It shows up in interfaces the same way it shows up in renders: a screen that is internally consistent, on-brand, accessible, and still wrong, because it answered the letter of the brief instead of the intent behind it. Nothing about that screen looks unfinished. That is precisely what makes it dangerous to skip past.

Three questions that separate a decision from a plausible guess

I ask the same three questions of AI output regardless of what it is — an interface, a paragraph of copy, a render of a room — and I don't let myself skip to the last one.

Would I defend this choice out loud, for a reason that isn't "it looked right"? Almost everything a model produces now clears the bar of looking right. That bar has stopped being useful. The question that still filters something is whether there's a reason behind the choice that would survive being said to another person.

Does it hold against the constraint that actually can't move? Every brief has one or two non-negotiables buried in it, and a model treats every constraint in a prompt as equally soft unless told otherwise. Usually most of them are. Finding the one that isn't, and checking the output against specifically that one, is the part a model can't do for itself.

What would I change if I'd built this from scratch, in the same time? If the honest answer is nothing, the model did real work and I checked it, which is the good outcome. If the answer is a list, that list was the actual job, and treating the first draft as the finished one skips it.

Taste doesn't show up in the training data

The deepest version of this problem is not about facts a model got wrong, which can be corrected with better input. It's about calls that were never going to be in the data at all, because they're not universal, they're specific to this product, this brand, this room. A model can give you five options that are all technically sound and has no way to rank them by what actually matters here, because "what matters here" is not a property of the request, it's a property of the person who understands the context the request came from.

That's the part of the work that doesn't compress. Everything mechanical around it, the drafting, the variations, the first pass at a hundred small choices, a model does well and I'm glad to hand over. The part where you look at five technically-correct options and reject four of them for reasons that would take a paragraph to explain and would not survive being explained to the model that generated them: that's not a task waiting to be automated away. That's the job.

What actually changes day to day

None of this is an argument against using the tools, and it isn't a plea for restraint for its own sake. It changes what I spend my attention on. Less time generating the fifth variation, more time on whether the second one is actually right and can say why. Less time defending a process, more time defending a specific decision when someone asks about it. The model does more of the layer that's mostly mechanical. I do more of the layer that's mostly judgment. That trade only works if the second layer is still happening somewhere, on purpose, and not quietly skipped because the first one now produces something finished-looking fast enough that nobody stopped to check it.