A while back I built a classroom exercise around a simple game benchmark, mostly because I wanted one distinction to become physical rather than theoretical. The exercise compares two ways of judging the same artefact. One asks: does this hit the frame-rate target, yes or no? The other asks: does this feel responsive? The first question can prove you wrong. The second one can't, which is precisely why it's so comfortable to reach for.
I designed the exercise expecting the junior learners to gravitate toward the vibes-based judgement. What actually happened surprised me: plenty of experienced professionals did the same when left alone. People who, in their own fields, would never accept "feels about right" as sign-off criteria, put an impressive-looking AI-generated artefact in front of them, and the discipline quietly dissolved. "Looks good to me" won, again and again, across cohorts.
I understand the pull because I feel it too. A test you can't fail never embarrasses you. "Does this feel responsive" doesn't require committing to a definition of success in advance, and it never produces the awkward moment where the thing that demos beautifully fails the actual acceptance criteria in front of colleagues. Vibes-based evaluation is socially frictionless. Which is more or less what makes it dangerous once everyone is doing it.
And everyone increasingly is, which is the shift I think capacity building hasn't caught up with. For the past two years, AI training has mostly meant teaching people to use AI, prompt it, steer it, fold its output into their work. That era is winding down. As AI-assisted building spreads into ordinary jobs, the marketing exec generating a landing page, the analyst producing a dashboard, the ops lead automating a workflow, a much larger share of the workforce now spends its time judging AI output, often in domains where they're not deep experts. The bottleneck has moved from generating to evaluating. And honest evaluation requires writing the test before you see the answer.
That ordering matters more than any single technique I teach. A test written after you've seen the output isn't really a test; it's a rationalisation with better formatting. The model produces something plausible, your brain quietly redefines success to match what's on screen, and you sign off on criteria you'd never have accepted an hour earlier. I know this intimately because I still do it. After years of lecturing rooms about falsifiable gates, I'll catch myself eyeballing a generated result, feel the "that looks right" reflex firing, and have to consciously drag myself back to the question of what I said success meant before I ran the thing. The reflex never fully goes away. The discipline turns out to be a practice rather than a personality trait, which I find oddly reassuring to admit, it means my lapses are normal, not disqualifying.
What this means for training is a change in what we drill. Less time on prompt phrasing; more on acceptance criteria. Before any generation exercise, I now make learners write down, in a line or two, what a passing result must do, in terms that can fail. Then we generate, then we check against the written line rather than against the demo's charisma. It's a small habit, and unglamorous. It's also the single highest-leverage thing I know how to install, because it survives tool changes, model upgrades, and whatever the next hype cycle turns out to be.
Exhibit · The ordering that matters
Write the test before you see the answer
- 01
Write down what a passing result must do
In terms that can fail — before you look at the output.
- 02
Generate the output
Now, and only now, run the exercise.
- 03
Check the output against the written line
Not against how convincing it looks.
If you take one thing: write down what success means before you look at the output. I've taught this for years and I still have to force myself to do it, which is rather the point.