The time a student caught me trusting a boxplot
The moment a tool flags something, most people stop questioning it. A teaching story about why a flag is a candidate for judgement, not a verdict, for boxplots and for AI systems alike.
Dr. Joshua Lau
AI Transformation Leader and Accredited Adult Educator · Praxora Lab
I run outlier detection exercises constantly, boxplots on Titanic passenger fares, tips dataset totals, whatever gives a class something concrete to argue with. And across cohort after cohort, the predictable failure is never the statistics. Learners get the IQR method fine. They can compute the fences and read the plot. The failure happens in the half-second after the tool renders: the moment a boxplot flags a point as an outlier, the learner treats it as settled fact. No second question about whether that flagged point is a data entry error, a genuinely unusual but valid observation, or just an artefact of a threshold too aggressive for this context.
Those are three different situations demanding three different responses, fix it, keep it, or adjust the method, and the boxplot cannot tell you which one you're in. Only domain judgement can. But the visual authority of the flag short-circuits exactly that judgement. The tool has spoken; the human stands down.
Exhibit · Before acting on any flag
A flag is a candidate for judgement, not a verdict
- 01
Data entry error
Fix it.
- 02
Genuinely unusual but valid
Keep it.
- 03
Threshold too aggressive for this context
Adjust the method.
I would dearly love to present this as a beginner's error I observe from a safe professional distance. Unfortunately, I have witnesses. Early in my teaching, I took a flagged outlier at face value in front of a full room. Built a small analytical argument on top of it, walked the class through the implications, felt rather pleased with the flow of the lesson. Then a sharp learner near the back raised a hand and asked, quite mildly, whether I'd checked if the point was just a data entry issue. I hadn't. I had done the precise thing I was there to train out of people, at the front of the room, with all the confidence of the person holding the marker. There's a particular flavour of silence that follows a moment like that, and I can still taste it.
To be fair to the moment, it did more for that class than any slide I've ever made. They got to watch the instructor fall into the trap live, which teaches its gravity in a way no warning can. Nobody is immune, least of all the person who's started to believe familiarity has made them immune, which described me rather well that morning.
The reason I keep retelling the story, at some cost to my dignity, is that the boxplot moment maps almost exactly onto how staff treat AI-flagged output in real operational settings. Fraud detection flags a transaction, and the analyst treats it as fraud rather than as a candidate for fraud. A quality model flags a defect, and the line supervisor pulls the batch without asking whether the threshold suits this product run. A screening tool ranks a candidate low, and the recruiter never opens the CV. In every case the tool did its job, surfacing a candidate for human judgement, and the human quietly resigned from the judging.
The training implication is uncomfortable for anyone who sells tool workshops, my past self included: teaching the tool is the easy half, and on its own it can make things worse. A workforce trained to operate a flagging system, without any structured scepticism toward its flags, is a workforce trained to defer at scale. What I build into sessions now is one standing rule for any flagged result: before acting, name which of the three situations you might be in, and say what evidence would tell them apart. Ten seconds of thought. It's roughly the ten seconds I skipped that morning, and I've been collecting interest on it ever since.
If you take one thing: a flag is a candidate for judgement, not a verdict. I learned that from a student in the back row, which is probably how I deserved to learn it.
Dr. Joshua Lau
AI Transformation Leader and Accredited Adult Educator
Managing Director leading AI transformation at a Malaysian manufacturer, and a WSQ and HRDC accredited trainer who has delivered generative AI and agentic AI training to more than 1,000 learners across Singapore and Malaysia.
Praxora Lab runs the AI Governance & ROI Executive Programme and the AI Masterclass, turning frameworks like this one into a deployment roadmap.
© 2026 Praxora Lab. Author: Dr. Joshua Lau. Read online at praxoralab.com/insights/outlier-detection-and-trust