MEMO · TO readers evaluating a workshop or a speaker · RE Governance
From human-in-the-loop to AI-on-the-loop: redesigning oversight architectures
How oversight structures need to change as AI systems take on more decision-making without a person approving every step, and why that shift is a design choice regulators already permit.
Terence Kok, Enterprise AI Strategist, Author, Keynote Speaker
Human-in-the-loop, a human reviewing and approving every individual AI decision, is a design choice, not a regulatory mandate. Regulators across the EU and Singapore require human accountability and a working intervention capability, requirements that are fully compatible with AI-on-the-loop architectures where an automated system performs the primary oversight function and a human governs the configuration, not every transaction inside it.
The gap most organisations run into is that current governance frameworks assume a human is ultimately responsible for a high-risk decision, without specifying how that responsibility is discharged when the oversight itself is automated. That ambiguity produces unnecessary operational constraints, organisations manually reviewing high volumes of low-risk decisions because nobody has redesigned what oversight actually requires at that volume.
A redesigned architecture rests on four elements: human oversight redefined at the level of configuration and governance rather than individual transactions, at least two independent AI control layers trained on distinct data so one does not simply validate the other's blind spot, full auditability across both the operational and the oversight AI's outputs, and a maintained human fallback path that can suspend or override the system at any point. In domains like anomaly detection across high-volume decision sets, automated oversight already outperforms manual review on consistency, where reviewer fatigue turns human sign-off into a symbolic approval rather than a real check.
The regulatory test is not whether a human clicked approve on each decision. It is whether an organisation can demonstrate that its control layers are explainable, monitored and subject to accountable human oversight. Smart city infrastructure is a working example: operational agents handle domain-specific functions, an independent oversight agent monitors compliance separately, and a human-in-command layer sets the escalation thresholds and adjusts policy, present in the system without being present in every decision it makes.
Exhibit ยท Redesigning oversight
A design choice, not a regulatory mandate
Human-in-the-loop
A human reviews every decision
- A person approves each individual AI decision
- Reviewer fatigue turns high-volume sign-off into symbolic approval
AI-on-the-loop
A human governs the configuration
- An automated system performs the primary oversight function
- A human governs the configuration, not every transaction inside it
- 01
Oversight redefined
At the level of configuration and governance, not individual transactions.
- 02
Two independent control layers
Trained on distinct data, so one does not simply validate the other's blind spot.
- 03
Full auditability
Across both the operational and the oversight AI's outputs.
- 04
A maintained human fallback path
That can suspend or override the system at any point.
Reference
This piece is adapted for Praxora Lab from the original. Originally published at terencekok.com ›
More in Governance
- Beginning your journey: identifying tasks for quality, traceable, auditable AI agents
The TRACE framework that structures how to evaluate whether a task is right for autonomous agent deployment, and how much oversight it needs.
- Foundations of dependable agentic AI
Why engineering reliability into agentic systems depends on bounded task specifications and trajectory-level observability in production, not on how capable the underlying model is.
- Beyond the pilot: a risk governance framework for scalable AI deployment
A governance framework for the step most AI programmes skip: moving from a working pilot to infrastructure that can be trusted at scale, built around a four-pillar risk model and a scoring method borrowed from industrial engineering.