Most organisations deploying AI agents choose a vendor and a model before they have defined what quality, traceability and auditability mean for the task in front of them. That sequencing error, not a gap in the technology, is what produces audit failures downstream. Task suitability has to be decided first.
The TRACE framework scores a candidate task against five criteria. Traceability requires every input and output to be logged, timestamped and attributed to a specific data source, disqualifying tasks that run on unstructured or unverified data until provenance controls exist. Reversibility and risk tier classifies the consequence of a wrong output as reversible with no harm, reversible with cost, or irreversible, with irreversible-consequence tasks requiring a human checkpoint before they proceed. Acceptance criteria demands a measurable, pre-agreed definition of success that exists independently of the agent's own output, using external ground truth rather than the model marking its own work.
Exhibit · TRACE, part one
What gets a task admitted for review
-
Traceability
Every input and output logged, timestamped and attributed to a specific data source.
-
Reversibility & risk tier
Reversible with no harm, reversible with cost, or irreversible. Irreversible tasks require a human checkpoint.
-
Acceptance criteria
A measurable, pre-agreed definition of success, using external ground truth rather than the model marking its own work.
Compliance mapping requires the applicable regulatory regime and internal policy to be identified before architectural decisions are made, not retrofitted once the system is in production. Escalation pathway requires a tested, working mechanism for a human to intervene, override or halt the agent mid-task, not a theoretical one described in a policy document nobody has run.
Exhibit · TRACE, part two
What gets a task cleared for deployment
-
Compliance mapping
The applicable regulatory regime and internal policy identified before architectural decisions are made.
-
Escalation pathway
A tested, working mechanism for a human to intervene, override or halt the agent mid-task.
A task that satisfies all five criteria is a reasonable candidate for agent deployment. A task that fails two or more stays manual or human-assisted until the failing criteria are remediated. The practical recommendation is to start with three to five tasks at most, reverse the usual procurement sequencing so risk classification happens before vendor selection, and mandate a fixed review interval, quarterly at minimum, that turns auditability from a design-time aspiration into an operating discipline a team actually runs.
Reference
This piece is adapted for Praxora Lab from the original. Originally published at terencekok.com ›