Foundations of dependable agentic AI
Why engineering reliability into agentic systems depends on bounded task specifications and trajectory-level observability in production, not on how capable the underlying model is.
Terence Kok
AI Governance & Assurance Practice. Enterprise AI Strategist and Keynote Speaker · Praxora Lab
The most common failure mode in agentic AI deployment is capability-led design: a system built around what the model can do, with controls retrofitted after the fact. Reliability depends primarily on the engineering controls placed around a model, not on the model's own capability, and a task only belongs to an autonomous agent if its success can be checked programmatically and the agent runs against a hard step and cost budget. Everything else is decision support, not automation.
Ten elements make up a specification-first architecture. A bounded task specification defines machine-checkable success criteria, explicit termination conditions, and a task contract covering input schema, side effects and cost budget before the agent runs. Closed-loop verification gives the system access to a deterministic checker, a compiler, a unit test, a constraint solver, so a plausible-sounding output can be distinguished from a correct one. Least-privilege authority scopes credentials to the task, with staged autonomy from autonomous reads through to human-confirmed irreversible writes, and an absolute boundary at operational technology and safety-instrumented systems that no autonomy tier crosses.
Exhibit · Before it runs
Three controls set before the agent starts
| Element | What it does |
|---|---|
| Bounded task specification | Machine-checkable success criteria, termination conditions, a task contract covering input schema, side effects and cost budget. |
| Closed-loop verification | A deterministic checker (compiler, unit test, constraint solver) that distinguishes plausible from correct. |
| Least-privilege authority | Credentials scoped to the task, staged autonomy, with an absolute boundary at OT and safety-instrumented systems. |
Every external input, including tool results, is treated as untrusted by default, structurally separated from instructions to prevent injection attacks. Trajectory-level observability logs every step, prompt hash, model version, tool call, arguments, result, latency and policy decision, because assurance requires a record, not a demo, and an incident can only be reconstructed from a log that was actually kept. Evaluation happens against known-correct trajectories with step-level scoring and adversarial test cases, not against whether the final answer merely looked right.
Exhibit · While it runs
Three controls active during execution
| Element | What it does |
|---|---|
| Untrusted input by default | Every external input, including tool results, structurally separated from instructions to prevent injection. |
| Trajectory-level observability | Every step, prompt hash, model version, tool call, arguments, result, latency and policy decision logged. |
| Evaluation against known-correct trajectories | Step-level scoring and adversarial test cases, not whether the final answer merely looked right. |
The remaining elements handle what happens when something goes wrong: checkpointed state, loop detection, spend ceilings and tested escalation paths that let a system degrade gracefully rather than fail silently, with deterministic conventional code handling routing, validation and sequencing so a model's judgment is reserved for the decisions that actually require it. The frameworks referenced, ISO/IEC 42001, the EU AI Act's Article 12, and comparable national standards, all converge on the same underlying requirement: a system has to be built to produce evidence of its own reliability, not just claim it.
Exhibit · When it fails
Four controls for graceful degradation
| Element | What it does |
|---|---|
| Checkpointed state | So a failed run can resume or roll back rather than restart blind. |
| Loop detection | Catches an agent repeating itself before it burns the budget finding out. |
| Spend ceilings | A hard cost and step budget the agent cannot run past. |
| Tested escalation paths | A system that degrades gracefully rather than fails silently. |
This piece is adapted for Praxora Lab from the original: Originally published at terencekok.com (https://terencekok.com/blog/foundations-dependable-agentic-ai).
Terence Kok
AI Governance & Assurance Practice. Enterprise AI Strategist and Keynote Speaker
Enterprise AI strategist and former Chief AI and Innovation Officer at Meinhardt Group, with twenty-five years leading transformation programmes across Asia and the Middle East, specialising in impact assessment, governance and deployment methodology.
Praxora Lab runs twelve half-day courses and three full-day integrated programmes, each led by a named practitioner, turning frameworks like this one into a deployment roadmap.
© 2026 Praxora Lab. Author: Terence Kok. Read online at praxoralab.com/insights/foundations-dependable-agentic-ai