: Report / Praxora Lab
Foundations of dependable agentic AI

Save as PDF: File → Print → Save as PDF  |  ← Back to the article

PRAXORALAB
Governance

Insight Report · Praxora Lab

Foundations of dependable agentic AI

Why engineering reliability into agentic systems depends on bounded task specifications and trajectory-level observability in production, not on how capable the underlying model is.

Terence Kok

Terence Kok

AI Governance & Assurance Practice. Enterprise AI Strategist and Keynote Speaker · Praxora Lab

The most common failure mode in agentic AI deployment is capability-led design: a system built around what the model can do, with controls retrofitted after the fact. Reliability depends primarily on the engineering controls placed around a model, not on the model's own capability, and a task only belongs to an autonomous agent if its success can be checked programmatically and the agent runs against a hard step and cost budget. Everything else is decision support, not automation.

Ten elements make up a specification-first architecture. A bounded task specification defines machine-checkable success criteria, explicit termination conditions, and a task contract covering input schema, side effects and cost budget before the agent runs. Closed-loop verification gives the system access to a deterministic checker, a compiler, a unit test, a constraint solver, so a plausible-sounding output can be distinguished from a correct one. Least-privilege authority scopes credentials to the task, with staged autonomy from autonomous reads through to human-confirmed irreversible writes, and an absolute boundary at operational technology and safety-instrumented systems that no autonomy tier crosses.

Exhibit · Before it runs

Three controls set before the agent starts

ElementWhat it does
Bounded task specificationMachine-checkable success criteria, termination conditions, a task contract covering input schema, side effects and cost budget.
Closed-loop verificationA deterministic checker (compiler, unit test, constraint solver) that distinguishes plausible from correct.
Least-privilege authorityCredentials scoped to the task, staged autonomy, with an absolute boundary at OT and safety-instrumented systems.

Every external input, including tool results, is treated as untrusted by default, structurally separated from instructions to prevent injection attacks. Trajectory-level observability logs every step, prompt hash, model version, tool call, arguments, result, latency and policy decision, because assurance requires a record, not a demo, and an incident can only be reconstructed from a log that was actually kept. Evaluation happens against known-correct trajectories with step-level scoring and adversarial test cases, not against whether the final answer merely looked right.

Exhibit · While it runs

Three controls active during execution

ElementWhat it does
Untrusted input by defaultEvery external input, including tool results, structurally separated from instructions to prevent injection.
Trajectory-level observabilityEvery step, prompt hash, model version, tool call, arguments, result, latency and policy decision logged.
Evaluation against known-correct trajectoriesStep-level scoring and adversarial test cases, not whether the final answer merely looked right.

The remaining elements handle what happens when something goes wrong: checkpointed state, loop detection, spend ceilings and tested escalation paths that let a system degrade gracefully rather than fail silently, with deterministic conventional code handling routing, validation and sequencing so a model's judgment is reserved for the decisions that actually require it. The frameworks referenced, ISO/IEC 42001, the EU AI Act's Article 12, and comparable national standards, all converge on the same underlying requirement: a system has to be built to produce evidence of its own reliability, not just claim it.

Exhibit · When it fails

Four controls for graceful degradation

ElementWhat it does
Checkpointed stateSo a failed run can resume or roll back rather than restart blind.
Loop detectionCatches an agent repeating itself before it burns the budget finding out.
Spend ceilingsA hard cost and step budget the agent cannot run past.
Tested escalation pathsA system that degrades gracefully rather than fails silently.
Reference

This piece is adapted for Praxora Lab from the original: Originally published at terencekok.com  (https://terencekok.com/blog/foundations-dependable-agentic-ai).

About The Author
Terence Kok

Terence Kok

AI Governance & Assurance Practice. Enterprise AI Strategist and Keynote Speaker

Enterprise AI strategist and former Chief AI and Innovation Officer at Meinhardt Group, with twenty-five years leading transformation programmes across Asia and the Middle East, specialising in impact assessment, governance and deployment methodology.

Want this applied to your organisation?

Praxora Lab runs twelve half-day courses and three full-day integrated programmes, each led by a named practitioner, turning frameworks like this one into a deployment roadmap.

Explore workshops →

© 2026 Praxora Lab. Author: Terence Kok. Read online at praxoralab.com/insights/foundations-dependable-agentic-ai