MEMO · TO readers evaluating a workshop or a speaker · RE Strategy

Insights Strategy

Why large enterprises are deploying AI at scale but measuring it incorrectly

Common measurement frameworks target the wrong layer of the technology stack, missing the business outcomes that actually matter.

Download branded PDF report
Terence Kok

About the author

Terence Kok

AI Governance & Assurance Practice. Enterprise AI Strategist and Keynote Speaker

Enterprise AI strategist and former Chief AI and Innovation Officer at Meinhardt Group, with twenty-five years leading transformation programmes across Asia and the Middle East, specialising in impact assessment, governance and deployment methodology.

Read the full profile ›

Capital expenditure on AI infrastructure across Southeast Asia, Greater China and the Indian subcontinent reached record levels in 2025, and much of it is being measured against the wrong yardstick. A model that is ninety-four percent accurate is a capability specification. It is not proof that the business it was deployed into has actually changed, and enterprises that conflate the two are optimising for a number that was never the point.

There are two distinct measurement layers, and most AI programmes only instrument one of them. The technical layer covers accuracy, F1 score, inference latency and model drift, useful for engineering teams, largely irrelevant to a board. The operational layer covers processing time per transaction, decision consistency, error and exception frequency, and cost per unit of output measured against a baseline set before deployment. Conflating the two produces decisions that optimise a model metric while the business outcome it was meant to move stays flat.

Exhibit · Two measurement layers

Where enterprises are pointing the instrument

Most AI programmes only instrument one layer. The board can only act on the other.

LayerWhat it tracksWho it's for
TechnicalAccuracy, F1 score, inference latency, model driftEngineering teams (largely irrelevant to a board)
OperationalProcessing time per transaction, decision consistency, error and exception frequency, cost per unit vs. baselineExecutive governance, the layer a board can actually act on

Three structural causes recur. AI sponsorship sits inside technology functions rather than the operating units that own the outcome. Vendor contracts are written around model performance benchmarks rather than business impact. And deployment speed keeps outpacing the measurement infrastructure needed to know whether any of it worked.

The corrective sequence is five steps: define the operational baseline before deployment, set operational KPIs with explicit attribution rules, instrument the workflow rather than only the model, connect model monitoring to operational alerting, and report the operational impact, not the model metrics, to executive governance on a fixed cadence. None of this replaces model-level monitoring. It sits above it, and it is the layer a board can actually act on.

Exhibit · The corrective sequence

Five steps, run in order

  1. 01

    Define the operational baseline

    Before deployment, not after.

  2. 02

    Set operational KPIs

    With explicit attribution rules.

  3. 03

    Instrument the workflow

    Not only the model.

  4. 04

    Connect model monitoring to operational alerting

    So drift surfaces where it matters.

  5. 05

    Report the operational impact to governance

    On a fixed cadence, not the model metrics.

Reference

This piece is adapted for Praxora Lab from the original. Originally published at terencekok.com ›

Dr. Jayarethanam Pillai

Before you go

This is an argument I recognise from a different discipline entirely. Economists spend a great deal of energy distinguishing an input metric from an outcome metric, and watching a board celebrate a ninety-four percent accurate model the way Terence describes here is not so different from a finance ministry celebrating disbursed budget while the outcome it was meant to fund never materialises. The corrective sequence he lays out, baseline first, then operational KPIs, is the same sequencing I would insist on for a public policy evaluation. Measure the thing you actually meant to change, not the thing that was easiest to instrument.

Signature, Jayarethanam Pillai