MEMO · TO readers evaluating a workshop or a speaker · RE Capacity Building

Insights Capacity Building

Implementing AI feels like walking a tightrope, and you're not the only one wobbling

Four independent studies put AI project failure somewhere between thirty and ninety-five percent depending on how you count it. The wobble a team feels on its first real pilot isn't a sign the project is failing, it's the statistically normal experience, and there's a specific set of moves that gets a team through it faster.

Download branded PDF report
Terence Kok

About the author

Terence Kok

Executive Director, AI Governance & Assurance Practice. Enterprise AI Strategist and Keynote Speaker

Enterprise AI strategist and former Chief AI and Innovation Officer at Meinhardt Group, with twenty-five years leading transformation programmes across Asia and the Middle East, specialising in impact assessment, governance and deployment methodology.

Read the full profile ›

Three weeks into our first real AI pilot, the model was doing something none of us had planned for, and the client was asking for a demo we weren't ready to give. That is the moment most teams read as evidence they picked the wrong project, hired the wrong people, or misread the technology. The data says otherwise. RAND Corporation interviewed sixty-five data scientists and engineers across government and industry in 2024 and put the AI project failure rate above eighty percent, roughly twice the failure rate of ordinary corporate IT projects, tracing the cause overwhelmingly to leadership and organisational factors rather than the model itself.

The rest of the research points the same direction from different angles. Gartner predicted thirty percent of generative AI projects would be abandoned after proof of concept by the end of 2025, over poor data quality, unclear business value and inadequate risk controls, factors visible before a line of code changes anything. MIT's NANDA project found ninety-five percent of generative AI pilots showing no measurable return as of 2025. BCG separately found seventy-four percent of companies still struggling to achieve and scale value from AI. Four different research bodies, four different methodologies, the same conclusion: if a first pilot feels shakier than the case study made it sound, the numbers say the team is in the majority, not the exception.

Most of the disappointment I have seen in AI projects traces back to a goal set before anyone had touched the tool. A polished case study shows the finished state, not the three weeks in the middle where the model did something nobody planned for. The gap between a working prototype and something a whole organisation can rely on is real, it is where most of that eighty percent lives, and mistaking the demo for the deployment is what turns a normal rough patch into a crisis of confidence.

The teams that get through that patch fastest tend to make the same first move: they make the first attempt smaller than it wants to be. Pick one narrow slice of the problem. Ship it to five users, not five hundred. Give the team a week to see what actually breaks before adding the next slice. A narrow pilot that wobbles is a Tuesday. A company-wide rollout that wobbles is a headline, and the difference between the two is scope decided in week one, not technology decided at any point.

Choose that first slice deliberately, too. The task that makes people groan when it lands on their desk, the report nobody wants to compile, the data entry nobody wants to do, is the one worth automating first, because the bar for success is low and the relief when it lifts is immediate. That early, unglamorous win buys the trust a team will need for the harder, more ambiguous project that comes after it.

None of it works solo. A team that admits together that the first attempt didn't work moves faster than one person trying to protect their own confidence by quietly patching around the problem alone. The wobble is easier to walk through when more than one person on the rope is willing to say out loud that it's wobbling, and it's harder for a project to recover from a failure nobody was willing to name until it was unavoidable.

Part of walking the rope steadily is knowing what it can and can't hold. AI is genuinely good at pattern-matching across data, at drafting, summarising and flagging what looks unusual against everything it has seen before. It is not good yet at holding judgment for a decision nobody has made before, the kind with no precedent in the training data and real consequences attached. Knowing that distinction in advance is what tells a team where to keep a human checkpoint in the loop, rather than discovering the gap the expensive way, after the model has already made the call.

The day our pilot finally worked cleanly, the room actually cheered, not because the project was finished, every real deployment keeps evolving long after the first success, but because everyone in that room had felt the same wobble together and come out the other side of it. That's the buy-in a business case built on headcount reduction never generates on its own: a team that owns the win because it survived the shaky part collectively, not one that was simply told the rollout succeeded.

Reference

This piece is adapted for Praxora Lab from the original. Originally published at terencekok.com ›

Dr. Jayarethanam Pillai

Before you go

Eighty percent is a figure I recognise from a different context entirely. When I reviewed new academic programmes as a dean, the ones that stumbled in their first term rarely failed because the curriculum was wrong, they failed because nobody had told the faculty running it that a rough first semester was the expected shape of the thing, not evidence the idea was mistaken. RAND's finding that the cause is overwhelmingly organisational rather than technical matches exactly what I watched happen with institutional pilots at UNDP: a leadership team reads the early wobble as proof of a bad decision and pulls support before the team has had the week Terence describes to find out what actually breaks. The discipline I would add from my own seat is smaller than his but just as easy to skip: name the wobble in a written report before someone above you notices it first. An institution that puts its own shaky start on the record is far harder to defund quietly than one that tries to hide it until the numbers force the conversation anyway.

Signature, Jayarethanam Pillai