Three weeks into our first real AI pilot, the model was doing something none of us had planned for, and the client was asking for a demo we weren't ready to give. That is the moment most teams read as evidence they picked the wrong project, hired the wrong people, or misread the technology. The data says otherwise. RAND Corporation interviewed sixty-five data scientists and engineers across government and industry in 2024 and put the AI project failure rate above eighty percent, roughly twice the failure rate of ordinary corporate IT projects, tracing the cause overwhelmingly to leadership and organisational factors rather than the model itself.
The rest of the research points the same direction from different angles. Gartner predicted thirty percent of generative AI projects would be abandoned after proof of concept by the end of 2025, over poor data quality, unclear business value and inadequate risk controls, factors visible before a line of code changes anything. MIT's NANDA project found ninety-five percent of generative AI pilots showing no measurable return as of 2025. BCG separately found seventy-four percent of companies still struggling to achieve and scale value from AI. Four different research bodies, four different methodologies, the same conclusion: if a first pilot feels shakier than the case study made it sound, the numbers say the team is in the majority, not the exception.
Most of the disappointment I have seen in AI projects traces back to a goal set before anyone had touched the tool. A polished case study shows the finished state, not the three weeks in the middle where the model did something nobody planned for. The gap between a working prototype and something a whole organisation can rely on is real, it is where most of that eighty percent lives, and mistaking the demo for the deployment is what turns a normal rough patch into a crisis of confidence.
The teams that get through that patch fastest tend to make the same first move: they make the first attempt smaller than it wants to be. Pick one narrow slice of the problem. Ship it to five users, not five hundred. Give the team a week to see what actually breaks before adding the next slice. A narrow pilot that wobbles is a Tuesday. A company-wide rollout that wobbles is a headline, and the difference between the two is scope decided in week one, not technology decided at any point.
Choose that first slice deliberately, too. The task that makes people groan when it lands on their desk, the report nobody wants to compile, the data entry nobody wants to do, is the one worth automating first, because the bar for success is low and the relief when it lifts is immediate. That early, unglamorous win buys the trust a team will need for the harder, more ambiguous project that comes after it.
None of it works solo. A team that admits together that the first attempt didn't work moves faster than one person trying to protect their own confidence by quietly patching around the problem alone. The wobble is easier to walk through when more than one person on the rope is willing to say out loud that it's wobbling, and it's harder for a project to recover from a failure nobody was willing to name until it was unavoidable.
Part of walking the rope steadily is knowing what it can and can't hold. AI is genuinely good at pattern-matching across data, at drafting, summarising and flagging what looks unusual against everything it has seen before. It is not good yet at holding judgment for a decision nobody has made before, the kind with no precedent in the training data and real consequences attached. Knowing that distinction in advance is what tells a team where to keep a human checkpoint in the loop, rather than discovering the gap the expensive way, after the model has already made the call.
The day our pilot finally worked cleanly, the room actually cheered, not because the project was finished, every real deployment keeps evolving long after the first success, but because everyone in that room had felt the same wobble together and come out the other side of it. That's the buy-in a business case built on headcount reduction never generates on its own: a team that owns the win because it survived the shaky part collectively, not one that was simply told the rollout succeeded.
Reference
This piece is adapted for Praxora Lab from the original. Originally published at terencekok.com ›
Sources
- RAND Corporation — The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed ›
- Gartner — Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025 ›
- MIT NANDA — The GenAI Divide: State of AI in Business 2025 ›
- Boston Consulting Group — Where's the Value in AI? (2024) ›