Why AI Projects Fail: It’s Not a Technology Problem, It’s a Delegation Problem
Most AI projects in the Mittelstand start the same way: a license gets bought, a pilot gets launched, and a young colleague who “has done some coding” gets handed the keys. Then, six months later, nothing has changed — except the invoice.
The technology worked fine. The delegation didn’t.
The Pattern I Keep Seeing
Over the past two years, I’ve run AI workshops at more than 50 companies. The failure pattern almost never varies. The topic lands in IT. There is enthusiasm at the start. There is a pilot. And then there is silence — not because the tools are inadequate, but because the people now responsible for the project cannot answer the question that actually matters: where is this business model going?
That is not an IT question. It never was.
I had this conversation recently with Christoph Maichel — a CTO who has led teams of up to 100 developers, taken a company public, and spent four decades writing code. His take was harder than mine: he believes the classical IT department, as most Mittelstand companies have structured it, needs to be fundamentally rethought. Not eliminated. Rethought. And he is saying this as someone who has lived inside that structure for most of his career.
When someone with that background reaches that conclusion, I pay attention.
AI Is an Operating System Question, Not a Software Question
Here is the distinction I keep coming back to: software you can buy. An operating model you have to build.
When a company treats AI as a software purchase — a line item in the IT budget — it is asking the wrong department to solve the right problem. AI changes how a company thinks, decides, and creates value. That is a boardroom conversation. Delegating it to IT is like asking your finance controller to set your five-year strategic direction because they are good with numbers.
Licenses are commodities. The competitive advantage comes from the gates you set around them — the points in the workflow where a human expert reviews the output before anything moves forward. In my own work, I call these Human-in-the-Loop checkpoints. They are not optional safeguards. They are where the real intellectual value sits.
At Pheraxia, for example, AI is not a project running alongside the business. It is baked into how we generate and validate ideas. But every meaningful output still passes through a domain expert before it becomes a decision. The AI accelerates. The expert judges. That division of labor is intentional — and it is a leadership design choice, not a technical one.
A Lesson From My Own Backup Script
A few weeks ago, hours before I open-sourced a repository on reliable AI agent systems, one of my own agents caught a bug in my backup script.
The script would have exited cleanly. It would have written a green log line. It looked healthy. But under one specific misconfiguration, it was saving nothing. I would probably have trusted it for months — until the day I needed a restore.
What made this uncomfortable was that the rule this bug violated was already written in the repo. In my own words. “An unwatched backup is a future silent failure.” I wrote that from experience. Then I wrote code that broke it.
The checklist caught its own author.
I share this not as a technical anecdote but as a leadership one. The gap between prediction and delivery is where most things go wrong — in AI projects, in operational systems, and in strategies. We are very good at naming principles. We are less disciplined about building the watchdogs that hold us to them.
The Mini-Framework: The Three-Layer AI Governance Stack
When I work with SME leaders trying to move from AI curiosity to AI operations, I use a simple three-layer model. It is not glamorous. It works.
Layer 1 — Ownership (Chefsache) AI direction is set at CEO or C-suite level. Not delegated entirely to IT. The business model question must be answered at this layer before any tool is chosen.
Layer 2 — Expert Gates (Human-in-the-Loop) Every output that feeds a real decision passes through a domain expert before it moves. This is not about distrust of the technology. It is about knowing where the intellectual value actually sits in your business.
Layer 3 — Watchdogs (Silent Failure Prevention) Anything that can fail silently — a data feed, an automated report, a backup, a scoring model — gets an independent verifier. Not the same agent that produced the output. A fresh context, no attachment to the work.
These three layers are not a technology stack. They are a governance stack. And that is exactly why they belong in leadership conversations, not IT tickets.
The Honest Sentence About AI Right Now
Christoph and I have been building together in public — through our 10xCEO videocast — precisely because we believe the most useful thing two practitioners can do right now is show their work while they are still figuring it out.
The most honest thing I can say about AI in 2025: nobody knows exactly where this lands. Not the researchers. Not the consultants. Not me.
But here is what I do know. The companies that will matter in three years are not the ones that bought the best licenses. They are the ones where leadership stayed in the room, asked the hard business model questions early, and built the governance infrastructure before they needed it.
The delivery gap — the space between a good prediction and a working system — is where most value is lost and most competitive advantage is quietly built. That gap does not close by itself.
What You Can Do Next Week
-
Audit your AI ownership. Write down every active AI initiative in your company. Next to each one, name who owns the business outcome — not the tool, the outcome. If the answer is “IT,” that is your first conversation to have.
-
Identify your Human-in-the-Loop gates. For each AI-assisted output that feeds a real decision, ask: who reviews this before it moves? If no one does, define that person and that moment this week.
-
List your silent failure risks. What in your operations looks healthy but could be quietly broken? Automated reports, data syncs, backup processes, scoring models. Pick one and add a watchdog — an independent check with its own verification logic.
-
Schedule one boardroom AI conversation. Not a demo. A conversation about where your business model is going and where AI either accelerates or threatens that direction. Thirty minutes. The right people in the room.
-
Separate prediction from delivery. In your current AI pilots, ask honestly: are we measuring whether the output is accurate, or whether it actually reaches the right place and produces a result? Most teams measure the first. Few measure the second.
I’m curious: in your experience, where does AI ownership currently sit in your company — and do you think that is the right place?