Skip to main content
Lahullier ConsultingLahullier ConsultingExecutive AI Strategy & Advisory

August 2, 2026 · 8 min read

AI Pilots Don't Fail Because the Model Is Bad. They Fail Because the Operating Model Is Missing

Read the original on LinkedIn

The post-mortem blames the vendor. The actual gap is ownership, escalation, and who signs off before it ships.

The Post-Mortem Always Blames the Wrong Thing

A pilot gets pulled. Someone schedules the retro, and within ten minutes the conversation drifts to the model. Maybe it hallucinated on an edge case. Maybe the vendor's roadmap slipped. Maybe a newer model came out and someone floats switching. That's usually where the conversation ends. That's usually the wrong ending.

I've sat in enough of these rooms to notice the pattern. Pull the model's outputs in isolation — same prompts, same data, judged on their own merits — and it usually performed fine. Not perfect. Fine. Good enough to be useful. The reasoning engine wasn't what broke. Everything wrapped around it broke: nobody had agreed on who was accountable when the output was wrong, nobody had built a path for a confused employee to escalate, nobody had defined what "working" was even supposed to look like before launch.

The real diagnosis doesn't come from re-scoring the model. It comes from asking who owned the outcome. Not who owned the project charter. Not who sponsored the initiative on a steering committee slide. Who was on the hook when the output was wrong and a customer or a colleague was standing in front of it, waiting. In most stalled pilots, the honest answer is nobody, exactly. That's the finding. The model is rarely the postmortem's actual subject. It's just the easiest thing to point at.

Only 1 in 10 Organizations Actually Made It to Production

Adoption of AI tools is close to universal at this point. Nearly every enterprise I talk to has licenses deployed, a pilot running somewhere, executives who've used a chatbot to draft something. That part of the story is done. Execution maturity is a completely different story, and it's a much smaller club.

Recent industry survey data puts the share of organizations operating AI at genuinely embedded scale — beyond pilots, running in real workflows with real accountability — at roughly 10 percent. Not 10 percent of AI initiatives. Ten percent of organizations, out of everyone who's tried. The other 90 percent have tools deployed, dashboards showing usage numbers, and very little that's actually changed how work gets done.

That gap is worth sitting with. It isn't a budget gap — most of the 90 percent have funded pilots, sometimes multiple rounds of them. It isn't a model-access gap either. Frontier models are commercially available to basically anyone with a credit card and a procurement signature. The 10 percent and the 90 percent are largely working with the same tools.

The difference is whether someone actually redesigned how the work gets done. The 10 percent didn't install a tool into an existing process and wait for adoption metrics to climb. They changed the process itself — who does which step, who reviews what, where the human checkpoint sits, what happens when the AI is uncertain. That's a harder, less glamorous project than picking a model. It's also the one most organizations skip.

What an Operating Model Actually Means Here

When I say "operating model," I don't mean a governance binder or a risk framework someone drafts and files away. I mean working answers to a small set of blunt questions: Who owns this output? Who gets escalated to when it's wrong? Who signs off before it ships to a customer or a claim or a contract? If you can't answer those in a sentence each, you don't have an operating model. You have a demo that's been given a production URL.

Across the failed and stalled pilots I've watched or heard about in detail, eight components keep showing up as missing, in some combination:

  • Ownership — a named person accountable for output quality, not a team or a committee
  • Escalation paths — what a confused end user actually does when something looks wrong
  • Review checkpoints — where a human looks at the output before it goes further
  • Accountability — consequences and correction loops when the process fails
  • Training — not tool onboarding, but training on when to trust the output and when not to
  • Success metrics — defined before launch, tied to business outcomes
  • Workflow design — the actual sequence of steps rebuilt, not just annotated with "AI-assisted"
  • Decision rights — who can pause, override, or kill the thing, and under what conditions

Most pilots define none of these before launch. They define all of them after something breaks, usually improvised, in a hurry, under pressure — exactly the wrong condition for designing accountability structures. That's the tell: this is a design problem, not a compliance problem. Compliance problems get solved with a policy document. Design problems get solved by someone sitting down before launch and deciding, on purpose, who does what.

Why the Developer Playbook Doesn't Transfer

AI coding tools are the go-to example when someone wants to argue that adoption can be fast and painless. Fair enough — AI-assisted coding did take off quickly, with real productivity gains and surprisingly little drama. But that's not because the model was uniquely suited to code. It's because software development already had an operating model AI could slot into.

Code review was already a checkpoint. Version control already gave you rollback. Pull requests already had an owner and an approver. Testing pipelines already caught a category of failure before it shipped. AI-generated code didn't need a new operating model built around it — it inherited one that had been maturing for two decades. It looked like the technology carried its own success with it. It didn't. It walked into a house that was already built.

Knowledge-work and operational workflows — underwriting, claims handling, customer service, procurement — usually don't have that same scaffolding. There's no standard "pull request" for a customer service response or a claims determination. There's often no clean rollback. And the review checkpoints, where they exist, were built for human error patterns — not for a system that's confidently wrong in a completely different way than a tired analyst is wrong.

Leaders who watched developer tools succeed sometimes assume the pattern repeats elsewhere by default. It doesn't. The absence of drama in the coding case was never the technology. It was the operating model already sitting there, invisible because it worked. Drop the same technology into a workflow with no equivalent scaffolding and you don't get the same result — you get the scaffolding's absence exposed, usually at the worst possible moment.

Agentic tooling raises the stakes considerably. Interest in agents that take action — not just draft text for a human to review, but actually execute a step, send a communication, update a record — is running well ahead of the governance needed to run them responsibly. This isn't a hypothetical risk. It's a sequencing problem. Experimentation is outpacing the operating model again, and agents make the gap more expensive, because there's less of a human in the loop to catch the mistake before it lands somewhere real.

The Questions That Should Precede Any Pilot

If you're the executive deciding whether to greenlight the next AI initiative, there's a short list of questions worth asking before the kickoff meeting, not after the first incident.

  • Name the single owner accountable for output quality — not the project. The output. If the answer is a team name or "the AI committee," that's not an answer yet.
  • Write the escalation path down, explicitly, for two scenarios: what happens when a human catches the model being wrong, and — the harder one — what happens when no human catches it. Most orgs have a rough answer to the first and nothing for the second.
  • Set the success metric before launch, and make it a business metric. Usage numbers, adoption rates, and query volume tell you people are clicking the tool. They don't tell you whether claims are processed more accurately, whether customers are better served, or whether any business outcome actually moved. Pick the business metric first, before anyone's bonus gets attached to a launch date.
  • Decide who has authority to pause or kill the workflow — and make sure it's not the person whose budget depends on keeping it running. This is the one teams skip most, because it's uncomfortable and because everyone assumes it'll never come up. It comes up.

None of these questions are about the model. You can answer every one of them without knowing which frontier model you'll use. That's the point. If you can't answer them, the model selection meeting is premature. You'd be optimizing a decision that was never the one holding you back.

Move the Model Selection Meeting to Week Four

Most AI initiatives start with a model bake-off. Compare vendors, run a few evaluation prompts, pick a winner, announce it, move to build. That sequence puts the least important decision first, and leaves the decisions that actually matter to get sorted out later, usually after something has already broken.

Flip it. Start with an ownership and escalation map. Start with the success metric. Start with who has the authority to pull the plug. If you can't name who owns the outcome and what happens when it's wrong, you're not ready to pick a model — and the model you eventually pick will matter less than you think. A well-run operating model can carry a mediocre model. A missing operating model will eventually break a great one.

The organizations operating ahead of the pack didn't get there with better models. Everyone has access to roughly the same models. They got there by building the operating model first and treating the model itself as the replaceable part: swap it out when a better one shows up, without redesigning ownership, escalation, and accountability from scratch every time.

Here's the concrete version for this week. Take one AI initiative you already have in flight. Write down, in a sentence each, who owns the output, what the escalation path is, and what the actual business success metric is. If you can write all three cleanly, you're already ahead of most organizations. If you can't, that gap is the finding — not a model score, not a vendor complaint, but the operating model you still haven't built.