Skip to content
Zhiyuan QidianFDE Outsourcing
Practice9 min read

Why AI pilots fail: 9 patterns and the early signals

Why AI pilots fail is rarely about the model. Nine failure patterns with their root causes and early signals, a scoping checklist, and how to win over users.

The short answer: almost nothing that kills an AI pilot is a model problem. More than 60% of enterprise AI attempts remain stuck in pilot. What stops them is system integration, permissions, data quality, process change and whether the organisation actually cooperates. All five are human problems, which is also why they are visible before you spend the money.

Below are the nine patterns that appear most often in public reporting and practitioner interviews. Each one is set out the same way: symptom, root cause, early signal, and what to do about it.

Why the blockers sit outside the model

Model capability improves by the quarter. Getting a model into one organisation’s systems has not got any faster. In a survey of 1,500 FDEs, the time split was 47% customer communication, 31% writing code and 22% internal coordination — far more time spent on interfaces, permissions, definitions and people than on making the model better.

Which is why “let’s use a stronger model” rarely rescues a stalled pilot. The two tables with similar field names will still collide, and the acceptance criteria nobody wants to sign off will still be unsigned.

Nine failure patterns

1. The demo data is not the real data

  • Symptom: the demo runs beautifully; the moment it meets live systems, it produces wrong answers.
  • Root cause: demos use a cleaned single sample. Real data has one field in several formats, large blocks of missing values, and rules that were abandoned years ago but never deleted.
  • Early signal: the person who built the demo will not run it against a read-only production account, and cannot say how many spellings a given field has.
  • What to do: make “runs on real data” an acceptance criterion for phase one, and budget separately for data preparation.

2. Technical metrics are reported; business metrics are not

  • Symptom: the review covers accuracy, latency and call volume. Nobody can say whose time was saved or what it cost before.
  • Root cause: accuracy and latency are intermediate metrics. Lower cost, higher throughput and more revenue are the point. Intermediate metrics can look excellent while nothing changes in the business.
  • Early signal: the goal reads “accuracy above X” and never “step Y in team Z goes from this long to that long”.
  • What to do: measure the manual baseline first — who does the task, how often, how long it takes, what an error costs. Without a baseline, no efficiency claim can be proven later.

3. The front line resists, and nobody saw it coming

  • Symptom: the system ships, the process does not use it. Training happens, accounts are issued, and the work continues in spreadsheets and chat groups.
  • Root cause: staff read an outside team as “the people the boss brought in to optimise my job”. When a system may remove a role, or dismantle the way a manager currently runs things, resistance is rational rather than emotional.
  • Early signal: the main contact only ever reports good news; key teams keep saying “our situation is different”; nothing is raised in review and everything surfaces after go-live.
  • What to do: describe how jobs will change before you describe what the system can do. Start with the team that already finds the task tedious.
  • Middle managers need a separate conversation. If the system changes their scheduling authority, sign-off power or headcount, they can keep a project moving on paper while it stands still in practice. Ask them directly how their own targets change — an all-staff email will not settle it.

4. Usage is forced through KPIs

  • Symptom: leadership mandates “everyone uses it, 30% more output this month”. Staff cannot remember their login, and answer the target with staged screenshots or piles of low-quality output nobody reads.
  • Root cause: usage rate became the goal. It is the easiest metric to fake, and when training investment is zero, form is all anyone can deliver.
  • Early signal: the appraisal sheet lists generation rates or conversation counts, and meetings and overtime both increase.
  • What to do: do not measure usage. Measure the quality and elapsed time of one specific step. Pay for the accounts and the compute centrally.

5. Definitions differ, so nobody trusts the numbers

  • Symptom: someone asks for last month’s sales. The generated query is valid and the number is wrong, because “sales” means something different in each system and each department.
  • Root cause: this is metadata governance, not model performance. Where several versions of an indicator exist, the model simply produces disagreement faster.
  • Early signal: two departments’ monthly reports do not reconcile, and someone asks which system the figure was exported from.
  • What to do: make definition alignment a phase-one deliverable, signed off by the business. Do not let the model compute metrics that have no agreed definition.

6. Nobody owns the outcome

  • Symptom: the post-mortem has a dozen attendees and everyone is proving it was not their fault. The business blames usability, IT blames unclear requirements, the vendor blames uncooperative workshops.
  • Root cause: the project was treated as a delivery, not a business change. Without an internal owner, nobody has the authority to stop it or to commit more resource when it underperforms.
  • Early signal: ask who owns it if it succeeds and if it fails, and the answer is “all of us together”.
  • What to do: name an internal owner before kick-off, in the project charter, with the authority to pull in the business, sign acceptance and call a halt.

7. No acceptance criteria, so the project never ends

  • Symptom: the build is finished, then “change this bit” and “the results still are not there” push acceptance back indefinitely.
  • Root cause: acceptance criteria were never front-loaded, and were written as subjective statements (“client satisfaction”, “significant improvement”) that are guaranteed to become an argument.
  • Early signal: no measurable acceptance clause exists in the contract, and requirements keep growing mid-build.
  • What to do: agree four things before kick-off — how the metric is defined, how data is collected, what the comparison baseline is, and what happens if the target is missed. Keep each phase stoppable.

8. The cost is pushed onto individual employees

  • Symptom: AI usage is written into performance reviews while the expenses form has no line for it. Heavy users buy their own compute.
  • Root cause: the budget covered tool purchase but not running cost, or decision-makers assumed staff would sort themselves out.
  • Early signal: people describe paying for tokens out of their own salary so the company can look more productive.
  • What to do: put inference cost, accounts and compute in the project budget and pay for them centrally. Ban company data from personal accounts — that is a cost control and a security control at once.

9. Go-live is treated as the finish line

  • Symptom: the launch is celebrated, and three months later a model upgrade, a shift in the data or a change in business rules leaves the system unattended.
  • Root cause: deployment is roughly 20% of total cost. The other 80% is keeping the system running. Most contracts price only the first part.
  • Early signal: no operating-period clause in the contract, no monitoring dashboard, and the answer on response times is “we would need to check”.
  • What to do: agree the owner, response times, evaluation cadence and model-upgrade policy before launch, and approve that budget alongside the build budget.

One-page summary

Pattern Early signal Response
Demo and real data diverge Will not run against production read-only Real data in phase-one acceptance
Technical metrics only Goal states accuracy alone Measure the manual baseline first
Front-line resistance “Our situation is different” Explain how jobs change first
Middle-manager drag Reports “everyone is using it” Align their own targets separately
Usage forced by KPI Appraisals list generation rates Measure output, not usage
Conflicting definitions Two monthly reports disagree Definition list signed off early
No owner “All of us together” Name an internal owner
No acceptance criteria “Client satisfaction” as a clause Metric, baseline and fallback first
Cost pushed to staff People buy their own compute Running cost in the budget
Launch as the end No monitoring, no operating clause Owner and budget set at launch

Projects that should not be approved at all

  • The goal can only be written as an adjective. “Improve efficiency”, “embrace AI” — nothing to accept against, nothing to judge progress by.
  • There is no baseline and no plan to measure one. You will not be able to show it worked.
  • Nobody can stop it and nobody will own it. The sponsor appears at the kick-off and the business agrees only to “support”.
  • Data may not leave the network, and no internal environment is provided. Security requirements and delivery conditions contradict each other, and the project proceeds on verbal assurances. Our advice is not to start.
  • The plan is to build first and decide the use case later. Without a specific workflow and named users, it will not be used.
  • The budget only covers people. If it buys capacity and nothing else, capacity is what you get — no assessment, no pilot.

Five things that make the front line willing to use it

  1. Describe how jobs change before describing capability. Left unexplained, people assume the worst, and the worst they imagine is usually worse than the reality.
  2. Start with volunteers, not with “everyone”. Remove the most tedious step for one team that already dislikes it.
  3. Make the tool useful to the employee first. Resistance falls when AI genuinely helps the person using it. If it only produces more reporting material, it is just another burden.
  4. Protect whoever says it first. The most valuable early information is “this does not work here”.
  5. Put the middle managers’ interests on the table. Ask how their targets, scheduling power and team size change. Leave it vague, and they will stop the project vaguely.

FAQ

How do we decide whether to keep funding a pilot? Three tests: somebody uses it daily; a business metric moved (not a technical one); and someone has approved the operating budget. If two of the three are missing, re-scope rather than spend more.

Is resistance just a training gap? Usually not. Training fixes “cannot use it”. Resistance is about “does not want to use it”, which comes from changes to roles, authority and targets. Deal with the interests before the training.

Would a better model fix it? Rarely. In stalled pilots the blockers sit in integration, permissions, definitions, ownership and people. A model upgrade also resets your baseline, undoing the tuning you already paid for.


Not one of the nine patterns has model capability as its root cause. That is why we insist on an assessment before delivery: settle the failure patterns, the acceptance criteria and the data boundaries first, then decide whether to spend. If a pilot has already stalled, look at how our FDE outsourcing service runs the work, or read how we deliver. New to the model itself? Start with What is a forward deployed engineer? and FDE vs outsourcing. If you are stuck on something specific, tell us about it.

Sources

  1. [1]前线共创,双向赋能:FDE 模式行业观察与实践报告 (Forward deployed engineering: an industry review) — Tencent Research Institute
  2. [2]那些被管理层请进公司,却不被一线员工接受的 FDE 们 (The FDEs management invites in and the front line will not accept) — Huxiu
  3. [3]把 FDE 送进企业之后:谁救火,谁背责,谁赚钱?(After the FDE arrives: who fights fires, who carries blame, who profits?) — 36Kr
  4. [4]Computerworld: reporting on enterprise AI deployment costs and FDE teams — Computerworld

Related reading

← Back to insights

Your situation is more specific than an article

Describe where you are stuck. We will tell you whether we can help and whether it is worth doing.

contact@mail.kchangai.com