Why Enterprise AI Pilots Stall in the GCC

Most large organizations in the Gulf no longer need convincing that AI matters. Boards have approved budgets, vendors have delivered proofs of concept and many groups can point to a portfolio of pilots. What far fewer can point to is an AI initiative that became part of how the business runs. That gap is where most of the difficulty with AI adoption in GCC enterprises now sits, and it is usually described as a single problem: pilots that never reach production.

My argument in this piece is that it is two problems, and that treating them as one leads organizations toward the wrong remedy. Some pilots fail because they should never have been approved in the form they took. Others succeed and still go nowhere because nobody decided in advance what success would require. The first is a problem of selection and the second is a problem of commitment. They call for different decisions, taken at separate points by people in quite different roles.

What the evidence says about AI adoption in the GCC

BCG’s January 2026 study of GCC organizations classified 39% of them as AI Leaders and 61% as Laggards, a split almost identical to its global sample. Given the scale of national investment in the region, that parity is worth noting. The Gulf has not yet turned its ambition into an advantage in value capture, but neither is it uniquely behind. This matters for reading the rest of the evidence, because most of the detailed research on why AI projects fail comes from North American and European samples. The regional data that does exist suggests the same patterns hold here.

That research is less flattering about the pilots themselves than the usual “stuck in pilot” narrative suggests. Gartner predicted in 2024 that at least 30% of generative AI projects would be abandoned after proof of concept, citing poor data quality, inadequate risk controls, escalating costs and unclear business value. S&P Global’s 2025 survey found that 42% of companies had scrapped most of their AI initiatives, up from 17% a year earlier, with respondents abandoning on average 46% of their proofs of concept before production. Cost, data privacy and security topped the list of reasons.

Read together, these findings do not describe an industry full of working pilots waiting for a signature. They describe a large share of pilots that did not deserve to proceed, alongside a smaller share that did and still stalled.

Some pilots should never have started

The first failure is one of selection, and it is the one the “stuck in pilot” framing tends to hide.

A pilot can be well run and still be the wrong pilot. Organizations often choose the use case that will be noticed rather than the one with the strongest economics, favoring customer-facing showcases over less visible back-office work such as document processing, where savings are easier to measure. When such a pilot is abandoned, the cause is not a lack of follow-through. The initiative was never likely to pay for itself.

The form of delivery matters as much as the use case. MIT’s NANDA initiative, in a widely cited 2025 report, found that AI tools delivered through external vendor partnerships reached deployment about twice as often as those built internally. The study drew on a modest sample of 52 organizational interviews and 153 survey responses and has been criticized for its methodology, so its figures are best treated as directional. The direction is still instructive. Commissioning a bespoke system for a problem that a mature product already handles is a decision that should be made deliberately, and it is often made before anyone has asked whether building is necessary.

Data is the third factor and, in this region, perhaps the most underestimated. PwC’s 2026 CEO Survey found that only 16% of GCC CEOs agree that their most-used AI tools have access to all the relevant documents and data, citing data silos, legacy systems and governance constraints. A pilot run on a curated extract can perform well in a demonstration and then degrade once it meets production data. Language adds a further layer. I wrote in May about why bilingual markets need language parity built into the architecture of an AI program, so I will only note the practical consequence here: a pilot evaluated on clean Modern Standard Arabic or on English has not been tested against the Gulf dialect, mixed-language text and scanned documents it will actually receive.

Not every failure fits neatly into selection. Some pilots address the right problem with the right approach and are simply executed poorly, whether through weak vendor delivery or tools that never improve with use. These are harder to spot early, which is why a pilot needs a realistic evaluation set and a stopping threshold from the start. Stopping a pilot for any of these reasons is not a failure of AI adoption. It is the selection process working, if somewhat late. The cost lies in how long such pilots run and how much they consume before anyone decides.

Adoption may be happening outside the program

There is a further possibility that formal AI programs tend to overlook. PwC’s 2025 Middle East Workforce Hopes and Fears Survey found that 75% of employees in the region had used AI tools in their roles over the previous year and 32% used generative AI daily, both above global averages. The survey does not distinguish sanctioned tools from personal ones, so it shows readiness rather than unofficial adoption. It does mean the workforce is not the obstacle.

Leaders should therefore check how much of their organization’s AI use already happens outside approved channels. Where it does, a business unit’s lukewarm response to a formal initiative may reflect a better alternative already in hand rather than inertia. That pattern is useful information about where demand exists and what standard an internal tool must meet. It is also a governance exposure, since customer records pasted into personal accounts sit uneasily with data residency obligations. Prohibition tends to push such use further out of sight. Sanctioned tools that are at least as convenient usually work better.

When a pilot works and still goes nowhere

The second failure is the one the “stuck in pilot” narrative describes accurately, and it is the more frustrating of the two because the hard part has already been done.

The most common cause I have encountered is ownership that ends with the vendor contract. Responsibility is typically divided so that IT owns the platform and the security review, an external vendor owns the build and a business unit owns the original request. When the pilot concludes, so does the vendor engagement. IT has no budget to operate the system, and the business unit was never told that the cost of running it from the second year would fall on its own line. The system remains technically functional and organizationally orphaned.

Procurement can compound the problem. Many procurement processes were designed to acquire assets with a defined scope, while AI products behave more like services: their quality depends on data the vendor has not yet seen, their pricing is often tied to usage and they need iteration after launch. A pilot bought as a fixed-scope engagement can succeed and leave the organization with no contractual route to scale it. Writing the production pathway into the pilot agreement, with conditional phases, indicative pricing at scale and handover terms, is usually far quicker than starting a new process afterwards.

How success is defined matters too. Pilots described by what the system can do, such as answering customer questions in two languages, can be proven in a demonstration. Pilots described by what will change in the business, such as handling time on a specific category of complaint, force a decision about production.

Pilot debt

Pilots that fail on their merits and are stopped promptly do little lasting harm. The ones that do harm are left in limbo, neither scaled nor cancelled, and I have come to think of their cost as pilot debt.

Pilot debt is what an organization learns when AI initiatives end without a decision: that AI is something the company announces rather than something it operates. Business units become less willing to offer their workflows for the next initiative, data owners take longer to approve access and each new sponsor inherits the skepticism left by the last. None of this appears in a project report, yet it shapes how the next program is received. The remedy is not to avoid failure. It is to make decisions, including the decision to stop.

Deciding before building

Most of the decisions that prevent both failures can be taken before any vendor begins work.

The first concerns form: whether the problem calls for a bespoke build, a purchased product or a capability that already exists in a platform the organization licenses. It is among the decisions that most affect the odds of success. Scope comes next. A pilot framed around a broad use case such as customer service AI is much harder to evaluate than one framed around a specific workflow, for example drafting first responses to Arabic-language card dispute emails for agent review. A workflow has an owner, a volume and a measurable baseline.

The organization should then agree on the result that would justify stopping, and name a business leader who will decide whether the system ships and who accepts its run cost from the second year. If no leader will take that role, the organization has learned something about the initiative’s real priority before spending on it. Data access should be confirmed against production systems rather than extracts, and the evaluation set built from several hundred real examples in the language and format users actually produce. The people whose work will change also need to be involved before launch. BCG found that the region’s AI Leaders are three times more likely than Laggards to run structured learning programs with protected time for staff, which suggests the investment tends to accompany success.

I find it useful to capture all of this in a single-page pilot charter, signed by the sponsor before work begins. It records the delivery form, the workflow, the baseline, the target and stopping threshold, the owner and budget, the data access arrangements, the evaluation set and the production pathway. A blank section is an early view of the risk register. Some sponsors will decline to sign, and that is a useful outcome in itself, because a sponsor unwilling to commit at the start would likely have withdrawn the production budget later.

The second-week test

The charter addresses the first failure by forcing the question of whether a pilot should exist at all. For the second, the question I find most diagnostic is what is expected to happen in the second week after the demonstration succeeds. A well-prepared plan answers in operational terms: who owns the outcome, what budget funds the system, how the contract covers scaling, when the first users begin and which measure will decide whether the work continues. A less prepared plan tends to answer with a further demonstration.

The organizations in the region that convert AI investment into results will not necessarily be the ones that run the most pilots. They are more likely to be the ones that are selective about which pilots they start, honest about stopping those that fail and clear, before anything is built, and about what success would commit them to doing next.

Stay Updated

Enjoyed this article?

Get more insights on AI, product strategy, and digital growth delivered to your inbox.

No spam. Unsubscribe anytime.

Share this article