A new report from Ness Digital Engineering, covered widely this month, puts a number on something a lot of business owners already suspected: almost every company says it plans to run AI agents in production, and almost none of them actually have. The report finds that roughly 99% of companies plan to deploy agentic AI, but only about 9-14% have fully done so. Ness calls the gap between proof-of-concept and real production use "Death Valley," and it is a useful phrase because it captures exactly what happens to most AI pilots: they launch with enthusiasm, run for a few weeks, and then quietly stop moving forward.

If you have a pilot sitting in that valley right now — a chatbot nobody uses, an automation that ran twice and got abandoned, a "let's test AI agents" initiative that never got a second budget line — this data says you are in the majority, not the exception. The more useful question is why the gap is so wide, and what separates the 9-14% that make it across.

The Death Valley Numbers, and Why They Matter to You

The report's core finding is stark by itself, but the reasoning behind it is more instructive than the headline number. Ness attributes the implementation gap to two specific causes: a loss of trust in probabilistic solutions built on generative AI, and a lack of visible difference in the daily user experience of the people who are supposed to benefit. In other words, pilots are not dying because the models are too weak. They are dying because the people using them do not trust the output enough to rely on it, and because even when the pilot technically works, nobody's day-to-day job actually feels different.

That second point deserves attention because it contradicts how most companies measure pilot success. A team automates a task, the automation completes the task, and the project gets marked "successful" in a slide deck. But if the person who used to do that task doesn't notice a change — if their queue is still full, their week still feels the same, their manager still asks for the same manual report — the organization has no internal evidence that AI is worth expanding. The pilot succeeded on paper and died anyway.

Editorial illustration of a glowing bridge crossing a dark valley, connecting a small AI pilot flag to a production platform on the far side
Most AI pilots don't fail from bad technology — they stall in the gap between "it worked once" and "people trust it enough to rely on it every day."

Why Pilots Stall Before They Scale

The Ness report notes that traditional financial institutions tend to point agentic AI at back-office efficiency and IT development, while digital-native fintech companies build it directly into customer-facing products. That split shows up outside financial services too. Companies that treat AI as an internal cost-cutting experiment run it quietly, measure it loosely, and let it die when attention moves elsewhere. Companies that treat AI as a product feature build accountability into the launch from day one, because a customer-facing failure is visible immediately.

Small businesses usually fall into the first category by accident, not by strategy. Someone tries a tool, likes the demo, rolls it out to a couple of people, and never defines what "working" means. Without a baseline, a target metric, and an owner accountable for the result, there is no mechanism that pulls a pilot out of the valley. It just sits there until someone forgets about it or a renewal invoice forces a decision.

The Small Business Edge Nobody Talks About

Here's the part the coverage of this report tends to miss: small businesses are actually better positioned to cross Death Valley than large enterprises are, if they run the pilot correctly. Large companies fail to scale AI agents partly because of organizational friction — procurement cycles, competing department priorities, and layers of approval that slow a working pilot down long enough for enthusiasm to fade. A 15-person company doesn't have those layers. If the pilot works and the owner decides to expand it, it can be expanded next week.

The tradeoff is that small businesses also have less room for a pilot to fail quietly and get a second attempt. That is exactly why the structure of the pilot matters more than the sophistication of the AI tool. Our workplace AI agent pilot playbook lays out the mechanics of running a 30-day pilot with a defined baseline and a scale decision built in — the same discipline the Ness report says most companies skip.

The Readiness Checklist Before You Scale Anything

The report argues that companies need a more structured approach before scaling agentic AI, built around a few concrete readiness questions: Is the business domain well understood? Is the existing technology infrastructure ready? Are the underlying processes documented? Is the data clean enough to trust? Are the APIs available to connect the agent to real systems? And, critically, is the organization actually willing to change how work gets done, rather than just bolting AI onto the existing process?

That last question is the one most pilots skip. It is tempting to point an agent at an unchanged process and hope the tool absorbs the mess. It rarely does. The businesses that cross into production are the ones willing to redesign the workflow around what the agent is good at, not the ones trying to make the agent mimic exactly how a human used to do the job. If you haven't defined the allowed actions and approval gates for your pilot yet, our AI agent governance guide walks through exactly that structure — what an agent can do alone, what needs a human sign-off, and what should stay off-limits entirely.

Stuck in your own AI Death Valley?

If a pilot stalled, or you're not sure how to structure the next one so it actually reaches production, we can help. Book a free strategy call and we will map the workflow, set the success metric, and build the pilot to cross the gap.

Book a Free Strategy Call →

Measure the Right Thing: Experience, Not Just Task Completion

The report's most important recommendation is also its most counterintuitive: companies should judge AI projects by whether they improve the experience of the people using the system, not by whether a single task got automated. A lead qualification agent that saves the sales team ten minutes but requires everyone to check a new dashboard, learn a new interface, and manually reconcile results against the CRM has not improved anyone's day. It has added a step. The report recommends using an interface people already work in — an existing messaging tool or inbox — as the single point of interaction, with the AI system pulling from multiple backend sources and using APIs behind the scenes. The person doesn't need to learn a new tool. They need their existing tool to suddenly do more. That is a very different design goal than "does the automation technically run," and it is the goal that determines whether a pilot survives contact with real daily use.

Cost Visibility Is Part of the Scaling Decision

The report also flags cost tracking as a scaling blocker: companies need to monitor LLM token consumption, AI subscription costs, and cloud spend at the application level so they can compare actual spend against expected return. Without that visibility, a pilot that looked cheap in testing can become an unpredictable line item once usage scales — and an unpredictable cost is one of the fastest ways to get a pilot killed by finance before it ever reaches full production. We covered this in more depth in our AI agent cost controls guide: track spend per workflow before you expand a pilot, not after.

The Agentic AI market in financial services alone is projected to reach $33.26 billion by 2030, according to the same report, which is part of why boards and investors are pushing harder for evidence that AI investments are producing results rather than generating demos. That pressure is coming to small and mid-size businesses too, just in a different form — it shows up as "why are we still paying for this tool if nobody uses it," asked by whoever signs the checks.

How to Skip the Valley Instead of Falling Into It

The companies in the 9-14% that reach production did not get there with a better model. They got there with a narrower, better-defined pilot: one workflow, a documented baseline, a person accountable for the result, an interface people already use, and a real decision point at 30 days to scale or stop. That is a project management discipline, not an AI capability. It is also exactly the kind of discipline a small business can apply faster than a large enterprise weighed down by approval layers.

If your last AI pilot went quiet without anyone officially killing it, that is Death Valley, and the fix is not a smarter tool — it's a better-run second attempt. Define one workflow, measure the baseline before you start, put the agent inside a tool people already use, and set a hard date to decide whether it scales or gets shut off. Do that, and you skip the 90% failure rate the rest of the market is currently living with.

If you want help getting your next AI pilot across the gap instead of stalling out in it, book a free strategy call at apolloagent.ai. We will help you define the workflow, set the baseline, and build a pilot designed to reach production — not just to demo well once.