Over the past two years, enterprise AI initiatives have shifted from experimental efforts to essential requirements, with many large organizations establishing pilot projects as a key step toward full deployment. While most pilots show positive results, only a few reach production.
To understand the challenge of scaling AI in the enterprise, I reviewed CXOTalk interviews with leading enterprise AI practitioners and executives.

We can see the patterns and gaps these experts identify across industries, companies, and use cases. Each of the following gaps appears harmless at pilot scale but becomes a cause of stalls when production scales.
1. Teams automate existing workflows without rethinking from first principles.
Most organizations build AI pilots on top of existing workflows without reimagining them for AI. In the pilot lab, this does not have a significant impact because the number of cases is small and workflow gaps have little effect. When teams push to production, the same process now meets far more volume at machine speed, exposing these gaps.
Bill Briggs, CTO of Deloitte, says the mistake happens before any technology decision is made. He urges people to ask, "How do we simplify based on first principles, which is the outcome, not the imagined constraints that we've got to do all 10 steps the way we've always done it?"
Sangeet Paul Choudary asks the question most teams skip entirely: "What was the problem in the organization this workflow was set up to resolve? And does that problem still exist now that AI comes in?" The pilot proves AI is ready. Production exposes why the workflow is not.
2. Investment in tools rather than the operating model.
For pilots, AI budgets fund only the tools. The tool only has to do the tasks it is scoped for, and nothing around it has to change. When teams push the same tool to work across functions at a much higher volume, it fails. The tool lacks support from the process and operating model that were never redesigned for AI.
Deloitte's Tech Trends research found the ratio behind this pattern: 93 percent of AI spend goes toward technology and tooling, only 7 percent toward the culture, process redesign, and operating model work underneath it. The tool is the primary investment for the pilot to work. For production, everything around the tool has to change, too, and nobody budgets for that.
3. Tokens are consumed before anyone defines clear outcomes.
Many organizations enable agent access before deciding what constitutes a good value for token spend. At low volume, the pilot incurs little cost. But as agents scale, they spend multiple times faster than we are ready to measure their value.
Praveen Akkiraju, Managing Director at Insight Partners, observes this pattern, called tokenmaxxing, recurring across many portfolio companies. Every token an organization makes available simply gets used, with no link back to value. His diagnosis is direct: "An agent deployed without a harness is ungoverned, technically and financially."
Raffi Krikorian, CTO of Mozilla, names the same failure from a different seat. Once a prompt is set off, everything that happens inside the harness "could just exponentially grow, and that is also a little bit out of my control." He calls the underlying arrangement renting, not owning. Infrastructure whose pricing and behavior can shift at the provider's discretion. Even where that spend is tracked carefully, it is not the same as knowing what it buys.
Tim Crawford, CIO Strategic Advisor at AVOA, puts the resulting gap in one number: 88 percent of companies use AI, fewer than 6 percent report measurable value. As organizations move from pilot to production scale, spend grows exponentially, and the value produced remains unmeasured.
4. Governance is built for a slower world.
Most governance processes are built around scheduled reviews and audits running on a calendar. A single-pilot agent, watched closely, fits within that cadence. As teams multiply agents and accelerate the decisions agents make, the governance cadence does not automatically speed up to keep pace.
Tim Crawford and Anthony Scriffignano (Distinguished Fellow at The Stimson Center) describe governance as the piece that every organization underbuilds relative to the speed at which agents are actually deployed. At scale, nobody is watching closely anymore, and the structures in place were never built to govern agents.
Where every unaddressed gap finally gets stopped:
Scaling exposes all four gaps, but deployment is where the gaps get caught.
Organizations rarely discover a weak foundation through examination alone. They encounter challenges when a deployment review halts the release to production due to questions nobody can answer.
Treating AI strictly as a technology rollout leads to these hurdles. Stated plainly, the fix is “Treating AI as an operating-model transformation,” which means:
- Rethink the problem before automating it
- Build the foundation and the harness before the agent goes to production
- Decide what value means before spending on tokens
- Treat governance as continuous, not a single gate.
Identifying these issues is easy compared to building by design, as evidenced by the experiences of organizations that have actually done it.
And that’s where this series goes next, so stay tuned!