Key Points
- Six months after a deployment that succeeded on adoption, compute was running at a 25 million dollar cost, and a re-architecture brought it to around 2 million.
- Many pilots died for one of two reasons, missing data to drive the agent or a problem too small to justify the investment, and selecting the few agents that matter now beats generating many.
- Legacy platforms stay in place as systems of record while the engagement layer around them is rebuilt, rather than being replaced outright.
Andy Baldwin is senior vice president of consulting offerings and growth for IBM Consulting, a business the host introduced as having "$21 billion in revenue." On CXOTalk episode 924 he returned repeatedly to one phrase that describes why AI programs stall even when the tools are in place. "AI is a contact sport," he said. "It is not something that you can then just, you know, throw over the fence and hope it's going to be used, hope it's going to be adopted."
Adoption requires changed behavior, not distributed licenses
Baldwin drew on a deployment at his previous firm: "we deployed at scale, you know, 150,000 Copilot licenses across the organization to really boost personal productivity." In a professional services business, he noted, "one or two hours of, you know, improvement actually does immediately add to the bottom line because you're effectively charging by the hour in that business model."
The lesson he took was about people rather than provisioning. Personal productivity "does require changes in individuals' personal behavior," so the work is "how you encourage that, how you incentivize that, how you make the training and development available." Usage is part of the target and not the whole of it: "You need people to be playing around with it. You need people to be using it, and you also need people to be getting to the point where, you know, they start to build their own personal productivity agents."
One dynamic is genuinely new. Teams can "very quickly create an MVP or a visual of this is what the agent can do," which he called "incredibly powerful to get people engaged," though faster prototypes do not mean faster production: "you can develop the agents very quickly, but some of the integration activity and the security provision still needs to happen."
Scale rewrites the economics
The most concrete part of the conversation was a cost story from Baldwin's prior organization. Unable to allow public chat tools, the firm deployed an internal version instead. "I deployed it to 3, 400,000 people," he said. "Actually, within 6 months, we were running at a $25 million compute cost."
He was clear that the spend reflected success, not failure: "the great news is everyone was using it. Nobody was using any tool that they shouldn't be using. But all of a sudden, the cost ramped up because that's scale." The architecture had simply never been built for the volume. He put the cause plainly: "we probably didn't architect it as efficiently as we could do for token consumption 'cause we never expected the consumption to be that high that quick." After a rebuild, the firm "brought the cost down to around 2 million."
That experience explains his framing of the pilot-to-production shift: "When it's proof of concept, it's relatively small amount of investment that you're making. Once you go to scale, it's completely different."
Match the model to the task
Baldwin's illustration of waste at scale was domestic. The phrase he uses is that "people are using a Ferrari to go down to the corner shop to buy a paper." As he put it, "You may feel great in the Ferrari, but it's probably not a particularly good use from a cost point of view." Sometimes the right answer is a smaller language model, or "some of, you know, machine learning or robotic," rather than "an expensive frontier model."
Doing that requires seeing what is actually running. "I can go onto my observability layer, and I can see the 60 models that we're currently using across IBM. 60 models," he said, adding that he can "see which models are being used by which tasks and activities," and can "dial down access to one of the models" when the spend is not justified. He does not expect users to solve this themselves, because "somebody needs to be directing the user."
Democratization moved the work out of IT
The reason visibility became urgent is that this technology reversed the usual sequence. Older systems were built centrally and pushed out, whereas AI "is being deployed and then being developed," applied "by line of business owners, by people in the finance function, the talent function, as well as in operations."
When the host suggested this resembled "the old shadow IT, but now on very significant steroids," Baldwin accepted the comparison and drew the governance conclusion: "Once you democratize something, you need to reimagine how you think about governance, oversight, you control." He put the coming volume in numbers, saying IBM research showed "most major enterprises will be deploying close to 15 hundred, I would call, enterprise agents by the end of 2026."
Why many pilots died, and what replaced them
Asked what separated successful pilots from those that quietly disappeared, Baldwin named two causes without claiming they were the only ones. "I think a lot of proof of concepts failed," he said, "because they either failed because they never really had the data to be able to drive the agent, or effectively they developed a proof of concept that was trying to solve a problem that when you got into it was not that material enough a problem to justify the investment."
The hunt for candidate workflows has since been automated away. Tools "will basically generate the agents based on your workflow. So that activity's gone." The scarce skill now is triage: choosing, out of twenty possible agents in a function, "what are the top 2 or 3 that are going to have the biggest impact?"
On legacy systems he described a pattern rather than a rebuild. Organizations keep the old platform "as a system of record because it's functionally, you know, it's efficient," while reimagining how people reach it. "So you're keeping the legacy. What you're doing is you're wrapping new technology, new ways of interacting with that technology through the engagement layer."
Accountability has no single owner
Asked who is accountable when an agent makes a costly mistake, Baldwin declined the clean answer. If a firm builds its own advice agent, responsibility sits with the firm. "They've done it, they've programmed it, they've processed it." Agents that arrive "out of the box from the vendor" leave the vendor responsible. His conclusion: the answer "is contextual depending on the agent and who produced it and the role it's playing."
He was equally candid about the planning horizon, quoting colleagues who told a room of clients: "In certain areas, it'll be hard for us to see beyond 12 months." He called that "incredibly uncomfortable to say."
Asked for the one thing CIOs should do now, he returned to where he started, saying "the CIO can deploy the tools, can deploy the models, but the business needs to be engaging, needs to be using it, and needs to be getting the value out of that," with the results fed back so the technology organization "can then, you know, make the appropriate adjustments."
Watch the full conversation with Andy Baldwin and read the complete transcript on the episode page.
CXOTalk prepared this article with AI assistance from the verbatim transcript of episode 924. Quotations are unedited from the broadcast.