Why AI Pilots Stall:
How to Make Enterprise AI Work
Most AI pilots never reach production, and the ones that do rarely scale. AI strategist Nate B. Jones explains why pilots stall and what the companies that succeed do differently.
-
AI Analyst and Advisor
-
How to make enterprise AI work is the question that matters in 2026, because most AI pilots never survive the transition to production and scale. Companies may prove value in demos, yet few turn that value into systems that run every day within real workflows, highlighting the need for strong organizational support to succeed.
This conversation examines why pilots stall, what must be in place around the model for AI to survive in production, who is accountable when an AI agent drifts, and what an executive should measure, emphasizing the importance of clear accountability for trust and responsibility.
Key Points
- Why most AI pilots stall before production, highlighting both the organizational and technical challenges that interfere with scaling.
- What AI needs to survive in production: context, memory, reusable procedures, and review gates that ensure ongoing oversight beyond initial pilots.
- How to ensure that pilots do not drift from their core mission.
Most enterprise AI pilots never survive the transition to production and scale. Even when interest is high, obstacles remain because pilots often don’t reflect real-world workflows, financials, permissions, or accountability scenarios.
In Episode 927, AI strategist Nate B. Jones examines why pilots stall and how to make enterprise AI work in production.
Jones led product for Amazon Prime Video and now advises Fortune 500 companies, growth-stage startups, and global banks on AI strategy. He publishes AI News & Strategy Daily to more than 300,000 subscribers on YouTube and nearly 500,000 followers on TikTok. This is his third CXOTalk appearance, after conversations on the AI vendor landscape in May 2025 and on prompting in June 2025.
What we will cover:
- Why AI pilots stall before production, and the difference between a successful demo and a system that endures real-world work
- What AI needs to survive in production: context, memory, reusable procedures, and review gates
- Why companies cannot switch to a cheaper AI model even when it wins on price
- How accumulated prompts and rules become procedural debt that quietly degrades results
- Who owns an AI agent in production, and who is accountable when it drifts
- What business and technology leaders should measure to ensure their pilots remain on track
Watch Episode 927 with Nate B. Jones live on Friday, July 31, 2026, at 1:00 PM ET on CXOTalk. Join this interactive discussion and ask your questions!
Episode Participants
Nate B. Jones is an AI–first product strategist and advises large organizations on AI strategy and is a popular influencer on social media with hundreds of thousands of followers and subscribers. This is his third appearance on CXOTalk.
Michael Krigsman is a globally recognized analyst, strategic advisor, and industry commentator known for his deep expertise in business transformation, innovation, and AI leadership.
In This Episode
Nate B. Jones: AI pilots fail because they're pilots. That is the critical error that most leaders I talk to run across after the fact.
Michael Krigsman: Every company wants to scale AI, but most pilots stall before production. Nate B. Jones advises Fortune 500 companies and global banks. He has almost a million followers across social media.
Why pilots fail and where to start
Nate B. Jones: We envision the pilot as a way to de-risk AI, but what we find in practice is that by naming and defining it as a pilot, you end up putting less resources behind it than you should. You pick more fragile and less important goals than you should, and you don't get the learnings you want ultimately, because the question of AI is a question of whole organization transformation. And it's not something that you can effectively and easily sandbox into a little pilot space.
And so that's one of the first things I actually tell leaders when they talk about pilots. I'm like, I know with most software you want to do pilots, and that made sense at the time. But, you have to think about AI differently because it's just a different technology.
Michael Krigsman: You said that AI is really full organization transformation, but yet you have to start someplace. We can have dreams of grandeur, but they may be delusions of grandeur. What do we do?
Nate B. Jones: I have two places I tell leaders to start. The first is to look in the mirror and look at themselves. I talk directly to CEOs to start with, and I talk with the rest of the C-suite. And I say transformation starts with you. You cannot just be talking about AI. You have to be living and practicing it. And that means that for most people who are not named CTO, you're going to have to become more technical.
And so we talk through that and we talk through what it means to go toward being comfortable with Claude code, being comfortable with Codex. I'm not saying commit code. I'm saying being comfortable using the tools you are expecting people to use on a daily basis and be comfortable publicly talking about what's working for you and what's not and being a public learner.
Number two, I say instead of calling it a pilot and picking an area that you think is a little bit de-risked, I want you to pick something where if it works and you are putting organizational momentum behind it, you are going to see real leverage and momentum in the business as a whole. So pick something where the business goal matters, and that's where you start to go after it.
Michael Krigsman: And then how do you determine the size, the scale, the scope?
Nate B. Jones: What I find is almost everyone has a charge from their board at this point to say, you know what? We got to do AI. Almost everyone has a list of projects that you want to get done with AI, and maybe you read it off LinkedIn, or maybe you came to it yourself, or maybe your CTO brought it up. There's a whole range of ways those ideas come into the company.
The thing that you do to pick the right project to work on, the right skies, the right scope, is you want to identify places where if this works, it's going to be transformational, and then you can name the inputs that would make that successful, and you can understand where each of those inputs in the business comes from. And so this is where it gets into a little bit of process mapping.
But if the process is tremendously ambiguous in terms of value creation in the company right now, that's an area where I'm like, okay, probably very important to get AI in there. I would not start there. I would start instead with something that's high leverage, but the processes are at least understandable to people in the company, and you can define the inputs so you can start to map them in and start to transform them with AI very deliberately.
Adoption is mostly a people problem
Michael Krigsman: Are we talking then primarily about technology issues, leadership issues, organizational issues, AI specifically? Help us parse that out.
Nate B. Jones: Honestly, what I find is it's about 20% a technology problem and 80% a people problem, which is sort of the reverse of what a lot of people expect because, of course, LLMs are a new tech. They are absolutely disruptive and different from traditional software. The working assumption is often, you're telling me to become more technical, Nate. My people need to be more technical. It's a technical issue. And what I find in practice is it's actually a people issue first. And I talk about leadership already.
So you have to talk with your leaders and be really honest with them. And I find that it's really important if you identify this project space to make sure that your sort of director-level, senior middle manager folks are going to be champions in this space for you, and you invest the time with them to make sure they feel really comfortable with the change and what we're talking about. Because on the ground every day, they're going to be the ones that are championing this with your own teams.
And so that piece has to get done. And I just dive in. I'm happy to spend the time there to make sure that they feel really good before we go any farther. Because if they're not champions in practice, nothing works. And so that's a key from a people perspective. And then once you get to on-the-ground teams who are actually implementing this change, what you find is a very predictable bell curve of adoption.
And you do have people on the right tail of that distribution who are top 10, top 20% of your folks who are going to jump in. They probably are experimenting with AI at home, and they're like, oh my gosh, finally we're doing something about this on my team. I'm so excited. Easy to get them on board, easy to get them going. They're immediately productive. And then the trick becomes, what do you do with the middle of the distribution?
Because everyone knows on the left-hand side, there are people who are resistant, and we all have our conversations where we're like, you got to have tough conversations, etc. And we know how that goes. Everyone doesn't have trouble with that decision-making process. The middle of the distribution only starts to shift when you are able to help them understand how this is something that they, one, need to do for their careers. Two, it's very doable from a technical affordance perspective.
And so this is where your toolset for the middle of the distribution, the larger teams, needs to be less technical. For most teams, unless they're engineers, it needs to be a less technical tooling so that they feel like they can use it. Then you need to define what success looks like for them in a way that they can feel like, this is achievable. I can get this done.
Michael Krigsman: The folks listening right now will feel a certain amount of frustration because you are describing the dynamics of traditional technology projects and project success and project failure. What technologists want to talk about is models, codecs, AI labs. And so, you're taking us down the enterprise rat hole, and it's not as interesting or as much fun.
Nate B. Jones: And I wish I could tell you differently. And I have been the person in the room when people are like, but what if, you know, ChatGPT-6 comes out and it's incredibly good and it just solves all these problems for us because people are so excited to use it? But, Michael, we are 2+ years into this enterprise transformation. I have seen model release after model release after model release. I have talked about the transformation we experienced in January around agentic tool use and how huge that is.
None of it has changed the people dynamics in these businesses.
Michael Krigsman: We are dealing with a set of people issues and the technology opportunity is there, but if you don't have things organized in the right way, it's not going to work.
Nate B. Jones: Well, and you can look at it as the overhang, the potential of technology impact versus where people are continues to grow really, really fast. And so what that means is if you're in an industry where overall organizations aren't moving fast and you can figure out how to make this jump, the rewards for doing so are much higher today than they were even a year ago because there is so much technological overhang that you can use to accelerate if you can get on that train.
Michael Krigsman: Folks, if you're watching on LinkedIn, just pop your question into the LinkedIn chat. If you're watching on Twitter/X, use the hashtag CXOTalk and me directly, @MKrigsman, to make sure I actually get it. Zeya Ottomone says, this is on LinkedIn, Enterprises don't have an AI pilot problem. They have an operating model problem. Technology is rarely what prevents AI from scaling. Leadership alignment, government, and execution usually are the issues. That's where long-term competitive advantage is created. So, we're back in his comment to the core issue of aligning the technology project. That's a really good point.
Data, outcomes, and undocumented knowledge
Nate B. Jones: The one thing that I would add to that in practice is that in addition to fixing your operating model, the technical issue that comes up most often that prevents these pilots, these initiatives from succeeding is data.
If you don't have a clear understanding of how data flows in and out of the business and what you do with the data in the middle when you're transforming it and doing things with it with AI, then you are going to be in trouble from a technical perspective because you're going to have effectively a Ferrari engine with your LLM, but you're not going to have anything to feed it. You're not going to have confidence that you have production data access to whatever you care about.
Maybe it's Salesforce and you're doing a sales outbound thing. Maybe it's something with your codebase and you have to be sure that you can confidently put code in and out of the LLM and securely sort of access it and securely put it back into the codebase when you've transformed it or you've added some additional code. You have to go through that process and map it out in order to ensure that once you finally get into the technical details.
This is from Pravarshi Reddy Rachamallu who says, any tips on how to define outcomes for AI projects?
Michael Krigsman: We deal with probabilistic outputs but are held accountable to deterministic outcomes.
Nate B. Jones: I would recast that question a little bit, Michael. I think it's a fair question. It's absolutely true that we have nondeterministic outputs from LLMs, but what I find at the level of outcomes, like when you're talking about a business goal, let's say your business goal, we're back in sales. Your business goal is around transforming your outbound motion, putting LLMs at the heart of it. So LLMs are doing the account enrichment. LLMs are helping you with drafting outbound that feels personalized, etc., etc.
In that world, if you're talking about outcomes, you're way, way above the level of noise that comes with non-deterministic outputs, and there are lots of tools that you have at your disposal technically to make sure that those outputs don't change enough to make the outcome relevant. And so business leaders, I find, basically look at it and say, okay, I want my outbound motion to either scale out so I can talk to way more accounts or to become much richer so that the outbound motions I have are more personalized.
But either way, when you get back into the details of how you make that happen in practice, if you're feeding today's LLMs very consistent inputs, like you get them Salesforce fields and whatever else you want, they are going to give you very, very consistent outputs that are personalized in the way you want. And so as much as that is true from a mathematical perspective, I don't find it's typically relevant from a business perspective.
Michael Krigsman: This is from Arsalan Khan on Twitter. Arsalan says, documented processes are not always the ones that people follow, hence the term institutional knowledge. How does AI extract the information in people's brains?
Nate B. Jones: I would call out a couple of threads that are worth pulling out. I think the first one is, we now have super actionable and effective personal computer use for AI. And that's something that really came along as the Sky team, which was acquired by OpenAI, was able to use the technology and the talent on that Sky team to launch a really effective computer use product for Codex. And other LLMs are following suit.
So I'm not just calling out Codex because they're the only ones at this point, but they were the first ones to really launch fast, rapid, effective computer use. And now we're seeing the rest of the industry start to catch up as they start to look at the possibilities. And what that enables you to do, and I do this all the time, is it enables you to go into the actual places where work happens, where exactly to your point. The work is rarely what is documented.
The work is what is actually done in the mess. And it can look at all of the tools where work happens. It can look at Linear, it can look at Jira, it can look at Slack, literally anything on your computer, which means it can document how work actually happens rather than how work is supposed to happen. And so that's the first piece I would call out. The second transformative tech, which is also something that has become much more useful in the last six months, is voice.
And so I can use something like a Wispr Flow. There's now a lot of open source alternatives that are similar voice tech, and I can just talk, and I can talk for ten minutes, for fifteen minutes, for however long I want about all the mess in my process in a completely unstructured stream-of-consciousness way, the way our brains naturally talk when we're trying to just get something out. And I can give it to an LLM verbatim.
And they now have context window and reasoning capabilities that allow them to organize all of that in a way that reflects the actual process and saves me days. And so I think that there's some tech that's come out that's actually made that much easier in the last few months.
Michael Krigsman: Yeah, some of the technology, Wispr Flow, is, I use that as well. There's tools like Granola, which serves a slightly different purpose.
Nate B. Jones: Plaud is also good. I find that Plaud picks up and romanizes multiple languages if you're in a multilingual context. It can then feed it to an LLM, and the LLM can read the romanized text and can easily transpose between languages when it's giving you notes.
Michael Krigsman: Granola, when you go into a meeting, it connects to your email and other sources, and it gives you the context of the meeting.
Nate B. Jones: That's right.
Michael Krigsman: In a good way, in an unusually helpful way.
Nate B. Jones: Yep. It's also very, very useful.
Learning from pilots and proving value
Michael Krigsman: When an organization abandons an AI pilot, do we consider that to be a failure or is it good judgment?
Nate B. Jones: The question of whether it's a failure or not is really a function of the organization's ability to learn. Can the organization look at something and say, this is what we learned from this pilot. This is why we are walking away from this pilot. This is our next move in this space. This is the lesson we learned from a people perspective. This is the lesson we learned from a technology perspective. And this is what we're going to do differently. If you're doing that as an organization, it's absolutely not a failure.
It's actually a great learning opportunity. If you walk away and you say this wasn't for us and you don't actually engage in that active sort of deliberate leadership reflection, then I think you kind of wasted the opportunity. You lost the chance to learn that you really needed.
Michael Krigsman: And this is on Twitter from Gus Bekdash, who says, are AI breakthroughs in average enterprises top-down or bottom-up? He now thinks that both must happen for success. And how can leaders, regardless of formal power, make this happen?
Nate B. Jones: The most powerful organization transformations that I've seen are when you have activity and energy that you're capturing at the leadership level, but then also it's fomenting, it's coming up from folks who are on the front lines and they're seeing opportunities. And I would say I've seen great examples of leaders who are team-level leaders or director-level leaders who are able to help facilitate both. And the way they do that is, one, they're actively encouraging a culture of experimentation with their frontline teams.
And sometimes that means figuring out ways to get work done in creative AI-forward means, even if that's not officially blessed yet, and you just are finding ways to get that done effectively regardless. And then you're showing off the progress and then you get the blessing and then that sort of breakthrough starts to happen. I've seen that happen a lot.
And then on the other side, you're talking up to your VPs, you're talking up to your CXOs, and you're saying, you know, this is what I see as the strategic opportunity for AI in my space. This is why. And depending on how open they are, you may not lead with, and this is what we're doing about it yet. You may just talk about it as an opportunity that you're keeping an eye on.
And then at the right moment, you sort of bring forward this grassroots effort you've been nurturing, and you say, and this is what I've done to sort of push that forward. And this is the results that we've had. We've saved so many hours or so many dollars, and this is why it's been important. And look, it's in line with the strategy we've been talking about.
Michael Krigsman: At the heart, you're really dealing with a set of either explicit metrics or clear heuristics for evaluating this pilot.
Nate B. Jones: I think that you can't get away from that. I know there are, and I've been in the room when people are like, we just know we need to do AI, and so we're just going to do the pilot regardless. But I find that, at the end of the day, when you are evaluating the pilot, if you are not looking at the dollars, you are not getting very far in the actual ROI-focused budgeting conversation.
Because like it or not, we're still in a firm, we're still talking about capital allocation, and capital gets allocated where you can get a return. And so if you can talk in dollars, you get so much farther with the CFO, you get so much farther with the people who are actually figuring out the budget. That makes it so much easier to actually get a scale-up.
The harness and AI fluency
Michael Krigsman: Nate, you have said that AI in production needs context, memory, reusable procedures, and review gates. Can you explain this harness, as you term it, and the strategy, therefore, of enterprise AI that we should be following?
Nate B. Jones: I would almost start with a harness conversation as a business object in and of itself. If you think about the harness, you can think about the harness as a way of encoding the organization's understanding of how to do business, its competitive advantage around the LLM, so you can maximize the value of the LLM. But you, because you are building the harness as a team, are effectively building your own intellectual property.
And what we've found, and this is why Anthropic and OpenAI and others have launched their own forward-deployed engineering programs, they've realized in the last few months that you can't just stick a raw LLM into an enterprise and expect magic to result. You actually have to put engineers on the ground who understand how to build these context systems, these memory systems, how to figure out tool calling, all of that, in order to get to useful outcomes for that particular business.
And because business is incredibly detailed, you have to do that individually for businesses. And so that makes me really optimistic. People sometimes ask me, are like OpenAI and Anthropic and eat the world? I tend to say no because business is detailed and business has a lot of value that is hard to capture that, if we can build into these harnesses, we're going to have competitive advantages. We're going to be able to encode this in a way that allows us to leverage AI to accelerate without giving away our secret sauce.
Michael Krigsman: This is from R.A. Matthews on LinkedIn who says, what does it mean to be AI fluent? Is this even the same across domains. And I think he's going back to the earlier conversation that we had that, in order to run a pilot or get involved with AI effectively, senior leaders need to gain a greater technical understanding.
Nate B. Jones: You can measure it a bunch of different ways. It is true that there is domain expertise that's relevant. What I find, though, is that in most cases, when you're talking about people who have been in their careers more than, say, five years, and that includes senior leaders, individual practitioners, etc., they have that domain knowledge. They have a bunch of effectively their own secret sauce that they're bringing to the table.
And so when I'm working with groups like that, what I tend to do is to say, you guys have a great advantage. You guys already know so much about your business. It's about understanding how to translate that into AI terms. And the good news is with voice, with other ways of kind of getting media and domains and documents into the AI, it's never been easier.
You can just throw a bunch of what you know into the AI and learn to ask for things that are going to be apparently challenging or difficult for you, and it becomes a way of learning. And so examples sometimes help here, but I tend to tell people, Take something that would take you two days, three days, a week, and you think, this is going to take me a long time. I want to see if AI can do it.
Then work backwards from that goal and say, what is all the information I typically need in order to get that whole piece of work done? Go get that information. That might be outta your head. It might be in files somewhere. It might be on the broader internet. Whatever it is, fetch it and then throw it at the AI. Throw it at a current frontier model. And say in one shot, get this done for me. And I find that just asking. It teaches people to ask bigger.
It teaches people to think about data as an input, and it teaches people to recognize how powerful these models are. And it means that they start to think differently about their work as a result, which is really what you want them to do. Think about AI as a colleague now. And what does it mean to think about it as a colleague where you go to AI in the morning, you're like, this is what I'm working on. Help me get this giant thing that I'm working on done.
Can you work on it while I do something else? That's now where we are with these models, but it's a mental shift.
Michael Krigsman: It's a good idea to just simply take a problem and put it in there, but that raises the question of figuring out what is an appropriate kind of problem. That's a challenge because AI is so new.
Nate B. Jones: I tend to say, okay, just throw some problems at the wall and let's talk about the problems that are in your space and let's see which ones are AI susceptible. And what I tend to find when we have those exercises is that people don't realize the capabilities of today's models. And so they tend to under-ideate and they tend to imagine less than they would.
And I have to be encouraging and kind of pulling on, I'm like, no, no, no, think a little bit bigger, think more about these ideas and we'll just throw them around. Yeah. And, you know, to be honest with you, it is absolutely true even now that not every problem is equally AI susceptible. But since we last talked, a huge number of new problem spaces have opened up and are now available to work with on AI.
And I am at a point now where I am very confident I can sit down with anybody, roll the dice in the enterprise, and we can talk. And over half of the problems they are facing are going to be things that I feel very comfortable saying, throw that at AI and see what it can do. That was not true a year ago.
Michael Krigsman: The models have gotten so much better. I think it's these blind spots that we all have, the things that we take for granted, the assumptions that we make that we may not, therefore, consider as an option to put into the model because we take it for granted. Yet, that may be where the greatest opportunity lies.
Why a culture of experimentation wins
Nate B. Jones: This is why I come back, and I think you said it really well at the top of the podcast. It's really boring to talk about the people problems because that's something we've talked about in management theory for a long time. But this is a people conversation. It's a conversation about our ability to open up our eyes to a new way of working. You know, for hundreds of years, the firm has been about people figuring out how to get things done in teams together, given capital.
And now the firm is about people figuring out how to partner with artificial intelligence to get things done given capital. And that is entirely new. It's a new thing that we all have to figure out together. And the companies that are doing the best at it and getting all the accolades are really just companies that have had a strong culture of experimentation for a long time and have been really relentless about rewarding that.
And so I like to joke, I'm like, look, I know everyone thinks Anthropic ships really fast, and OpenAI is obviously very good at AI, and there are other companies out there as well. There are YC and Silicon Valley startups, etc. You think, well, they've had more time with this, but the reality is they kind of haven't. AI has been out there for all of us for the same amount of time. What they're doing differently is they are finding people and encouraging a culture of experimentation and rewarding that.
Michael Krigsman: Sometimes you hear the term an AI-first culture to describe exactly what you just mentioned.
Nate B. Jones: That's right. I think that that is a piece that I really talk about a lot is that if you're not fostering that cultural piece, then you don't get what we talked about earlier. You don't get that bottoms-up experimentation, that idea that I can just foment an idea and come and talk to my manager or talk to my leader and say, hey, I have this. I think it works. I'll be honest.
My favorite engineers in the world over the course of my career have been those kinds of people, and not just engineers, others as well. They've come and they've said, I was just messing with this. I had a shower thought and I'm working on this idea. It's not done yet. What do you think? I'm like, eight times out of 10, it's like, oh my gosh, this is a really good idea. I don't quite know where it fits yet, but we're going to work it out together.
Michael Krigsman: On LinkedIn, Renee M. Gagnon points out that adoption can be supported if you have what she calls agentic users. I think those are those folks who will gravitate to the new technology and be early adopters and supportive.
Nate B. Jones: That is absolutely true as long as those people are really rewarded for spending time teaching as well as doing. Because when they're doing, they're going to be incredibly productive. And so when you have people who really understand how agents work and they've figured out these are the problems in my space that are agent susceptible and I can go after them and get a phenomenal amount done. And suddenly, everyone's like, oh my gosh, how is this person getting so much done? This is amazing. We love this.
And we celebrate the doing, rightly so. But we want to also celebrate the teaching because immediately what you need then is for them who are on the front lines typically to go to others in their space and say, Let's just sit down and chat together. Let's look over my shoulder. Let's just talk really candidly about our spaces. This is how I solve these problems and start to socialize that out peer-to-peer. If that happens, then you're really in business.
Michael Krigsman: This is a traditional technology innovation adoption problem to help avoid the anti-innovation antibodies that exist in so many organizations from attacking the new thing and saying, no, no, no.
Nate B. Jones: It definitely is. And I think that's one of the great ironies of this moment is that LLMs are absolutely a generational technology. They're transformative in a way deterministic software isn't. We can all see it in the organizations that have figured it out. But the way you get there is still sort of very human-denominated. And so there's some aspects of very traditional leadership culture that we have to talk about to make that jump. Yeah.
Open weights, costs, and team fluency
Michael Krigsman: This is from Kat Duffy. Given the current debate between open-weighted versus closed models and the lack of predictability recompute costs as the frontier companies shift pricing models, how are you thinking about the pros and cons for SMEs in working with open/on-prem builds versus, and she says, an admittedly more user-friendly enterprise strategy.
Nate B. Jones: I think about the larger trend lines that we can depend on and bet on when we make these kinds of decisions. I'm just going to name them, and then we can get into the dynamics of the decision. Number one, I think that we need to start to think about a frame of cost per completed action. That is a business-specific frame. You're going to find actions that you need completed in your business that are different from mine, et cetera.
You have to understand what is the dollar cost I'm paying in intelligence for that completed action? And then you're in a position where you can go back to both open weights models and frontier models and figure out what is the most economical choice. And we will get into the UX because that's a great part of the question. But I think it starts there.
And what's surprising about that is that people often assume if it's open source and open weights, it's just always going to be cheaper to go with the open source approach. And it depends. It depends on your willingness to invest in the upfront capital of a technical stack to serve those models. It depends on your willingness to maintain that stack over time to update the models as they come along. And it depends on if you're finding a Hugging Face sourced or other sourced open weights model.
It depends on the cost per token that you get there, and not just per token, but per token per task. And so one of the things people don't realize is that models will take dramatically different token lengths to solve the same problem. And so Kimi K3 came out and takes many more tokens, something like 50% or more, to solve business tasks than Fable does. Now, Fable's a very expensive model, but Fable's much more token efficient. And so you have to get into those level of details to understand the dynamics.
And then when you bring it to the team, and this is where the SME piece comes in, it's going to be a function.
Michael Krigsman: Yeah.
Nate B. Jones: of your team's willingness to be technically fluent. And so if you have a situation where you're like, I can afford the stack, my team is technically fluent, and I feel really good about maintaining it, absolutely, an open weights model is going to give you a more economical alternative to traditional frontier.
But if your team is not as technically fluent, if you don't want to invest in bringing the tech to them in a really user-friendly manner, and if you don't want to invest in the technical stack and the GPUs and everything you're going to have to have to run it, then it may effectively be cheaper to use a lab because the lab is just going to be there. The lab is going to effectively be bearing the cost of the UX for you and bearing the cost of serving for you.
And that's what's wrapped up in the price. And you're just going to decide that that is what you can do to leverage AI without formidable costs on SME. And SMEs are famously undercapitalized. And so that's why I go into that level of detail, because when I talk with leaders there, that's what they're thinking.
Michael Krigsman: We recently had Aaron Levie, the CEO of Box, as a guest on CXOTalk, and he made the comment that he can predict AI success based on your technology stack.
Nate B. Jones: Yep.
Michael Krigsman: Not too much different from what you were just saying.
Nate B. Jones: And I think that part of why is that technical stacks are proxies for the talent in your technical teams.
And so if your technical team is able to have a really intelligent conversation with you about a monorepo versus a services-based approach, or they're able to have a conversation with you about which GPU they chose and why, or they're able to have a conversation with you about how they think about MCPs versus API availability and agent versus human access, you're in a great spot no matter what you choose because the team is fluent.
Whereas if the team is like, oh, you know, we have this on Oracle, we've had it on Oracle for a really long time, and good luck. Your team is not in a position to get there, and that is reflected in the technical stack.
Michael Krigsman: I will also mention that we recently had, as a guest, the chief technology officer of Mozilla. If you're interested in open-source models and this discussion in-depth, go to CXOTalk.com and search for Mozilla. This would be an excellent time to subscribe to the CXOTalk newsletter so we can notify you about upcoming shows, and you can participate and ask more and more questions because we love your questions. All right. Here's a question from Swami Vaidyanathan on LinkedIn.
He says, enterprises would initially need to redesign their operating models as part of AI transformation, but do you see a future where enterprises get to an AI steady state where model releases are absorbed more steadily than being treated as a fundamental shift in the way they operate?
When new model releases matter
Nate B. Jones: I think we're already getting there with some companies, actually. I've seen that, where if you have a good harness and the harness is relatively thick and the LLM is relatively thin from a conceptual perspective, I don't mean thin as in less capable. I mean it's less shaping of the overall business impact, then you're in a position where it's your harness that is determining the outcome to the business overall. And the LLM is just the utility intelligence inside it.
And then at that point, what's transformational becomes choosing to adjust your harness to enable an LLM to do more. And this is exactly what we see with Frontier Labs when they say we are adjusting our harnesses as LLMs have the ability to work for longer periods of time and to do longer running agentic tasks. They choose to adjust the harness to enable the LLM to do more. But it's not that they are fundamentally changing how AI transformation works.
They have an understanding of the trajectory of the model and how it's growing over time and how new versions are affecting things. And they're able to say, generally speaking, the models are getting smarter, they're getting better at longer running goals. We are anticipating that and we're just adjusting the harness as a result. And smart organizations are already doing that. So I actually,
Michael Krigsman: that's something that has been different in 2026. Well, there's no doubt that tokenomics, the cost of these tokens, and the need to have an efficient, cost-effective token strategy is driving the design of the harness of prompts of your software so that you can easily switch between models. And the models are changing all the time.
Nate B. Jones: They're changing all the time, and that part's not going to stop. One of the things that I've just gotten used to is that the time between model releases is continuing to get shorter as you get players who are entering the race for different applications, as you get acceleration from the frontier labs. And you just learn to say, okay, a new model is out. Any given new model may not change how I do business. I just need to understand what outcomes I'm driving with AI.
Michael Krigsman: I need to understand if there's a particular model that's coming out that has attributes that are relevant to me, and then I can make changes. It may not change how you do business, but it may change or serve as a valuable input into some of the decisions that you're making because some of these new models, I mean, for example, Fable at various times displays a level of what, if it were a person, one would call insight. That's extraordinary.
Nate B. Jones: That's right. That's a place where it's like, it's not that I'm saying ignore the new model releases. That I'm saying have a mental map where you understand whether a new model release is relevant. And I think that what we're getting at is two different pieces that are often confused, so it's worth separating them.
There are utility models that we would use for business processes that are not likely to change a lot day to day, and then there are new model releases that represent a jump in the frontier intelligence capability that we all have access to, and those are models where you're going to have emergent properties like you're talking about treating it as a colleague, treating it as a partner, being genuinely insightful.
And you're going to want to be in a position where you have a fingertippy feel in your business context of what those frontier models are capable of. So you can think differently, think bigger. Imagine what you can do with intelligence that is like that in ways that you couldn't do a month ago. And that's the part of the new model race that's really energizing for me. Because I get to try Fable and I'm like, oh my gosh!
I'm getting results here that I haven't been able to get from other models before. What could I do differently as a result? That's really fun.
Michael Krigsman: One technique I've been using lately is I will go through a prompting process with Fable. It comes up with output. I then take that output and modify it in my own words. Give it back, and now Fable learns from the difference. And so, the expertise is in that diff, and Fable can incorporate it. It's pretty amazing.
Nate B. Jones: It's really remarkable. You know, another one that I found that's really fun is you have Fable look back at past work that you've done, and you have Fable understand how your own work has evolved in partnership with AI. And you will find most people's computers have this now, a whole litter or trail of documents, Excel files, PowerPoints, and other things you've built with AI. You can talk to Fable and say, look through this past. Look through this history.
How have my own working patterns with AI changed, and what can I learn about working more effectively with AI? Fable is smart enough to do that.
Job fear, AI costs, and accountability
Michael Krigsman: This is from Jessica Baker who says, what is the best way to drive AI adoption and innovation in a company where the knowledge workers are still frightened that AI is coming for their job?
Nate B. Jones: I tend to have a really honest conversation with the C-suite about that fear because you're right. To quote Claude, you're absolutely right. The knowledge workers I tend to talk to, that is the number one fear they have, particularly in the United States. And you have to address that as an elephant in the room if you want to have a conversation with frontline teams that's productive.
If you are not able to sit there and say, honestly, this is why we're adopting AI, we're not adopting AI to take your job away, we're adopting AI because of the leverage that we get as a business, because we want you to be more productive, because you need it for your careers long-term, all of which are true. Then frankly, your team is not going to be incentivized to work and to work well.
And I've had to have those conversations because there are some cases where I've talked to leaders and they're, well, you know, I do want to cut staff. I want this to be a labor-saving efficiency. And then I'm like, look, if you want to do that, you're not going to get a lot of buy-in from the team. Teams know how to smell that kind of fear, and they will sniff it out. They will be as resistant as they possibly can.
And it's going to be a real hard road for you because frontline teams have a lot of power in this. They can choose not to share what's in their head. They can choose to keep doing their old process secretly. They can choose to sabotage the AI effort in dozens of small ways that are very hard to notice. And they do if they feel like that fear is something that's real.
Michael Krigsman: This is from Anna Tatar. She says, how would you recommend to incorporate increasing cost of tokens to assess the true efficiency of an AI tool?
Nate B. Jones: We are told cost per token a lot. I hear it a lot. Of course, the labs talk about it a lot. But we need to think about cost per task because that's a much more actionable measure of the value of AI. Because if you look at cost per task, it's not necessarily getting more expensive.
In fact, in many cases, it's getting much, much cheaper over time because the frontier keeps moving forward and dumber models come behind and are able to do things that previously required a frontier model and frontier model costs. And so look at it as I have an expanding universe of tasks that are now available to AI. There are some that are going to be extra hard that are at the frontier today. This is the cost per task for that. It's going to be more expensive.
And then I have a widening array of what I call sort of utility tasks, things that are not particularly hard for AI to do anymore. Then it's about saying which model, which serving stack is most efficient to get that task done. It's going to, on the whole, be a cheaper curve over time.
Michael Krigsman: This is from Mohamed Chergui on LinkedIn. It's a real enterprise question. He says, many AI pilots fail not because of the models but because organizations struggle to make consistent decisions about ownership, governance, and adoption. Which organizational decision do you see as the biggest bottleneck to scaling enterprise AI?
Nate B. Jones: I go back to it as a people problem, and I think about accountability and ownership for AI-driven outcomes residing with specific members of the C-suite. Think about it as, if the CMO needs to be accountable to an AI-driven outcome, well, let them be accountable to that AI-driven outcome, but now they're the single throat to choke, and they're the ones that you can actually drive, and then everything else gets simpler.
Michael Krigsman: This is from Tancredi De Pretto who says, what are the ways we can avoid AI sycophancy influencing business decisions as people become more susceptible to automation bias?
Nate B. Jones: Either you're telling computers what to do or computers are telling you what to do. And I'm pretty blunt with leadership when I say this is something that is going to help you do your job, but it is still your job to push back. And I have seen it go both ways, Michael. I have seen leaders effectively subsumed delegate everything to AI, and it tends to have pretty immediate business consequences within the next three or four months. And I've also seen leaders who use it as a thinking board.
And so that's what I encourage. And I think you have to be honest when you initially talk about it.
Michael Krigsman: Yeah.
Nate B. Jones: Or else you do get into that pitfall.
Michael Krigsman: Let's talk again about tokenomics because it's so important. How can we manage tokenomics and manage the out-of-control costs? You've mentioned this, but it's so important for many of us.
Nate B. Jones: It's just budgeting. And I know that sounds unsexy, but if you're setting $5,000 as your budget, just notionally, and the frontier model is eating up a lot of that budget in a given month, just push people and push your stack toward cheaper models and open-source models and make people make do within that budget. You will be surprised at how much of the work still gets done. You can get 80% or 90% of that value without the frontier model these days, and you get a tremendous cost savings.
Michael Krigsman: R.A. Matthews says, models are trained on English. And so therefore, Kimi K3, the per-token task cost is way more expensive. At the core, meaning is what needs to be translated, which is intent.
Nate B. Jones: Kimi's cost structure is probably not a function of English per se. It is a function of how that model searches the solution space over time. I'm super familiar with the idea that some languages are represented in more expensive compute terms. So if you're using Tamil or if you're using Hindi, it may be more literal bytes per character in some cases. So that's true. But from an LLM utility perspective, you still get a relevant conversation about task efficiency regardless of what language you're operating in.
And really from a language perspective, what you need to think about is you need to think about the ability of the user in that language to have a complete, fast, clean experience. I tend to walk back into experience really fast when I talk about language because, if we don't, even internally, adoption stalls.
Michael Krigsman: All right. Then a related tokenomics question or cost question from LinkedIn from Nelson Almanzar who says, the cost per task. I think it becomes a question of value to the enterprise, not just cost to the enterprise at that point.
Nate B. Jones: If it costs you $100 an hour, but that $100 is going to give you $500 in return because maybe it's a frontier lab model that's using it and they get extraordinary value out of it, then you're going to pay that cost all day.
Michael Krigsman: Yeah.
Nate B. Jones: Whereas, if you are getting marginal return for that additional cost and you can go down to the $10-an-hour task and use your intelligence there and you get double the return on ROI, well, you're going to do that. You kind of have to do that math and not just look at it from a cost perspective. This is from Jadranka Berger who says, when your ops and data are tightened, is the model you're using.
When your ops and data are tighter, it's true that you have more option to change the model out, which gets right back to tokenomics and what we've been talking about in this hour, where you can trade it down to a cheaper model in many cases and get equivalent performance. That's a kind of leverage. But you also have the option, as I talked about, to loosen up your harness a little bit with a frontier model.
Michael Krigsman: Right.
Nate B. Jones: You get a different kind of leverage. You may have the option to do longer-running tasks or more tasks in one shot than you did before. You have to weigh the return on investment there.
Agent owners, evals, and production gates
Michael Krigsman: You have said that every agent needs one named owner.
Nate B. Jones: Yes.
Michael Krigsman: Who is that person? The developer, the engineering manager, the CIO, the business requester, the risk manager? Who should own the agent in the enterprise?
Nate B. Jones: There's sometimes a misperception of that statement that it's like, okay, so one person in the enterprise needs to own all the agents. I actually mean that if you have an agent and you don't want a tragedy of the commons effect where there's this little ghost agent running around and no one's taking care of it, no one's working on data and maintaining it, then you have to think about for any given agent you launch, who's the owner?
And so I actually think about the agent River and Shopify and Tobi, has an agent named River. The agent is in a Slack channel and everybody at Shopify can talk to that agent. But Tobi is the person responsible for that agent. Now, that's certainly not the only agent at Shopify. There's lots of other agents. Tobi is not responsible for all of the agents. Tobi is responsible for River. And so I think about it as a culture of ownership where you now have effectively digital employees.
And if you are going to have an agent on your team, you should know that either you, as the manager, or some employee that is working for you, they're the DRI. They're the person responsible, and that's the culture we want.
Michael Krigsman: This is from Arturo F. Munoz on LinkedIn who says, is it possible ever to expect to constrain a probabilistic AI so effectively that it will never drift and will behave as a deterministic system exclusively. And, rather than put this into the realm of what today is science fiction, let me ask, what can we do to help constrain hallucinations and that probabilistic diffusion that happens that can infect our results?
Nate B. Jones: I find that that is a question that boards will ask sometimes, that non-technical people will ask sometimes, and that when you talk to engineers who are familiar with today's models, you talk to CTOs familiar with today's models, it never comes up. And the reason it never comes up is that the hallucinations problem in actual production systems today is largely not an issue. And it's largely not an issue for a variety of reasons.
One, the labs care about it, and so they've been reinforcement learning like crazy, and they've been really working on validation. If you have been annoyed by 5.6 SOL talking to you a lot about checking its work, well, that's what they're doing to deal with hallucinations. That's part of the result of that work. But the other reason is quite simply that harnesses help you address a lot of that, and evals help you address a lot of that. And evals are just a part of the harness. You're checking the work.
And if you are implementing evals effectively, if you're implementing your tool calling so you call validated data effectively, you are in a position as a CTO or as an engineering manager building these systems where you have largely zeroed out that problem space.
Michael Krigsman: Nate, going back to pilots, how many experiments and failed pilots are necessary before an AI in production at scale can become a realistic outcome? In other words, how much do we have to experiment with failure in order to achieve. I think that you shoot straight for production.
Nate B. Jones: This is what I mean about picking a high-leverage task, and that you make sure that the gates along the way are really high so that when you get to production, you know that you've tested it. Then it becomes a question of effectively jumping through those gates and saying, okay, we have a gate around user experience. It has to be great. This is how we're testing it. This is how we know. We have a gate around serving and inference quality. This is how we know. Costs.
This is how we know this works. And so by setting your gates up high, you actually are just focused on reaching production from the get-go. And you're making sure along the way that you set up requirements that you may wash against like a rock a lot. Like you may try and try and try, and there are n failures along the way, but you know the bar is right, and you know you're headed toward a bigger goal. And then it's just the ability of your technical team to jump through those gates.
That gets you there. Some teams jump way, way faster because they're more fluent in AI and some jump slower, but they're going in the same direction regardless.
Advice for CIOs and when to stop
Michael Krigsman: All right. Speaking of jumping through hoops, what advice do you have for chief information officers when it comes to all of this set of issues?
Nate B. Jones: I tell chief information officers and I sort of sit with them and I'm like, you have the hardest job in the world right now. I have a lot of empathy for you. It's a really tough spot to be in. I think number one, you cannot keep AI out of the business, and so don't set out to try and over guardrail and think that you can control exactly where AI is in your business because people bring their own AI to work whether you like it or not.
And shadow AI is a massive issue at enterprise level, and CIOs worry about it all the time. And so I say you have got to find ways to say yes to your teams on AI as quickly and easily as you possibly can to minimize the risk of shadow AI in the business.
Then, on the other side, you have to have a real investment and a real conversation on cyber defense because one of the things we've seen, especially in the last couple of weeks, is that you have a tremendous scale-up in attack capabilities for these models. You have to assume that your enterprise could both be under attack itself and be leveraged to attack others if you're not careful.
Michael Krigsman: At what point do you stop a pilot and just acknowledge the approach, the process being automated, the models being used, or something else will never work? When do you throw in the towel on a pilot?
Nate B. Jones: I would throw in the towel on the pilot when, and this is what I find actually happens, when the goal changes and you find that whatever you were setting out to do is not as relevant or worthwhile or rewarding as you thought. That often happens either when business circumstances change or when you understand more about AI during the course of the pilot and you're like, actually, I picked the wrong goal. That happens a surprising amount of the time. In that situation, absolutely, you want to cut bait.
You want to walk away and you want to say, I have a better goal now because I learned.
Michael Krigsman: The learning is really that key part because if you've learned, as you said earlier, then you gain something out of it. If it's just,
Nate B. Jones: It wasn't wasted.
Michael Krigsman: Hopefully.
Nate B. Jones: It was worthwhile. Yeah, absolutely.
Michael Krigsman: Well, Nate, we're out of time. This has been amazing, and thank you so much for coming back to CXOTalk again, Nate.
Nate B. Jones: I had so much fun. The audience asked such great questions. You all are wonderful in terms of how you think about the business. It's been a great conversation.
Michael Krigsman: It has been awesome. You guys who asked questions, thank you so much. It's been like a torrent of questions, and if I missed anybody's questions, I apologize. Now, before you go, please subscribe to the CXOTalk newsletter and join us again. We really have amazing, amazing shows that are coming up. We have shows scheduled now into September. We're booking shows in October, so join us. And, connect with us on LinkedIn, with Nate and with me, and we'll see you again next time, everybody, and I hope you have a great day! Thank you!

