Key Points
- The sanctioned route has to be the easiest one. Snap focuses its engineers on 14 blessed agents so the same hard problem isn't solved a thousand times over, and every agent passes through a single central MCP gateway.
- Judgment cannot be taught to an agent. Snap constrains agents by personas, permissions, and risk tiers rather than trusting prompts to keep them safe.
- Model choice ranks below context and evaluation. Jain says model choices change every three weeks anyway.
Most engineering leaders describing AI adoption talk about what their agents build. Saral Jain spends more of his agent budget on stopping them. "At Snap, more than 90% of all code is written by AI agents right now," he told CXOTalk. Jain is head of engineering and CIO at Snap, and in host Michael Krigsman's introduction, "his team built those agents."
"AI is not about a bunch of disconnected agents," he said. "It is a managed production system." Snapchat "reaches almost a billion people every month," running on "thousands of microservices running across thousands of repos." At that scale, the system that catches what an agent gets wrong matters at least as much as the agent that wrote it.
The golden path, not a thousand experiments
Snap does have thousands of agents. Fourteen of them are the sanctioned production path. Snap has "created these 14 agents that cover almost 90% of what an engineer's day-to-day workflow looks like," Jain said, describing them as a "golden path" of blessed, managed agents covering the core loops: writing code, reviewing it, investigating crashes, reading A/B results.
"What we don't want to do is every employee do demos and prototypes and create a bunch of disconnected agents that do not solve the problem in a safe, secure, privacy-safe way," Jain said. The fix is to make the sanctioned route the path of least resistance, which he summarized as "the safe way should be the easy way."
Two of them recur throughout his account. "Our code reviewer, CodePal, reviews 90% of all of this code within the first 5 to 10 minutes, and it has found tens of thousands of bugs in new code that is being written." The other is Casper, a remote coding agent that listens to Slack and Jira threads: "Casper, which is our remote coding agent, is generating thousands of PRs every month."
Guardrails before building
The number he keeps returning to is a budget split. Jain said "more than 50% of our investment actually goes into the guardrails, not in terms of actually building stuff, but putting guardrails in place," listing the eval layer, the code review layer, and the rollback layer.
He argues the spending "actually does not slow us down. It actually helps us move faster because people have confidence in the system now." Snap's own history makes the point: "the first agent that became popular at Snap was CodePal. It is actually not an agent that helps you write code, it's an agent that helps you review code. It's a guardrail." His advice to other leaders: "invest first in the guardrails, not in the building."
Jain says judgment itself is not teachable. "You don't teach agent judgment, you just put the guardrails in place so that the agents just cannot do what you don't want it to do essentially."
Context beats model
Asked what an organization new to agents should focus on, Jain made his second principle "context beats model 100 out of 100 times." Which model is best, he noted, "changes every three weeks anyway."
Krigsman pressed him on what context actually means. Jain's answer was Snap's own history: "the information that has existed at Snap for the last 15, 16 years" in documents, codebases, and meeting notes. Older review bots saw only the diff; CodePal reads "not just the piece of code that is changing, but literally all of the surrounding code that is not changing as well." And the scoping cuts both ways: Snap does "not give our CodePal access to our Oracle financial data, because it does not need that data."
One consequence: Snap's code search, built for engineers, now sees agents with "60 times more traffic on those code search systems internally than humans."
Gameable inputs and the metrics Snap refused
Jain is blunt about metrics that flatter. Snap does not "obsess about activity metrics, things like number of PRs or number of tokens. I think these are all gameable input metrics." Snap deliberately skipped the token-maxing phase the industry has been talking about. Jain counts it among the best decisions they took that "we never created these vanity dashboards where people were projecting the amount of tokens they used."
"Our per-engineer commits are up by 75% year over year while reducing the number of severe outages by 57%." Shipping more while breaking less is, as he put it, "music to my ears."
Coordination shifts too. A change touching 20 services once meant convincing 20 sets of owners to reprioritize. An agent can now open all 20 pull requests with context attached. As Jain puts it, "if the cost of building is cheaper than the cost of having the meetings, then have fewer meetings."
The risk is unowned code
Pressed repeatedly by Krigsman on accountability, Jain located the risk somewhere other than authorship: "the risk is not AI-generated code or human-generated code. The risk is unowned code."
Ownership at Snap is tiered by blast radius. An internal dashboard with no sensitive data "can actually be 100% AI-generated with very little human reviews, maybe even auto-approved in some cases." A privacy-sensitive change reaching a billion users is not. There, humans "understand the spec, understand the code, and are able to vouch for the outcomes of this code."
Asked what worries him at 3:00 in the morning, he gave two answers in order. "AI agents can be wrong in not obvious ways. They can be wrong in very subtle ways, and they can be wrong very confidently." First, detection: "the biggest worry I have is ensuring that the guardrails are put in place such that we can detect when something is bad." Second, culture: "I worry that over a period of time it would be easy for people to ship changes they do not understand."
Asked for advice for CIOs, Jain answered, "Four pieces of advice very quickly": "start with a problem, not with an agent," focus on context rather than the model, invest in evals and guardrails, and start at the top. That last one he called "maybe the most important thing": "You cannot ask thousands of people to change the way they work without doing it yourself."
Watch the full conversation with Saral Jain and read the complete transcript on the episode page.
CXOTalk prepared this article with AI assistance from the verbatim transcript of episode 929. Quotations are unedited from the broadcast.