Meet Src: Your engineering org, instantly queryable

Meet Src: Your engineering org, instantly queryable

Stack Trace Podcast

Insights

A Newer Model Isn't Always Better: Lessons From Vanta's Engineering Org

A Newer Model Isn't Always Better: Lessons From Vanta's Engineering Org

Span Team

Iccha Sethi, Senior Vice President of Engineering at Vanta, on model upgrades, token spend, and how the engineering discipline is changing with AI. The full audio episode is available on Spotify.

Three Takeaways

  • A newer model is not automatically better for your product. Before you bank on the frontier model, you need your own evidence that it beats what you're already running for specific use cases.

  • Adding AI is a tax before it's a multiplier. More tokens don't produce proportionally more throughput, the same way adding more engineers never did.

  • Off-the-shelf tooling only gets you so far. The tools that move the needle are built around how a specific team already works.

Iccha Sethi leads engineering at Vanta, an agentic trust platform powering security for over 16,000 customers. She joined Stack Trace to dig into what it takes to build thoughtful AI products, including how Vanta approaches everything from evaluating the latest models to the internal practices that help its teams get more out of their agents.

How to Test and Evaluate the Latest Models

Engineers are often trained to treat the newest model as a strict upgrade, which makes the latest release tempting on arrival. Sethi says this is a common misconception, and that "a newer model is not always better than the older model." When the first reasoning models were released, they underperformed for some of Vanta's key use cases, which led the org to selectively swap back to older models.

Vanta now runs offline evals against what the team calls "golden datasets," which Sethi describes as the non-deterministic version of the tests engineers already know how to write. Every new model gets a row in a large internal table: latency, accuracy, similarity score, measured across use cases. Different product features require different models, so the question is never which model is best, only which is best for a specific goal or use case.

If a model passes those benchmarks during offline evals, Vanta rolls it out to a percentage of customers and gathers signal from online evals. In some cases the team also runs shadow mode, continuing to serve the old model while comparing the new one alongside it.

The same logic shapes how Vanta reads its own eval maturity. The team built an internal matrix that runs from capturing traces, to holding golden datasets, to running evals and experiments, to feeding results back into the product. The goal, Sethi says, "isn't to get to green everywhere, but to highlight where we want to get to green and where we're okay with the yellow."

Adding AI Is a Tax Before It's a Multiplier

Sethi is adamant that "input multiplier is not output." Headcount comes with a coordination tax and an onboarding tax, AI comes with a learning tax, and what teams get on the other side is a fractional multiplier. The key to working through those challenges was less about using the latest models and more about improving how effectively teams worked inside their agentic workflows and environments.

That approach also shapes how the engineering org looks at token spend. As companies work to rein in token costs after facing increasingly large bills, Sethi emphasizes that burning tokens to demonstrate effort isn't the same as producing results. Some of her most effective engineers aren’t the heaviest token users, for example, because they've worked out when a single file of context beats twenty tool calls.

Engineers and their managers are notified if they cross certain spend thresholds, and the conversation that follows usually focuses on which use cases and goals warrant the cost. Sethi's main goal is to build that awareness rather than clamp down on engineers, and to treat those conversations as learning opportunities.

Build Tools Fit for Your Team

One Vanta engineer was getting more code reviews than he could keep up with, and the team's existing review bot wasn't matching his needs. So he built his own, trained on his past reviews and on the reviews of engineers he considered particularly skilled at code review, and gave it a deliberately warm personality.

The bot has since become popular across the org. The engineer had already tried the general-purpose review tools and found they didn't cut it for him. What moved the needle was that this one was built specifically around how the team already thinks about and reviews code.

More generally, Vanta holds internal tooling to the same bar as any customer-facing product. The team tracks much of its AI usage on a Datadog dashboard: which skills get invoked most, thumbs up and thumbs down on review comments, evals running against the results. Internal efforts, in Sethi's view, require the same kind of discipline as shipping external products.

What This Means for Leaders

None of the individual pieces makes a team effective with AI on its own. What produced results at Vanta was the practice built around all of it: an eval loop that tests models against real workloads rather than public benchmarks, enablement that treats the learning curve as a cost worth paying, and tooling shaped around how a specific team already works. That layer is where effectiveness lives, and it has to be judged on output rather than on the inputs that are easier to count.

Vanta's engineering team has been publishing its own account of this work in a series called Trustcraft. Their post on their AI Quality Eval Maturity Model details the five-dimension framework Sethi described here along with additional benchmarks used by Vanta’s engineering teams. 

Ultimately, the maturity score was never the point, and shipping AI features customers can trust is.

Everything you need to unlock engineering excellence

Everything you need to unlock engineering excellence