Software engineering has had a mature discipline around safely shipping changes โ automated tests, staged rollouts, monitoring, fast rollback โ for a long time. Support AI teams, by contrast, have often shipped prompt and policy changes the way engineers shipped code in the 1990s: manual review, then straight to production, then wait and see. Thatโs changing, and the teams making the change fastest are seeing it directly in incident rates.
The Three Stages
| Stage | What it catches | Customer risk |
|---|---|---|
| Offline evals | Known regressions against curated scenarios | None |
| Shadow mode | Emergent behavior on real live traffic | None (response never shown to customer) |
| Canary release | Anything the first two missed | Limited to a small traffic percentage |
Offline Evals: Catch Regressions Before Theyโre Live
An offline evaluation suite โ a curated set of representative conversations with known-correct outcomes, run against every proposed change before it ships โ is the equivalent of a unit test suite. Building a genuinely representative eval set is harder than it sounds, though โ it needs to include not just typical conversations but the specific edge cases and past incidents that have caused problems before.
Shadow Mode: See Real Traffic Without Real Risk
Shadow mode runs a proposed change against live incoming traffic in parallel with the current production agent, comparing outputs without ever showing the shadow agentโs response to the actual customer. This surfaces exactly how the change would have behaved on real, current traffic โ including edge cases no offline eval set anticipated โ with zero customer-facing risk.
Canary Releases: Limit the Blast Radius
Once a change clears both offline evals and shadow mode, rolling it out to a small percentage of live traffic first โ with the same monitoring used for the audit trail โ lets a team catch anything shadow mode missed while limiting how many real customers are affected before a rollback. Choosing the right canary percentage and duration matters: too small or too brief a window wonโt accumulate enough real interactions to catch a rare failure mode, while too large or too long a window unnecessarily extends the exposure window if something does go wrong.
Building the Monitoring That Makes Canaries Useful
A canary release is only as good as the monitoring watching it โ without clear, automated alerts on the metrics that matter (resolution rate, escalation rate, sentiment, cost per conversation), a canary can quietly underperform for hours before anyone notices, defeating the purpose of limiting exposure in the first place. This monitoring layer deserves the same design attention as the eval suite and shadow mode setup, rather than being treated as an afterthought once the more visible testing stages are in place.
None of This Is New โ Just Newly Applied
None of these three practices is novel in software engineering generally; whatโs changing is support teams finally adopting them for AI agent changes with the same rigor theyโd expect for any other production system. Teams that resisted this discipline, treating conversational AI as somehow exempt from the testing standards applied to the rest of their stack, are the ones most likely to show up in incident postmortems a year later.
Building an Eval Set That Earns Its Keep
A genuinely representative eval set is harder to build than it sounds โ it needs to include not just typical conversations but the specific edge cases and past incidents that have caused problems before, since a generic eval set drawn from average traffic will systematically underrepresent exactly the rare, high-stakes situations where a regression matters most. The practical habit worth building is adding every real incident to the eval set the moment itโs resolved, so the suite grows more representative of your actual failure modes over time rather than staying frozen at whatever it looked like on the day it was first assembled.