This month’s release addresses two of the most common requests we’ve heard since shipping the Agent Memory API: a safe way to test agent changes before they touch real customers, and a straightforward way to prove what the agent did and why, after the fact.
Agent Simulation Sandbox
Before now, testing a prompt or policy change meant either a manual QA pass on a handful of hand-picked scenarios, or shipping to a small live traffic percentage and watching for problems. The new Simulation Sandbox runs proposed changes against a library of thousands of historical, anonymized conversation replays — surfacing exactly which past interactions would now resolve differently, and flagging any that would now give an answer that contradicts current policy.
Teams can build custom scenario sets from their own edge cases and re-run them on every future change, which means a hard-won lesson from a past incident doesn’t have to be re-learned the next time someone touches the same part of the configuration. The sandbox also supports side-by-side comparison, showing the current production agent’s response and the proposed change’s response to the identical historical input.
Compliance Audit Trail
Every autonomous decision — a resolution, a refund approval, an escalation, a declined request — now generates a structured, exportable audit record:
- Input: the customer message and any extracted entities that triggered the decision
- Policy version: the exact rule set in effect at that moment
- Model version: which model and configuration handled the turn
- Reasoning summary: a plain-language trace of the decision path
- Outcome: the action taken and its result
Records are queryable by customer, date range, decision type, and policy version, and are immutable once written — a requirement for any account that may need to produce them as evidence in a dispute.
Also in This Release
Read receipts are now available across all messaging channel integrations, not just web chat, closing a gap that had been a recurring request from teams running high-volume WhatsApp and WeChat deployments. The reasoning engine shipped in April now supports streaming partial responses in the widget, reducing perceived latency on longer answers by showing the response as it’s generated rather than waiting for the full answer to complete.
The audit trail is queryable by customer, by date range, by decision type, and by policy version, so a compliance team investigating a specific complaint can pull the exact record in seconds rather than searching through weeks of transcripts. Records are retained according to each account’s configured retention policy and are immutable once written, which matters for any account that may need to produce them as evidence in a dispute.
Availability
The Simulation Sandbox is available now on Growth and Enterprise plans; the Compliance Audit Trail export is Enterprise-only and can be requested through your account team. We expect both to become foundational tools for any team running frequent agent configuration changes at scale.