Returns and refunds are usually the intent category teams are most nervous about automating, and for good reason — it’s the category where a mistake costs real money and where customers have the strongest incentive to push an agent past its intended limits. It’s also one of the highest-volume, most template-shaped categories, which makes getting the guardrails right worth the extra care.

Step 1: Separate Eligibility Logic From Conversation Logic

The most durable architecture keeps policy — what’s eligible, within what window, up to what amount — in a rules engine the agent calls, not embedded in a prompt the agent might be argued out of mid-conversation. A customer pushing back with “but the website said 30 days” should trigger a lookup against the actual policy record, not a negotiation with the language model.

Step 2: Design for the Pressure-Testing Conversation

Customers who want an exception will try multiple framings in one conversation — an appeal to fairness, then a threat to leave a bad review, then a claim their situation is unique. A well-guardrailed agent responds identically to policy regardless of framing, because the eligibility check runs once against verifiable data, not once per argument the customer raises.

Step 3: Route Genuine Edge Cases to a Human With Full Context

Genuine edge cases — a product that arrived damaged outside the stated window, a documented service failure — deserve a fast human review, not an automated denial. The guardrail isn’t “never make an exception,” it’s “the agent doesn’t make exceptions; a human with authority does, and only after the agent has gathered the relevant facts.” Build this handoff so the human reviewer receives a clean summary — order date, stated reason, what policy check failed and why — rather than a raw transcript to re-read from scratch.

Step 4: Invest in How Denials Are Communicated

When a request is genuinely ineligible, how the agent communicates that denial affects the outcome almost as much as the decision itself. A denial that cites the specific policy clause and explains the reasoning plainly is received very differently from a vague “unfortunately we can’t process this,” even when the underlying decision is identical — customers are considerably more likely to accept a denial they understand than one that feels arbitrary, even if they’re disappointed by it either way.

Test the Pressure-Testing Conversation Explicitly

It’s worth building the multi-framing pressure conversation described in Step 2 into your simulation suite as its own named scenario category, rather than hoping it shows up naturally in a broader test set. A realistic test run — fairness appeal, then a review threat, then a claim of a unique circumstance, all in one thread — is a strong, repeatable way to confirm the eligibility check genuinely runs once and holds, rather than eroding one argument at a time the way an unguarded language model tends to.

Why This Architecture Also Makes Policy Changes Safer

This separation also makes policy changes much safer to ship: updating a rule in the rules engine takes effect immediately and consistently across every conversation, without needing to verify that a prompt change didn’t subtly alter how the agent reasons about unrelated cases elsewhere in the flow. A team that embeds policy directly in prompt text instead often finds that tightening one rule accidentally loosens another, simply because the two lived as loosely related instructions in the same block of text rather than as two independently testable rules.

A returns and refunds agent earns trust by being consistently, boringly correct — not by being persuadable. Build the policy check as infrastructure the agent consults, not a stance it argues for, and invest as much care in how denials are communicated as in the underlying eligibility logic itself.