A pattern we see repeatedly in customer accounts: per-token pricing keeps falling, generation after generation, and yet the monthly AI spend line item keeps climbing. The falling unit cost isn’t the problem — unmanaged growth in both conversation volume and model-tier usage is.
Separate Volume Growth From Tier Drift
When a bill rises, the first question to answer is which of two very different things caused it: more conversations overall (a business-health signal, usually good news), or more conversations being routed to an expensive model tier than the intent actually requires (a routing inefficiency, entirely fixable). Without per-intent cost visibility, these two causes are easy to conflate.
Set a Budget Per Intent Category, Not Just a Global Cap
A global monthly spend cap creates a bad incentive: once approached, teams often respond by degrading service broadly rather than addressing the actual inefficiency. Budget per intent category instead:
| Intent category | Suggested cost ceiling | Basis |
|---|---|---|
| Order status, FAQ | 2-3x current median | Low complexity, high volume |
| Billing, basic troubleshooting | 2x current median | Moderate complexity |
| Multi-step technical, complaints | Case-by-case, monitored not capped | High variance is expected |
Revisit Routing Thresholds Every Model Generation
Every time a new model generation ships — cheaper, faster, sometimes both — the calculus for which intents deserve the expensive tier shifts. Building a recurring calendar reminder to re-evaluate routing thresholds against the current model landscape, rather than waiting for a cost anomaly to prompt the review, catches savings opportunities before they’ve cost months of unnecessary spend.
Watch for Cost Hiding in Retries and Fallback Loops
A less obvious source of cost drift is retry and fallback behavior — a conversation that fails an initial attempt and retries against the same or a higher model tier can silently multiply the effective cost of a single customer interaction. This is easy to miss in a per-conversation cost view that only shows the final successful attempt, which is why it’s worth specifically tracking retry rate and its associated cost as its own line item rather than assuming it’s negligible.
Cost Management Is Ongoing Operations, Not a One-Time Setup
AI cost management isn’t a one-time architecture decision; it’s an ongoing operational discipline that needs the same regular review cadence as any other significant line item in your budget. Teams that set up per-intent budgets once and never revisit them tend to drift back toward the exact global-cap blind spot this framework is meant to avoid, simply because the categories and their appropriate ceilings shift as your product, traffic mix, and the underlying model landscape all keep moving.