What Sentiment Routing Actually Solves

Most support teams use sentiment analysis as a reporting tool: understand how customers feel, track CSAT trends, flag concerning conversations for review. That’s valuable, but it’s retrospective.

Real-time sentiment routing uses the same signals to make routing decisions during the conversation — before a customer has a chance to churn, post a negative review, or escalate to a manager.

This guide covers the full implementation: model selection, routing logic, integration patterns, and tuning considerations.

Architecture Overview

Incoming Message
      ↓
Sentiment Model (inference < 50ms)
      ↓
Sentiment Score + Intensity + Trend
      ↓
Routing Decision Engine
      ↓
  ├── Negative + High Intensity + Escalating → Senior Agent Queue
  ├── Negative + Moderate + Stable → Standard Queue (priority bump)
  ├── Neutral/Positive → AI Handling or Standard Queue
  └── Repeated Negative across sessions → Retention Queue
      ↓
Assignment + Context Package

The key architectural principle: sentiment is an input to the routing decision, not the decision itself. Route on sentiment in combination with conversation history, customer tier, issue category, and current queue depth.

Step 1: Model Selection

Options

Lexicon-based models (VADER, SentiWordNet):

  • Fast (< 1ms)
  • No training data needed
  • Poor on domain-specific language (“this is broken” → scored as ambiguous)
  • Good for: simple positive/negative/neutral triage

Fine-tuned transformer models (DistilBERT, RoBERTa fine-tuned on support data):

  • 10-50ms inference on CPU, < 5ms on GPU
  • Requires 5,000+ labeled examples
  • Handles domain-specific language well
  • Supports multi-class output (angry, frustrated, confused, satisfied, neutral)

LLM-as-classifier (GPT-4o-mini, Claude Haiku via API):

  • 100-300ms latency
  • No training data needed
  • High accuracy on nuanced text
  • Cost: ~$0.002 per conversation turn
  • Good for: low-volume or high-stakes routing

Our recommendation: fine-tuned DistilBERT for real-time routing (latency < 10ms, accurate on support language), with LLM-as-classifier as a fallback on low-confidence outputs.

Labels to Train For

Standard positive/negative/neutral is insufficient. Train for:

  • frustrated — repeated complaints, expressions of lost time
  • angry — profanity, explicit threats to cancel or escalate
  • confused — multiple questions, “I don’t understand” patterns
  • satisfied — expressions of thanks, resolution confirmation
  • urgent — time pressure language, “I need this today”, “this is blocking me”

Multi-label is correct for most messages (a customer can be both frustrated and urgent).

Step 2: Real-Time Inference Integration

Integration Pattern

For a webhook-based support platform:

import httpx
from fastapi import FastAPI, Request

app = FastAPI()
sentiment_client = httpx.AsyncClient(base_url="http://sentiment-service:8000")

@app.post("/webhook/message")
async def handle_message(request: Request):
    payload = await request.json()
    message_text = payload["message"]["body"]
    conversation_id = payload["conversation"]["id"]

    # Async sentiment inference — don't block the message receipt
    sentiment_task = asyncio.create_task(
        sentiment_client.post("/infer", json={"text": message_text})
    )

    # Enqueue message immediately
    await enqueue_message(payload)

    # Await sentiment result (target: < 50ms total)
    sentiment_response = await sentiment_task
    sentiment = sentiment_response.json()

    # Update routing with sentiment
    await update_routing_context(conversation_id, sentiment)
    await maybe_reroute(conversation_id, sentiment)

    return {"status": "accepted"}

Key design choice: enqueue the message immediately, don’t block on sentiment inference. Update routing context after. If sentiment inference fails, fall back to standard routing.

Inference Service (DistilBERT)

from transformers import pipeline
from fastapi import FastAPI

app = FastAPI()
classifier = pipeline(
    "text-classification",
    model="./models/support-sentiment-distilbert",
    return_all_scores=True,
    device=0  # GPU; use -1 for CPU
)

@app.post("/infer")
async def infer(body: dict):
    scores = classifier(body["text"], truncation=True, max_length=512)[0]
    labels = {s["label"]: round(s["score"], 3) for s in scores}
    primary = max(scores, key=lambda x: x["score"])
    return {
        "primary": primary["label"],
        "confidence": primary["score"],
        "scores": labels,
        "high_intensity": labels.get("angry", 0) > 0.6 or labels.get("frustrated", 0) > 0.7
    }

Step 3: Routing Decision Logic

async def maybe_reroute(conversation_id: str, sentiment: dict):
    conv = await get_conversation_context(conversation_id)
    current_queue = conv["queue"]
    customer_tier = conv["customer"]["tier"]
    sentiment_history = conv["sentiment_history"]  # list of last N sentiments

    # Escalation trigger: current message angry/frustrated AND intensity rising
    trend = calculate_trend(sentiment_history + [sentiment])
    if (
        sentiment["primary"] in ("angry", "frustrated")
        and sentiment["high_intensity"]
        and trend == "escalating"
        and current_queue != "senior"
    ):
        await reroute(conversation_id, "senior", reason="sentiment_escalation")
        await notify_agent(conv["assigned_agent_id"], "Conversation rerouted to senior queue — customer sentiment escalating")
        return

    # Priority bump: frustrated but not at escalation threshold
    if (
        sentiment["primary"] == "frustrated"
        and not sentiment["high_intensity"]
        and current_queue == "standard"
    ):
        await bump_priority(conversation_id, delta=2)
        return

    # Retention queue: Enterprise customer with repeated negative sessions
    if customer_tier == "enterprise":
        recent_sessions = await get_recent_session_sentiments(conv["customer"]["id"], days=30)
        negative_sessions = [s for s in recent_sessions if s["primary"] in ("angry", "frustrated")]
        if len(negative_sessions) >= 2 and current_queue != "retention":
            await reroute(conversation_id, "retention", reason="repeat_negative_sessions")
            await flag_for_csm_review(conv["customer"]["id"])

Step 4: Context Package for Receiving Agent

When a conversation is rerouted, the receiving agent needs context fast. Send a pre-computed summary:

{
  "routing_reason": "sentiment_escalation",
  "sentiment_summary": {
    "current": "angry (0.82 confidence)",
    "trend": "escalating over last 4 messages",
    "trigger_phrases": ["this is completely broken", "I've been waiting three days"]
  },
  "suggested_opener": "I can see this has been a frustrating experience. I'm taking over and I have full context — let's get this resolved right now.",
  "resolution_authority": ["refund_up_to_500", "extend_trial_7_days", "escalate_to_engineering"]
}

The resolution authority list is critical. A frustrated customer routed to a senior agent who can’t actually resolve their issue is worse than no routing at all.

Step 5: Tuning and Calibration

Expected Outcomes After 30 Days

Monitor:

  • Escalation rate by sentiment trigger (should be 5-15% of conversations)
  • False positive rate — angry-flagged conversations that resolved with standard handling (target: < 20%)
  • Rerouting CSAT lift — conversations that were rerouted should show higher CSAT than non-rerouted negative conversations

Common Tuning Issues

Too many reroutes: Lower the intensity threshold or require 2 consecutive negative messages before triggering. Rerouting every frustrated customer overloads senior queues.

Missing real escalations: Increase sensitivity for specific customer tiers (Enterprise customers warrant lower thresholds). Add keyword triggers for explicit escalation phrases (“I want to cancel”, “I’m calling my account manager”).

Latency too high: Profile your inference service. DistilBERT on GPU should be < 5ms. If using CPU, batch nearby inference calls. Consider caching sentiment for repeated identical phrases.

Real-time sentiment routing typically produces a 12-18% CSAT improvement on escalated conversations within the first 60 days. The mechanism: by the time a frustrated customer reaches a senior agent, they haven’t waited through a standard queue. The first thing they hear is “I know this has been difficult, I’m taking over.” That framing changes the conversation.