Lumina Beauty Group, a mid-size cosmetics and skincare retailer selling across four markets, came to us with a specific, expensive problem: nearly a third of its returns were shade-mismatch or “expected texture was different” complaints — issues that a human agent could usually resolve with a photo, but that were taking two to three back-and-forth messages to even get to that photo.

The Problem: Friction in the Ask

Lumina’s text-first chat flows asked customers to describe the problem before offering to attach an image, and most customers described rather than photographed, leading to guesswork on both sides. The brand’s own data showed conversations that included a photo in the first message resolved in one round-trip; conversations that started with text-only description took an average of 2.4 exchanges to reach the same resolution. Agents were effectively re-asking the same question — “could you send a photo?” — in a majority of return-adjacent conversations.

What Changed

Lumina rebuilt its opening flow for any conversation tagged with return-adjacent keywords to lead with an image upload prompt, paired with a visual AI model trained on the brand’s own product catalog images to compare the customer’s photo against reference shade and texture data. For genuine mismatches, the agent could immediately offer a replacement in a closer shade rather than a refund — preserving the sale. For cases where the photo showed correct application but a skin reaction, the agent escalated immediately to a human with the image already attached.

A general-purpose vision model without Lumina’s own reference data struggled to distinguish subtle shade differences that the brand’s purpose-trained model caught reliably — a case where a narrow model clearly outperformed a general one.

Handling Ambiguous Cases

Not every photo resolves cleanly — lighting conditions, camera quality, and genuinely borderline shade differences all produce cases where the model’s confidence is too low for an automatic decision. Rather than forcing a guess, the flow routes these to a human reviewer with the photo and the model’s best assessment attached, saving the reviewer the step of requesting the image themselves.

Results After One Quarter

Metric Before After
Avoidable returns (return-adjacent conversations) Baseline -55%
Average exchanges to resolution 2.4 1.3
Single-exchange resolution rate 31% 74%
Customer satisfaction on return conversations Baseline Up, driven by “didn’t feel like I had to argue my case”

For any product category where “does this actually look/feel wrong” is the real question, a visual-first flow beats a text-first one — not because AI vision is flashy, but because it removes an unnecessary translation step between the customer’s problem and the agent’s understanding of it.

Why a General Vision Model Wasn’t Enough

Lumina initially trialed a general-purpose vision model before commissioning its own training run, and the results were noticeably weaker on exactly the cases that mattered most: subtle shade differences between adjacent SKUs in the same product line, which a model without brand-specific reference imagery tended to score as near-identical even when a human eye could tell them apart under normal lighting. The lesson generalized beyond beauty — a narrow, purpose-trained model reliably outperforms a general one whenever the distinguishing detail is specific to a brand’s own catalog rather than a broad visual category any model would have seen widely during its own training.