AI agents in retail: what actually works in production.
Nine systems, three different states, and why a portfolio where everything is 'in production' isn't a portfolio.
Every time someone shows me their AI portfolio, I count how many cases are "in production."
If it's all of them, I stop listening.
Not because it's necessarily a lie. Because a portfolio where everything works equally well isn't telling you the state of anything, and the state is the only interesting part.
So I'll start with mine.
The three states
In production. IRIS runs the store network over WhatsApp. Cerebro verifies SAP prices against the storefront every day across three business units. Meridian runs eight Bayesian marketing mix models, one per brand, retrained weekly. Octavius Flow splits seven scheduled agents across four marketplaces in Chile and Mexico. Andrea runs for eight brands. Chispeza is in use at three chains.
Live pilot, which isn't the same thing. Draper. The data scope covers 17 banners, but the live operation today is Hoka: three stores and cl.hoka.com. That's a pilot. Calling it a rollout would be the first lie in a series.
Built, not deployed. Cifra consolidates five Excel sources into a 38-category cash-flow model, tested against more than 27,000 real records. It isn't anybody's system of record.
Nine systems, three states. The distance between the first and the third is where everything we learned lives.
The channel decides adoption
IRIS runs on WhatsApp. Delphine reaches the footwear buyer through a Telegram bot. Neither is an aesthetic decision.
I've told this before: I replicated a SaaS product in five days and the lesson came back the opposite of what I expected. The weak link isn't the tool's horsepower, it's how many people on the team can operate it. Since then the first question stopped being "which model?" and became "where does this person already work?"
An excellent agent in a channel nobody opens returns zero. Not a little: zero.
Business rules first, model on top
Octavius Flow's pricing agent runs every four hours. It applies business rules first; then it layers LLM scoring over that.
That order isn't negotiable. Rules are auditable and the model isn't. When a price comes out wrong — and it does — you want to be able to say "the rule says this" before you have to say "the model decided that."
The other way around, business judgment is the first thing to evaporate, and you notice three weeks later, once the margin is gone.
A person approves before anything reaches a customer
In Draper the AI decides the next best action per person — who to contact, what to offer, when, on which channel — and dispatches with human approval from its own operations console.
You can read that as a lack of ambition. It's the reverse. The bottleneck in a decisioning system isn't decision quality: it's how much the commercial team trusts it enough to let it run. Human approval buys that trust, and it can come out later. Starting without it is how you get the system switched off in the first bad week.
And there's a part that isn't optional. Chile's Ley 21.719 lives in Draper's data model from day one, not as a patch. Here, that isn't a design preference.
No holdouts, not ready
Every Draper decision is logged against a control group. Incremental revenue gets proven by measuring, not by declaring.
This is the one that costs most to skip, because the cost doesn't show up early. A system without holdouts works perfectly for two good quarters and dies in the first bad one, when somebody asks how much of this was incremental and the honest answer is "we don't know."
Meridian exists for the same reason, one level up. Without causal measurement, media budget gets split by history and instinct, and platform dashboards attribute sales to themselves. Eight independent models, one per brand, fitted on 137 weeks of real point-of-sale revenue from SAP. Not on platform-reported conversions.
Ranking well isn't deciding
Delphine predicts footwear sell-through by combining the product photo with its attributes and its price. Out-of-sample AUC is 0.723.
That's honestly good and honestly limited. It ranks risk well across a buying portfolio; it doesn't replace the buyer's judgment on any individual decision. That's why the verdict ships with a calibrated uncertainty band: the model also says how confident it is.
An agent that can't express its own uncertainty isn't ready to touch money.
And I publish the AUC on purpose. If the number embarrasses you, publishing it isn't the problem.
All nine are still alive
This is where I expected to write the list of the ones that didn't make it. It's the part that makes an essay like this credible.
I don't have one. All nine are still alive.
I could sell that as good aim. I think it's something else: when the channel is the one the team already uses, when rules come before the model, and when someone approves what goes out, the failure mode stops being "the system died" and becomes "the system is moving slower than I promised." Draper has been in pilot longer than I'd have liked. Cifra is built and waiting.
Neither fell over. Both are slow. I'll take that, but I'm not going to sell it as a win.
What we don't delegate
After nine systems, the list is short and it's firm:
- The message that goes out to a real customer, without a person approving it.
- The final price, without a business rule above it.
- The individual buying decision, which Delphine informs and the buyer makes.
- The reply to a negative review with a serious theme, which IRIS escalates instead of answering.
None of the four is a technical limitation. We chose all four.
If you're building your own portfolio, the question isn't what an agent can do. By now, nearly everything.
Whose effect can you prove?
The ones still alive in our operation are the ones that can answer it. The ones that can't get defended with enthusiasm, and enthusiasm has a fairly short shelf life in a management committee.
- The Retail Brief · 8 min · essay
How we deploy AI across a large retail store network.
No app to install, no training, no six-month project. What we learned putting AI to work across 160+ stores over WhatsApp.
- The Retail Brief · 6 min · essay
Build vs Buy in the age of AI.
"Your website is terrible," a customer told me. They were right. And fixing it changed how we decide what software we buy.