How we deploy AI across a large retail store network.
No app to install, no training, no six-month project. What we learned putting AI to work across 160+ stores over WhatsApp.
I replicated a SaaS product in five days.
Five. And instead of being impressed with the AI, I was left with something a lot less fun: if building is this fast, then I spent years solving the wrong problem.
The problem was never getting the tool. It was getting anyone to use it.
Nobody can be everywhere
We run 160+ stores across 10+ core retail brands, each with its own implementation standards, with a lean team.
The arithmetic doesn't work. A supervisor can't walk their zone every week. And brand standards don't break all at once: they degrade slowly, one window at a time, until one day you walk into your own store and don't recognize it.
Alongside that, two more things that don't scale by hand. Maintenance issues travel through ownerless WhatsApp threads — whoever reported it doesn't know if anyone picked it up, and whoever could fix it finds out three days later. And hundreds of Google reviews waiting for a reply, spread across 196 listings, each needing a different brand's voice.
All three are the same problem. The operation produces more signal than we can read.
Before this, there was a vendor
This didn't start from nothing. Before IRIS, store operations ran on software from an outside vendor.
I don't tell it as a criticism of the product. I tell it because the replacement wasn't a feature decision, and that's the part people read wrong.
When building stopped being the bottleneck, the question stopped being which tool has more features. It became which tool gets opened. A portal loses that fight however good it is, because the problem was never the portal: it was asking a store manager with six urgent things on top of them to log into one.
In Chilean retail there's one answer, and it isn't sophisticated: WhatsApp. They're already there.
We chose the channel before the model. That was the architecture decision, and nobody on the technical team made it.
Photo, voice note, review
Three things, all through the same channel.
Visual validation. The store manager photographs the window. In under ten seconds the verdict comes back, against that brand's implementation standards — not a generic retail standard. Every brand has its own criteria and the system knows them.
Voice-note routing. They send an audio note, the way they send any audio note. The system transcribes it, classifies it as maintenance, checklist or commercial, and routes it to the right team with an SLA. What changes isn't the transcription: it's that the problem now has an owner and a clock.
Review agent. Replies in each brand's voice — 157 style profiles — across the group's 196 Google listings, and escalation when a negative review says something a person needs to see.
On top sits a health score per store: visual, zonal and commercial compliance in one number, with traffic-light alerts. Supervisors stopped asking "how's the zone doing?" and started asking "why did this store drop?"
How we measured whether it worked
This is where I nearly got it wrong.
The measure of success wasn't model accuracy. It was how many store managers used it without anyone reminding them.
A model at 95% accuracy the team abandons in March is worth less than one at 80% they're still using in December. Accuracy is a property of the model. Use is a property of the deployment, and it's the one that dies on you without warning.
Four rules that came out of it
IRIS wasn't the only one. Then came Cerebro on pricing, Draper on customer decisioning, Meridian on marketing mix, Delphine on buying forecasts, Octavius Flow on marketplaces. Nine systems, all with the same DNA.
- Channel before model. If the user has to learn a new interface, you've already lost. Delphine, our footwear forecasting model, reaches the buyer through a Telegram bot — because that's where the buyer already is.
- A person approves before anything reaches a customer. In Draper the AI proposes the next best action; someone approves it before the message goes out. It's what makes the commercial team trust it enough to let it run.
- Rules first, model on top. Octavius's pricing agent runs business rules and layers LLM scoring over them. The other way around, business judgment is the first thing to go, and you find out three weeks later.
- No holdouts, not ready. Draper logs every decision against a control group. A system that can't demonstrate its effect ends up defended by faith, and faith runs out in the first bad quarter.
An operator has to lead this
I've said it before and I'll repeat it because this case is the example: AI is our generation's biggest profitability lever if operators lead the adoption. Not consultants. Not evangelists.
The decision that made IRIS work — WhatsApp instead of an app — doesn't come out of a technical evaluation. It comes from having watched a store manager at seven on a Tuesday evening.
I'm not writing this from a solved position. We have systems moving slower than I promised and reviews that still hurt to read. But of everything we tried, this is what worked — and the part that worked wasn't the part I expected.
If you're looking at where to put AI in your operation, don't start with the use case. Start here:
What signal is your operation already producing, every day, that nobody has time to read?
For us it was three. Windows, voice notes and reviews. All three existed long before any model showed up. The only thing missing was someone to read them.
- The Retail Brief · 6 min · essay
Build vs Buy in the age of AI.
"Your website is terrible," a customer told me. They were right. And fixing it changed how we decide what software we buy.
- The Retail Brief · 7 min · essay
AI agents in retail: what actually works in production.
Nine systems, three different states, and why a portfolio where everything is 'in production' isn't a portfolio.