Multimodal AI
Why Retailers Need Conversation Intelligence and Computer Vision Together
Conversation intelligence and computer vision solve different problems, and retailers who buy one usually discover the other's problem within a quarter. The camera shows that 60 percent of Saturday visitors left without billing and cannot say why. The conversation tool explains every objection and cannot say how many people never got a conversation. This guide sets out the decisions that need retail video + conversation intelligence together, and what joining them actually requires.
Two half-answers most retailers already own
Almost every organised retailer in India has cameras, and many have footfall counters or a video analytics layer on top of them. A growing number have begun capturing consented sales conversations. Each investment was justified on its own, and each produces a report the other cannot check. The camera report says conversion fell. The conversation report says objections rose. Whether those are the same customers, at the same hour, at the same counter, is a question nobody in the building can answer.
The reason is structural. The conversation is the evidence of what happened between the customer and the advisor. The video is the evidence of what happened around them. Voice holds the intent, the objection, the offer and the outcome. Vision holds the footfall, the wait, the group, the path and the customers who walked out unattended. Joining them takes a common key, and the natural one is the advisor's tap to record: a time, a store and a zone. Not a face.
The questions each half answers alone, and together
| Question | Conversation alone | Computer vision alone | Joined |
|---|---|---|---|
| Why did this visit not convert? | Yes, in the customer's words | No | Yes, with the wait and group as context |
| How many visitors never got a conversation? | No | Partly: unattended walk-outs | Yes, as a share of entries, by hour |
| Is our capture rate honest? | No denominator | Entries only | Yes: consented conversations over entries |
| Does wait time change the objection? | No | No | Yes |
| Which display leads to which question? | No | Hotspots only | Yes |
| Was the advisor speaking to the decider? | Sometimes, from the transcript | Group size only | Yes, more often |
| What did the customer ask for that we lacked? | Yes | No | Yes |
Five decisions that need both
- Staffing at peak. The camera says waits climb after 5 pm. The conversation says price objections climb after 5 pm too. Joined, the store learns whether the second causes the first, and whether one more advisor on Saturday evening is worth more than a discount.
- Unattended walk-outs. Vision counts them. Only the conversation data, by comparison, shows what those customers would likely have asked and objected to, so the store can decide whether to chase coverage or fix the pitch.
- The coverage denominator. Every conversation score depends on how many walk-ins it represents. Entries from the camera make capture rate a real number instead of an estimate.
- Display to conversation. If customers who linger at the premium display then raise a price objection, the display is doing its job and the offer is not. If they linger and ask nothing, the display needs an advisor nearby.
- Group visits. A family visit in a two-wheeler showroom or a jewellery store is a different sale from a solo one. Conversion by group size, with the transcript showing who spoke, tells the trainer what to teach.
What the join needs, and does not need
- Needs: a zone map of the floor, cameras covering the zones where conversations happen, an edge device per store to turn video into events, clock sync between phones and cameras, and consented capture already running.
- Needs: a confidence score on every join, and reporting that separates strong joins, weak joins, vision-only visits and voice-only visits.
- Does not need: facial recognition, loyalty lookup, or any identity for the shopper. Time and zone are enough.
- Does not need: audio from the cameras. The conversation is captured on the advisor's phone after consent.
- Does not need: a single model that does everything. Two pipelines and a join are more honest, more auditable and easier to explain to a data protection officer.
The cost of running them apart
The visible cost is two dashboards and two vendor relationships. The real cost is decisions made on half-evidence. A regional manager reading the camera report sends a trainer to fix conversion at a store whose problem is one advisor for three counters. A trainer reading the conversation report coaches the finance step at a store whose customers leave before any step is reached. Both reports were correct. Neither was sufficient.
Where Borentis stands
Borentis today is the conversation half: consented capture on the advisor's phone in Hindi, English and Hinglish, scored against the retailer's playbook, live as Borentis Floor. ShopperDNA, which would add camera events from the store's existing cameras and join them to the conversation by time and zone, is on the roadmap. Retailers that want to prepare for it can start with two things now: a zone map of each store, and capture rate measured against a real entry count.
Frequently asked questions
Can I get conversation and computer vision from one vendor today?
Vendors that do both at the visit level in Indian assisted retail are rare. Borentis runs the conversation half today and has ShopperDNA, the joined product, on the roadmap. Footfall vendors offer vision alone.
Does joining conversations to video mean identifying customers?
No. The join uses the advisor's tap time, the store and the zone. It matches events, not people. Identity exists only where a customer gave a number in the conversation for follow-up.
Which should a retailer buy first?
If the question is why customers leave, the conversation. If the question is how many leave unattended, vision. Most chains have both questions in different stores, which is the case for the pair.
Related reading
- Voice vs video vs multimodal retail intelligence
- Why CCTV alone cannot explain retail conversion
- How to use CCTV for retail analytics
- ShopperDNA, on the roadmap
Where Borentis applies this
- Walk-in Recovery: The customer who left is still yours.
- Execution Scorecards: See the floor before the P&L does.
Borentis is the Agentic Operating System for Customer Interactions, built for Indian retail floors: consented one-tap capture on the advisor's phone, every conversation scored against your playbook with the evidence behind every number, leads created when a number is heard, and coaching from your own best conversations.