Multimodal AI
How Voice and Vision AI Can Understand the Complete Customer Journey in a Store
Voice and vision AI understand a store visit by each covering the stages the other cannot. A camera sees a customer arrive, browse, wait and leave, and never hears a word. A consented recording on the advisor's phone hears the need, the objection and the offer, and starts only when the advisor taps. Multimodal customer intelligence is the discipline of laying the two on one timeline, so the four silent minutes before the conversation and the two spoken minutes inside it are read as one journey.
The journey as the customer lives it
Most store visits in assisted retail follow the same shape. The customer enters, often with someone. They orient, then move to a display or counter. They wait, or they are approached. A conversation happens, or it does not. A decision follows, and the customer either bills, leaves with a promise to return, or simply leaves. Each stage leaves a different kind of trace, and only some of those traces are audible.
In plain terms: the conversation is the record of the interaction, and the camera is the record of what surrounded it. Voice carries intent, objection, offer and outcome. Vision carries footfall, wait, group, path and the walk-outs that nobody spoke to. The join happens on the advisor's tap to record, by time and zone, and never by who the person is.
Which signal sees which stage
| Stage | Vision sees | Voice hears | Only both can answer |
|---|---|---|---|
| Arrival | Entry count, time of day, group size | Nothing yet | Do groups convert differently from solo visitors after the same conversation? |
| Browsing | Zone dwell, display hotspots, path | Nothing yet | Does time at display A before engagement change the objection raised? |
| Waiting | Minutes unattended before an advisor arrives | Nothing yet | Does a wait over four minutes raise price objections? |
| Engagement | That two people are together at the desk | Need, model asked for, offer quoted, objection, rival named | Was the advisor talking to the decider? |
| Decision | Movement to billing, or to the exit | Yes, no, or a number and a date | Did the conversation end well and the queue undo it? |
| Exit | Walk-out, attended or unattended | Follow-up brief if a number was shared | How many walk-outs were never spoken to at all? |
One visit, told both ways
A furniture store in Pune, Saturday, 4.10 pm. The camera pipeline records: two people enter together; they spend six minutes in the sofa zone; nobody approaches for the first four; at 4.16 an advisor arrives; the group moves to the desk at 4.21; they leave at 4.29 without passing billing.
The consented conversation, started when the advisor tapped at 4.16, records: the customer asks about a three-seater in fabric; the advisor quotes the festive scheme correctly; the second person asks about delivery to a fourth-floor flat without a lift; the advisor says it will be checked; the customer says they will come back Sunday and shares a number.
Read separately, the camera reports an unconverted group visit with a four-minute wait, and the conversation reports a delivery objection with a lead captured. Read together, the visit says something more useful: the wait was long, the objection was logistical not price, the decider was the second person, and the store has until Sunday to confirm fourth-floor delivery. That is the record the store manager needs, and neither system alone produces it.
What multimodal customer intelligence adds by stage
- Arrival and group: conversion by group size, so the playbook can add a step for addressing the second person.
- Browsing: which display the customer stood at before the conversation, so the advisor's opening can start from what they were looking at.
- Waiting: wait time as a variable next to adherence and outcome, so staffing decisions use evidence from the floor.
- Engagement: the conversation scored as today, with the wait and the group as context columns.
- Exit: unattended walk-outs counted as a share of entries, the one number the voice half can never produce, because nobody tapped.
Confidence and failure modes
A join is a hypothesis with a score. The strongest case is one visual track in the desk zone at the moment of the tap. Real floors are messier. Two shoppers cross in front of a camera and one track splits into two, so a six-minute dwell becomes a two and a four. Two advisors stand near one group at a busy counter, and it is not certain whose tap belongs to which conversation. A customer walks out of frame to the parking area and back, and arrives as a new entry.
The right response is not to hide this. Each joined visit carries a confidence, weak joins are reported separately, and the store-week aggregates lean on strong joins. A retailer reading a multimodal report should expect to see how many visits were joined with confidence, how many were vision-only because nobody tapped, and how many were voice-only because the camera did not cover the zone.
Privacy at each stage
The camera half runs under signage, processes video on an edge device in the store, and emits events (an entry, a dwell, a walk-out), not frames or faces. No facial recognition, no re-identification across visits. The voice half starts on the advisor's tap after the customer agrees, and audio is deleted after transcription. The join uses time and zone, so no identity is needed to connect the two. Under India's DPDP Act, that is the difference between processing anonymous events and processing personal data, and the design keeps the second to the conversation alone.
Borentis today runs the voice half, live as Borentis Floor. ShopperDNA, which would add the camera events and the join described here, is on the roadmap.
Frequently asked questions
How does voice and vision AI know which shopper the conversation belongs to?
It does not identify the shopper. It matches the advisor's tap, with its time and zone, to visual tracks in that zone at that time, and scores the match. When the match is ambiguous the record says so.
Can vision AI hear what customers say?
No, and it should not try. CCTV audio is unconsented. The conversation comes from the advisor's phone after the customer agrees, which is a separate, consented signal.
What if the advisor never taps?
Then the visit is vision-only: counted, timed and, if unattended, flagged as a walk-out. Coverage, the share of walk-ins with a consented conversation, is reported next to every number for this reason.
Is multimodal customer intelligence available in India today?
The voice half is. Borentis Floor captures and scores consented conversations in Hindi, English and Hinglish. The vision-plus-voice product, ShopperDNA, is on the roadmap.
Related reading
- What is multimodal AI for retail?
- What is ShopperDNA? (on the roadmap)
- What is shopper intelligence?
- Is recording customers legal in India?
Where Borentis applies this
- Walk-in Recovery: The customer who left is still yours.
- Objection Intelligence: The reason they did not buy, in their own words.
- Playbook Adherence: Your playbook, finally observed.
Borentis is the Agentic Operating System for Customer Interactions, built for Indian retail floors: consented one-tap capture on the advisor's phone, every conversation scored against your playbook with the evidence behind every number, leads created when a number is heard, and coaching from your own best conversations.