BORENTIS

Multimodal AI

What Is Multimodal AI for Retail? Voice, Vision and the Complete Store Visit

Multimodal AI in retail is the practice of reading a store visit through more than one sensor at once, most usefully a microphone and a camera, so that what was said and what was done are analysed as one event. The conversation carries the intent, the objection, the offer and the outcome. The camera carries the arrival, the wait, the group and the path. Neither alone describes the visit. This guide explains what the term means on a real Indian shop floor, how the two signals are joined, and what the term does not mean.

A definition that fits a store

In research, multimodal AI usually means one model trained on text, images and audio together. On a retail floor the useful meaning is narrower and more practical: two separate pipelines, one for speech and one for video, each producing timestamped events, fused after the fact at the event level. The speech pipeline transcribes a consented conversation and scores it against the retailer's playbook. The video pipeline counts people, times waits and tracks movement between zones. A fusion layer lines the two up by time and place.

This matters because the honest version of multimodal retail intelligence is not a single giant model watching and listening. It is two well-understood systems and a join. The join is where the value and the difficulty both live.

What each modality measures

ModalityMeasuresCannot see
Voice (consented conversation)What the customer came for, the objection raised, the offer quoted, the rival named, the number shared, the outcomeHow long they browsed before an advisor arrived, whether they came alone, whether they left because nobody came
Vision (camera events)Entries and exits, group size, time in each zone, wait before engagement, unattended walk-outs, display hotspotsWhy they left, what they asked, what was promised, whether the finance step was explained
Multimodal (voice + vision, joined)The complete visit: the wait, the conversation and the exit as one record; conversion by path, by wait, by groupAnything outside the store, and any visit with no consented conversation is vision-only

The complete store visit

A visit has an inside and an outside. The inside is the interaction between the customer and the advisor. The outside is everything around it: the three minutes at the display before anyone approached, the second person who did the deciding, the queue at billing that turned a yes into a maybe. Conversation intelligence describes the inside with precision. Video analytics describes the outside with precision. Multimodal AI for retail is the decision to stop reading them separately.

The practical test of whether a retailer has multimodal intelligence is simple. Can it say, for one store, in one week, how many customers waited more than four minutes before an advisor engaged, and what share of those conversations then ended on a price objection? If the answer needs two reports and a guess, the retailer has two modalities, not one.

How the two signals are joined

  1. Session start. When the advisor taps to record on their phone, that tap is a timestamped event in a known store. It is the anchor.
  2. Zone. The store's floor is divided into named zones (entrance, display A, desk, billing). The advisor's zone at tap time comes from where the phone or the desk is; the camera pipeline tags visual events with zones too.
  3. Time window. Visual events in the same zone within a tolerance before and after the tap are candidates for the same visit.
  4. Confidence. Each join carries a score. One visual track in the zone at tap time is a strong join. Three tracks and two advisors is a weak one, and the record says so.
  5. Aggregation. Joined visits roll up by store, zone and week. Nobody looks at one shopper; the unit of analysis is the store-week.

What multimodal retail intelligence is not

  • It is not facial recognition. The camera pipeline produces events (an entry, a dwell, a walk-out), not faces or identities.
  • It is not a customer profile. No record follows a person across visits or stores. Identity exists only where a customer gave a number in the conversation, with consent, for follow-up.
  • It is not audio from the CCTV. Camera audio is unconsented and unusable. The conversation comes from the advisor's phone after the customer agrees.
  • It is not one model. Two pipelines, one join. Anyone selling a single model that watches and listens should be asked how consent works for the listening half.

Where this stands today

Borentis today captures consented conversations on the advisor's phone in Hindi, English and Hinglish, and scores them against the retailer's playbook with the transcript line behind every score. That is the voice half, and it is live as Borentis Floor. ShopperDNA, the product that joins camera events to those conversations, is on the roadmap. The design decisions above, events not identities, joined by time and zone on the advisor's tap, edge processing for video, are the decisions it is being built on.

For the argument that CCTV alone cannot explain conversion, and a worked visit told both ways, read the companion guides linked below.

Frequently asked questions

Is multimodal AI for retail the same as smart CCTV?

No. Smart CCTV is vision alone: counting, heatmaps, queue detection. Multimodal adds the consented conversation and joins the two. The camera never explains why a visit ended without a sale; the conversation does.

Do I need new cameras for multimodal retail intelligence?

Usually not. Most existing IP cameras produce a stream that an edge device can process for events. The constraint is coverage of the zones that matter, not camera quality.

Does multimodal AI need customer consent?

The camera half runs on signage and produces anonymous events, as CCTV already does. The conversation half needs consent before recording, evidenced per conversation. Both are notice-and-purpose questions under India's DPDP Act.

Can a small chain use this?

The voice half needs only the advisor's phone. The vision half needs cameras that cover the zones and an edge device per store. A ten-store chain can run both; the question is whether the visit-level questions are worth answering at that scale.

Related reading

Where Borentis applies this

Borentis is the Agentic Operating System for Customer Interactions, built for Indian retail floors: consented one-tap capture on the advisor's phone, every conversation scored against your playbook with the evidence behind every number, leads created when a number is heard, and coaching from your own best conversations.