BORENTIS

Multimodal AI

How to Combine Customer Conversations With In-Store Behaviour, Step by Step

To combine customer conversations with in-store behaviour, a retailer needs two event streams that share a clock and a floor plan, one anchor that both can see, and a join that admits when it is unsure. The conversation stream comes from consented capture on the advisor's phone. The behaviour stream comes from camera events processed at the edge. The anchor is the advisor's tap to record. This guide sets out the steps in the order they should be done, and the places where audio + video analytics usually goes wrong.

Before you start

  • Consented conversation capture already running, with coverage measured. If advisors are not tapping, there is nothing to join.
  • A floor plan for each store with named zones: entrance, each display cluster, each counter or desk, billing, exit. Six to twelve zones is typical.
  • Cameras that cover the conversation zones, not only the entrance. Footfall cameras at the door are not enough; the join happens at the counter.
  • An edge device per store that turns video into events and keeps frames inside the building.
  • A store manager who agrees that the aim is the store-week, not any single visit or advisor.

The eight steps

  1. Map the zones. Draw them on the floor plan, then check that each camera's field of view is labelled with the zones it covers. A counter that two cameras see from opposite sides needs one zone name, not two.
  2. Sync the clocks. Phones, the edge device and the cameras must agree on time to within a second or two. Most join failures in early pilots are clock drift, not model error. Use network time on all three and check weekly.
  3. Anchor on the tap. When the advisor taps to record, the app logs the time, the store and the advisor's default zone (their counter). That tap is the only bridge between the two streams. The customer's consent is captured at the same moment.
  4. Define the visual events. From the edge device: entry, exit, zone enter, zone leave, group formed (two or more tracks moving together), engagement (a staff track joins a visitor track in a zone), and unattended walk-out (a visitor track exits with no engagement event).
  5. Join by window. For each tap, look for visitor tracks in the tap's zone within a window before and after it. A typical window is 60 seconds before and 30 after; tune it per store. The visit for that conversation is the track's full path from entry to exit.
  6. Score the join. One candidate track in the zone: strong. Two or three: weak, and record all candidates. No candidate track: voice-only visit. Tracks with no tap: vision-only visit, and if unattended, a walk-out.
  7. Aggregate by store-week. Wait before engagement, group size and path become context columns on each scored conversation. Unattended walk-outs sit next to capture rate. Nothing is reported at the individual level.
  8. Review the failure cases. Every week, read a sample of weak joins and unmatched taps with the store manager. Most have a floor explanation: a camera moved, a counter was relocated for a promotion, two advisors shared one desk.

The join keys

KeyFrom the conversationFrom the camera eventsTolerance and note
StoreAdvisor's assigned store on the appEdge device ID mapped to storeExact match
TimestampTap to record, network timeEvent time, network timeA few seconds; drift over five seconds breaks joins
ZoneAdvisor's default counter, or the phone's assigned deskZone tag on the visitor track at tap timeExact, with adjacent zones as a fallback flagged as weak
Session start and endTap start, tap stopEngagement event, exit eventThe camera session is usually longer; the conversation sits inside it
IdentityNone usedNone usedNot a key. The join works without it, by design

Confidence, and what to do with weak joins

A join is a claim that this conversation and this visual track belong to the same visit. Treat it as a claim with evidence. Strong joins drive the store-week aggregates. Weak joins are kept, counted and shown separately, because the pattern of weakness is itself information: a counter that is always ambiguous is a counter with too many advisors on it, or a camera in the wrong place. Never promote a weak join to strong by hand to make a number look better.

The failure modes worth naming in advance are three. A crossing, where two shoppers pass each other and one track splits into two, so a nine-minute visit becomes a four and a five. Two advisors near one shopper, so the tap cannot be assigned with certainty. And a customer who leaves the frame and returns, so one visit is counted as two entries. Each has a fix on the floor, and each should be visible in the weekly review rather than buried in an average.

Privacy checklist

  • The camera pipeline emits events only: entries, dwells, engagements, walk-outs. No frames, no faces, no facial recognition, no re-identification across visits or stores.
  • Video is processed on the edge device in the store. What leaves the building is a list of timestamped events.
  • Audio is captured only on the advisor's phone, only after the customer agrees, and is deleted after transcription.
  • The join uses store, time and zone. It creates no new personal data.
  • Reports are by store and week. The customer is anonymous everywhere except the lead they chose to leave.
  • Under India's DPDP Act, notice at the entrance covers the camera events, and per-conversation consent covers the audio. Write both down, and keep the consent evidence with the transcript.

What the combined record tells you

The conversation is the account of what happened between the customer and the advisor. The camera is the account of what surrounded it. Voice supplies the intent, the objection, the offer quoted and the outcome. Vision supplies the footfall, the wait, the group, the path and the customers who left without a word. Joined on the advisor's tap, by time and zone, they produce a visit record that a store manager can act on the same week.

Borentis today runs the conversation side of this: consented capture on the advisor's phone in Hindi, English and Hinglish, scored against the retailer's playbook, live as Borentis Floor. ShopperDNA, which would add the camera events and the join described in these steps, is on the roadmap. The zone map and the clock sync are worth doing now, whichever vendor supplies the vision layer.

Frequently asked questions

How do you match a conversation to the right person on camera?

You do not match to a person. You match the advisor's tap, with its time, store and zone, to visitor tracks in that zone at that time, and you score how certain the match is. The shopper is never identified.

Do we need to replace our cameras to combine audio and video analytics?

Usually not. The constraint is coverage of the counters and display zones where conversations happen, and an edge device to process the streams. Door-only footfall cameras will not support the join.

What happens when the advisor forgets to tap?

The visit is vision-only. It is counted, timed and, if nobody engaged, flagged as an unattended walk-out. Capture rate, conversations over entries, is reported for exactly this reason.

How accurate does the join need to be?

Accurate enough that store-week aggregates rest on strong joins and weak joins are reported separately. A pilot should expect a large share of strong joins at quiet counters and fewer at crowded ones, and should fix the floor before tuning the model.

Related reading

Where Borentis applies this

Borentis is the Agentic Operating System for Customer Interactions, built for Indian retail floors: consented one-tap capture on the advisor's phone, every conversation scored against your playbook with the evidence behind every number, leads created when a number is heard, and coaching from your own best conversations.