BORENTIS

Video and vision

What Is Retail Video Analytics? What It Measures and How It Works

Retail video analytics is software that reads the video from store cameras and turns it into counts and events about people and stock: how many entered, where they stood, how long they waited, whether a shelf was empty. It works by detecting people in each frame, following them across frames, and testing their positions against lines and zones drawn on the camera view. It does not understand what shoppers want. It measures where bodies were and for how long, and that is a useful thing to know as long as nobody pretends it is more.

How it works, layer by layer

  1. Ingest: the software reads a live stream from each camera, usually RTSP from the DVR or NVR, at a few frames per second. Analytics does not need the 25 frames per second that security recording uses.
  2. Detection: a neural network finds people in each frame and draws a box around each one. Modern detectors do this well in daylight at ordinary store distances.
  3. Tracking: a second step links the boxes from frame to frame so that one person is one track with a start time and an end time. This is the fragile layer. Tracks break when people cross, overlap or leave the view.
  4. Geometry: lines and zones are drawn on the image. Crossing the entrance line is an entry. Being inside a zone for longer than a threshold is a dwell.
  5. Event rules: entries, exits, dwells, queue length, two tracks together for a while (a group), a staff track near a shopper track (an engagement), a shopper track that ends without any staff track nearby (an unattended walk-out).
  6. Aggregation: events roll up into counts by hour, zone and store, which is what a manager actually reads.

What retail video analytics measures

Event typeDerived fromCommon useCaveat
FootfallTracks crossing an entrance line, direction-awareConversion denominator, staffing by hourRe-entries and staff need excluding
Zone dwellTime a track spends inside a drawn zoneWhich displays hold attentionCrowding breaks tracks and shortens dwell
HeatmapTrack positions accumulated over a dayLayout and display placementNeeds a near top-down view to mean much
Queue length and waitTracks inside the counter zone and how long each staysBilling staffing, service standardA browser near the counter looks like a queue
Group sizeTracks entering together and staying closeFamily versus solo visits by hourPairs split and re-merge
Staff-shopper proximityA staff track within reach of a shopper track for a few secondsWas the customer attended, and how soonProximity is not a conversation
Shelf gap or planogramComparison of a shelf image to a referenceOn-shelf availability, display complianceNeeds a close, well-lit camera
Unattended walk-outA shopper track with no staff proximity before exitThe lost-customer countDepends on staff exclusion working

What it does not measure

  • Intent: whether the person by the display wanted to buy it or was waiting for a friend.
  • Reason: why a shopper left. The camera sees the leaving, not the cause.
  • What was said: cameras have no useful audio and analytics does not transcribe.
  • Satisfaction or mood: emotion recognition from faces is unreliable and, for a store, unnecessary.
  • Who the person is: the systems described here do not, and should not, identify anyone.

Edge or cloud

Video analytics can run in the cloud, with streams uploaded, or on an edge device in the store. For an Indian store on a shared broadband line, edge is the practical choice: a small computer reads the local streams and sends only events upstream, a few kilobytes an hour. Frames never leave the building. That is also the cleaner position under the DPDP Act, because there is no video of customers sitting on a server to protect, retain or explain.

Where it fits with conversation intelligence

Video analytics measures the outside of the interaction: arrival, wait, attention, attendance, exit. Conversation intelligence measures the inside: what the advisor asked, what the customer objected to, what was offered. Conversion is explained only when the two are read together by time and zone, the walk-outs of 5 to 6 pm beside the conversations of 5 to 6 pm, and never by tying a face to a voice.

Borentis captures consented in-store conversations on the advisor's phone today. ShopperDNA, on the roadmap, is the vision product that would add the outside view from the store's existing cameras and join it to those conversations by time and zone.

The privacy line

  • Behavioural events are anonymous counts: a track, not a person. Biometrics are the opposite, and nothing here needs them.
  • No facial recognition, no age or gender estimation, no emotion reading in the design.
  • Entrance notice covering security recording and anonymous analytics; footage retention per security policy; events kept as aggregates.

Frequently asked questions

Is retail video analytics the same as facial recognition?

No. Facial recognition identifies people. Retail video analytics as described here counts and times anonymous tracks. Some platforms offer both; a retailer can and should run only the first.

How accurate is retail video analytics?

Entrance counts on a well-placed camera are typically within a few per cent after validation. Dwell and group measures are fair in open areas and degrade with crowding. Paths across many cameras are the least reliable.

Does retail video analytics need new cameras?

Not for counting and queues on most modern recorders. Heatmaps, paths and shelf checks usually need a dedicated top-down or close camera.

Can video analytics tell me why conversion is low?

It can tell you how many were unattended, how long they waited and where they left from. The why is in the conversation, which is a separate and consented instrument.

Related reading

Where Borentis applies this

Borentis is the Agentic Operating System for Customer Interactions, built for Indian retail floors: consented one-tap capture on the advisor's phone, every conversation scored against your playbook with the evidence behind every number, leads created when a number is heard, and coaching from your own best conversations.