Video and vision
What Is Retail Video Analytics? What It Measures and How It Works
Retail video analytics is software that reads the video from store cameras and turns it into counts and events about people and stock: how many entered, where they stood, how long they waited, whether a shelf was empty. It works by detecting people in each frame, following them across frames, and testing their positions against lines and zones drawn on the camera view. It does not understand what shoppers want. It measures where bodies were and for how long, and that is a useful thing to know as long as nobody pretends it is more.
How it works, layer by layer
- Ingest: the software reads a live stream from each camera, usually RTSP from the DVR or NVR, at a few frames per second. Analytics does not need the 25 frames per second that security recording uses.
- Detection: a neural network finds people in each frame and draws a box around each one. Modern detectors do this well in daylight at ordinary store distances.
- Tracking: a second step links the boxes from frame to frame so that one person is one track with a start time and an end time. This is the fragile layer. Tracks break when people cross, overlap or leave the view.
- Geometry: lines and zones are drawn on the image. Crossing the entrance line is an entry. Being inside a zone for longer than a threshold is a dwell.
- Event rules: entries, exits, dwells, queue length, two tracks together for a while (a group), a staff track near a shopper track (an engagement), a shopper track that ends without any staff track nearby (an unattended walk-out).
- Aggregation: events roll up into counts by hour, zone and store, which is what a manager actually reads.
What retail video analytics measures
| Event type | Derived from | Common use | Caveat |
|---|---|---|---|
| Footfall | Tracks crossing an entrance line, direction-aware | Conversion denominator, staffing by hour | Re-entries and staff need excluding |
| Zone dwell | Time a track spends inside a drawn zone | Which displays hold attention | Crowding breaks tracks and shortens dwell |
| Heatmap | Track positions accumulated over a day | Layout and display placement | Needs a near top-down view to mean much |
| Queue length and wait | Tracks inside the counter zone and how long each stays | Billing staffing, service standard | A browser near the counter looks like a queue |
| Group size | Tracks entering together and staying close | Family versus solo visits by hour | Pairs split and re-merge |
| Staff-shopper proximity | A staff track within reach of a shopper track for a few seconds | Was the customer attended, and how soon | Proximity is not a conversation |
| Shelf gap or planogram | Comparison of a shelf image to a reference | On-shelf availability, display compliance | Needs a close, well-lit camera |
| Unattended walk-out | A shopper track with no staff proximity before exit | The lost-customer count | Depends on staff exclusion working |
What it does not measure
- Intent: whether the person by the display wanted to buy it or was waiting for a friend.
- Reason: why a shopper left. The camera sees the leaving, not the cause.
- What was said: cameras have no useful audio and analytics does not transcribe.
- Satisfaction or mood: emotion recognition from faces is unreliable and, for a store, unnecessary.
- Who the person is: the systems described here do not, and should not, identify anyone.
Edge or cloud
Video analytics can run in the cloud, with streams uploaded, or on an edge device in the store. For an Indian store on a shared broadband line, edge is the practical choice: a small computer reads the local streams and sends only events upstream, a few kilobytes an hour. Frames never leave the building. That is also the cleaner position under the DPDP Act, because there is no video of customers sitting on a server to protect, retain or explain.
Where it fits with conversation intelligence
Video analytics measures the outside of the interaction: arrival, wait, attention, attendance, exit. Conversation intelligence measures the inside: what the advisor asked, what the customer objected to, what was offered. Conversion is explained only when the two are read together by time and zone, the walk-outs of 5 to 6 pm beside the conversations of 5 to 6 pm, and never by tying a face to a voice.
Borentis captures consented in-store conversations on the advisor's phone today. ShopperDNA, on the roadmap, is the vision product that would add the outside view from the store's existing cameras and join it to those conversations by time and zone.
The privacy line
- Behavioural events are anonymous counts: a track, not a person. Biometrics are the opposite, and nothing here needs them.
- No facial recognition, no age or gender estimation, no emotion reading in the design.
- Entrance notice covering security recording and anonymous analytics; footage retention per security policy; events kept as aggregates.
Frequently asked questions
Is retail video analytics the same as facial recognition?
No. Facial recognition identifies people. Retail video analytics as described here counts and times anonymous tracks. Some platforms offer both; a retailer can and should run only the first.
How accurate is retail video analytics?
Entrance counts on a well-placed camera are typically within a few per cent after validation. Dwell and group measures are fair in open areas and degrade with crowding. Paths across many cameras are the least reliable.
Does retail video analytics need new cameras?
Not for counting and queues on most modern recorders. Heatmaps, paths and shelf checks usually need a dedicated top-down or close camera.
Can video analytics tell me why conversion is low?
It can tell you how many were unattended, how long they waited and where they left from. The why is in the conversation, which is a separate and consented instrument.
Related reading
- How to use CCTV for retail analytics
- Can AI analyze CCTV footage?
- Why CCTV alone cannot explain retail conversion
- ShopperDNA, on the roadmap
Where Borentis applies this
- Walk-in Recovery: The customer who left is still yours.
- Execution Scorecards: See the floor before the P&L does.
- Objection Intelligence: The reason they did not buy, in their own words.
Borentis is the Agentic Operating System for Customer Interactions, built for Indian retail floors: consented one-tap capture on the advisor's phone, every conversation scored against your playbook with the evidence behind every number, leads created when a number is heard, and coaching from your own best conversations.