BORENTIS

Video and vision

Can AI Analyze CCTV Footage? What Works in a Retail Store Today

Yes, AI can analyze CCTV footage, and in a retail store it does five things reliably: find people, follow each one for a few seconds, count them across a line, time them inside a zone, and notice when a shelf or display changes. What it cannot do on the footage a store actually has is tell you who someone is, why they left or what they said. Most disappointment with AI on retail CCTV comes from the gap between those two lists.

What works today on ordinary footage

  • Person detection at store distances in daylight and under normal shop lighting, including from wide-angle security cameras, with weaker results at the frame edges.
  • Short tracks: following one person for the seconds or minutes they are in one camera's view without crossing anyone.
  • Line counting: entries and exits across a threshold, with direction.
  • Zone timing: how long a track stayed within a drawn area, and how many tracks were there at once.
  • Change detection on a fixed view: a display that moved, a shelf that emptied, a door left open.

What works only with the right camera

  • Paths across the floor: needs overlapping, near top-down views so tracks hand over between cameras. Security mounts along a wall do not give this.
  • Product-level shelf checks: needs a camera within a couple of metres of the shelf, facing it squarely, in even light.
  • Reliable group detection: needs the entrance camera to see the pavement or lobby side so groups arrive together in one view.
  • Night and low light: works if the camera has a decent sensor; cheap infra-red night modes lose the detail detectors rely on.

Accuracy caveats by task

TaskTypical result on 2 to 8 security camerasMain failureHow to check
Entrance countWithin a few per cent after tuningRe-entries, staff, delivery peopleHand count two hours on two days
Zone dwellFair in open zones, worse when crowdedTrack breaks and ID switchesFollow ten shoppers by eye and compare
Queue lengthGood at a clear counterBrowsers standing near the counterPhotograph the queue at five random times
Group sizeRoughPairs separate and rejoinSample twenty entries by eye
Staff exclusionGood with a rule, poor withoutStaff in plain clothes, promotersCompare staff count to roster
Path across camerasPoor on wall mountsNo overlap between viewsDo not promise it
Product interactionPoorHands and products too smallDo not promise it
Face, age, gender, emotionNot attemptedExcluded by designConfirm it is switched off

How to analyze retail CCTV footage: recorded or live

  1. For a baseline, export a week of recordings from the DVR and run the analytics on the files. This costs nothing on the floor and shows what the cameras can support before anyone buys hardware.
  2. For ongoing use, read the live streams on an edge box in the store. Frames are processed locally and discarded; only events leave. Never upload footage to a general-purpose AI service to describe it; that moves customer video off the premises for no analytical gain.
  3. Draw lines and zones once per camera and keep them under version control, because moving a display moves the truth.
  4. Validate each measure against a manual count before it appears on any report.
  5. Report counts by hour, zone and store, never by person.

What AI cannot read from a frame

A frame shows a person standing in front of a television for ninety seconds and walking out. It does not show that they asked for the 55-inch model in stock at a rival, that the advisor did not mention the exchange offer, or that they left to consult a spouse. The camera captures the surroundings of the sale. The conversation captures the sale. Conversion makes sense only when the two are laid side by side by time and zone, and the identity of the shopper is not needed to do it.

Borentis captures consented in-store conversations on the advisor's phone today and scores them on the playbook. ShopperDNA, on the roadmap, would join camera events from existing CCTV to those conversations, so a ninety-second dwell followed by a walk-out is read beside what was said on that floor at that time.

The privacy line

  • The systems worth running produce events, not identities. No facial recognition, no age or gender estimation, no emotion reading.
  • Under India's DPDP Act, anonymous counting under notice is the conservative position; profiling customers by face is not.
  • Keep footage retention as the security policy sets it; keep analytics outputs as aggregates; process on the edge so no frame leaves the store.

Frequently asked questions

Can AI analyze old CCTV recordings?

Yes, exported DVR files run through the same detection and tracking as live streams. It is the cheapest way to find out what your cameras can support.

Can a general AI chatbot analyze my CCTV footage?

General multimodal models can describe a frame or a short clip. They do not run continuous counting across a day, and uploading store footage to a third-party service raises a privacy question that a store does not need to raise.

Does AI CCTV analysis work at night or in low light?

It degrades. Detectors need detail that cheap infra-red modes remove. Stores that trade after dark should check the evening count separately.

Can AI tell why a customer walked out?

No. It can flag that a shopper left unattended after a long dwell. The reason lives in what was or was not said, which is a consented conversation, not a frame.

Related reading

Where Borentis applies this

Borentis is the Agentic Operating System for Customer Interactions, built for Indian retail floors: consented one-tap capture on the advisor's phone, every conversation scored against your playbook with the evidence behind every number, leads created when a number is heard, and coaching from your own best conversations.