BORENTIS

Store operations

Visual Merchandising Audit Software With AI: Scoring Windows and Displays From a Phone Photo

Visual merchandising audit software with AI takes a photograph of a store window, a mannequin group, a display car, a demo bay or an accessories wall from the store phone and returns a compliance score against the brand's docket in seconds, with the miss circled on the photo. It replaces the VM manager's three-week wait to learn how a docket landed across 300 stores, and the area manager's habit of judging a window from a blurry group forward. This guide explains what the model actually checks, how a verdict is produced, where the model is strong, where a human still decides, and why the models built for general trade give poor verdicts on a brand's own showroom.

What the software does

A docket goes out on a Friday: window W39, the navy set on mannequin 1, the brown coat on mannequin 2, the campaign backdrop, the festive bunting along the facade, the new accessory on the wall at the entrance. Three hundred stores are meant to set it by Saturday opening. Before the software, the VM manager finds out how it landed from photographs posted to groups over the next fortnight, judged by eye, one by one, with no record of which stores were checked.

With the software, each store photographs its window from a marked spot on Saturday morning inside the app. The model compares the photo to the docket and to the reference image and returns a score per element: mannequin 1 correct, mannequin 2 wearing last week's set, backdrop in place, bunting missing on the left. The misses are circled, a redo is requested in one tap, the store reshoots, and by 11:00 the VM manager has a compliance map of 300 stores rather than 40 forwards.

The software is one module of a store operations platform, alongside checklists, scored audits and tickets, described at /learn/store-operations-software-for-retail-chains-india/. It can run on its own, but it works best when a failed verdict becomes a ticket automatically.

What the AI checks, element by element

ElementThe question the model answersTypical miss it catches
Facade and signageIs the signage lit, clean and current; is the festive dressing in place?One letter of the fascia out; last festival's bunting still up
WindowAre the mannequins, props and backdrop those of the current docket?Previous docket left in place; backdrop missing; mannequin count wrong
MannequinsIs each mannequin in the named look, styled and complete?Brown coat on the wrong mannequin; missing accessory; price ticket showing
Display cars and demo unitsAre the display cars clean, priced, positioned and lit as the standard says?Missing price board; wrong variant on the turntable; door left open
Accessories wallIs the wall stocked to the layout, with the launch accessory at the marked position?Empty hooks; last season's hero in the hero slot
Campaign elementsAre the standees, danglers and counter cards of this campaign present?Old campaign standee beside the new one
Billing counterIs the counter clear, the counter card current, the tent card present?Clutter; personal items; wrong offer card
GroomingIs the team in the current uniform and badge?Wrong colour day; missing badge

How the verdict is produced

  1. The brand loads the docket: the reference photo, the element list and the display standards for that window type and store format. A flagship window and a 400-square-foot EBO window have different dockets.
  2. The store photographs the window from the marked spot inside the app. The camera is live only, with GPS and time locked in; the app rejects a photo that is too dark, too far or too tilted before the model sees it.
  3. The model finds the elements in the photo: it locates each mannequin, the backdrop, the props, the signage and the campaign pieces.
  4. It compares each element to the docket: is the look on mannequin 2 the brown coat; is the backdrop the W39 backdrop; is the campaign standee the current one.
  5. It scores each element and the window overall using the brand's weights, and circles each miss on the photo.
  6. A question can be asked of the photo in plain language, "is the brown coat on mannequin 2?", and answered in seconds.
  7. A fail becomes a redo request to the store and, if the redo does not arrive in the window, a ticket to the area manager. Every verdict, redo and override is logged with time and user.

Where the AI is strong and where a human still decides

  • Strong: presence and absence. Is the backdrop there, is the standee the current one, is the hero accessory in the hero slot. This is the bulk of VM compliance and the model does it in seconds across hundreds of stores.
  • Strong: identity of a look. Which named outfit is on which mannequin, which variant is on the turntable, when the reference images exist.
  • Strong: repeatability. The same photo gets the same score in every store on every day, which no team of area managers achieves.
  • Weaker: quality of styling. Whether the scarf is draped well, whether the lighting flatters the coat, whether the whole reads as the brand. The model can flag "differs from reference"; a VM manager decides whether the difference is a fault or a good local call.
  • Weaker: unusual stores. A corner window, a mezzanine display bay or a mall kiosk needs its own reference and its own marked spot, or the model will score against the wrong docket.
  • Human always: the override. When the store says "Sir, brown coat is out of stock in our size run, we used the camel", a person accepts the deviation and the log records who accepted it and why.

General trade models and showroom models

Most image AI for retail was built for FMCG and general trade, where a brand's product sits inside a retailer's store. Those models are described by their vendors in the vocabulary of that world: planogram match, facings, share of shelf. That is the right vocabulary for a biscuit brand auditing a kirana, and the wrong one for a brand auditing its own window, because a showroom has nothing of that kind to match. It has a facade, a window, mannequins, display cars, demo units, an accessories wall, campaign elements, a billing counter and people in uniform, and the standard is the brand's own docket, not a retailer's layout.

When you evaluate a vendor, ask it to score one of your own store photos against one of your own dockets in the demo. A general trade model will find the products and count them; a showroom model will find the mannequins and tell you which one is wearing the wrong coat. The difference is visible in the first minute.

Festive dockets: the fortnight the software pays for itself

The Indian festive window, from Navratri through Diwali and again for the wedding season and the year-end sale, is when dockets change weekly, when every store is busiest, and when a window set three days late costs the most. It is also when a VM manager physically cannot see 300 windows. In those weeks the model does what the team could not: a compliance map of the whole network by mid-morning on docket day, the repeat offenders named by the second week, and the redo requested before the weekend footfall arrives rather than after.

The same weeks produce the honest measure of a docket itself. If 40 percent of stores miss the same element, the element is hard to execute or the kit did not arrive, and that is a supply chain and design finding, not a store discipline finding. Only a network-wide score, taken the same way in every store, makes that distinction possible.

Drishti in BorentisOps

Drishti is the visual merchandising AI inside BorentisOps at /solutions/products/borentisops/. It is trained on the brand's own dockets and display standards for windows, mannequins, display cars, demo units, accessories walls, signage, billing counters and grooming, and it scores every element of the docket rather than the overall impression. A store photographs the window from the store phone, no fixed cameras, and receives a compliance score in seconds with the miss circled. A question can be asked of the photo in plain words and answered the same way, and a redo is accepted the same way.

A failed verdict becomes a ticket with an owner and an SLA, the VM manager sees a compliance map of the network within the hour of docket day, and the VM score lands on the store scorecard beside the operations score and the sales-conversation score from Borentis Floor. How the model is trained and checked is described at /learn/how-ai-verifies-visual-merchandising-photos/.

Frequently asked questions

What is visual merchandising audit software with AI?

It is software that takes a live photograph of a window, a mannequin group, a display car, a demo bay or an accessories wall from the store phone and scores it against the brand's docket in seconds, element by element, with each miss circled. It replaces judging windows from group forwards and gives the VM manager a compliance map of every store on docket day.

Does visual merchandising AI need cameras in the store?

No. The models in this category work from a photograph taken on the store phone from a marked spot, inside the app, with GPS and time locked in. Fixed cameras are not required for docket compliance. Camera-based signals such as footfall and dwell are a separate capability and, at Borentis, are on the roadmap rather than in the product today.

How accurate is AI at checking visual merchandising?

On presence and absence, and on identifying a named look or variant when reference images exist, the models are reliable and, above all, consistent across hundreds of stores. On styling quality and unusual store layouts they flag a difference and a person decides. Ask any vendor to score one of your own store photos against your own docket in the demo before believing an accuracy figure.

Can general trade image AI be used for a showroom window?

Poorly. Models built for FMCG and general trade look for a brand's products inside a retailer's store and count them against the retailer's layout. A showroom window has mannequins, props, a backdrop and campaign pieces, and a display bay has cars or demo units, so the right model is one trained on the brand's own dockets and display standards for those elements. The difference shows in the first minute of a demo.

Related reading

Where Borentis applies this

Borentis is the Agentic Operating System for Customer Interactions, built for Indian retail floors: consented one-tap capture on the advisor's phone, every conversation scored against your playbook with the evidence behind every number, leads created when a number is heard, and coaching from your own best conversations.