BORENTIS

Store operations

How AI Verifies Visual Merchandising Photos: From Docket to Score in Seconds

AI verifies a visual merchandising photo by comparing what a store phone captured against a reference standard, the docket, and returning a verdict per check: is the window dressed to this fortnight's plan, is the right outfit on mannequin two, is the display car priced, is the facade lit. The model does not judge taste. It answers a list of yes or no questions the visual merchandising team wrote down, with a confidence for each, and turns the answers into a score in seconds, so the store can redo a failed item while the manager is still standing in front of it. This guide explains the steps in plain language, what image AI can and cannot judge from a phone photo, how the redo loop works, and the honest limits every buyer should hear on the demo.

What a docket is, and why it is the reference

A docket is the visual merchandising team's standard for a period: the reference photo of the window and each display zone, the outfit and accessories list per numbered mannequin, the props, the signage creative and its text, the display cars or demo units that should be on the floor and the rule for their price boards, the accessories wall layout. It changes on a cycle, fortnightly or monthly in apparel, and around festive periods everywhere: Navratri, Diwali, Pongal, Eid, the wedding season.

Image AI is only as good as the standard it is compared against. A model asked "does this look compliant?" with no reference will return a confident opinion about nothing. A model asked "is the brown jacket on mannequin two, as in this reference photo?" can answer, and a person can check the answer. The docket comes first; the model second. That is also why a model trained for a different reference, a shelf in someone else's store with its planogram and facings, is not the right tool for a showroom; the words appear here once because that is the whole difference.

From docket to score in seven steps

  1. The VM team publishes the docket as structured checks, not a PDF: one zone per check, a reference image, a question with a yes or no answer, and a weight. "Mannequin 2: brown jacket, cream trouser, tan bag. Weight 10."
  2. The store opens the checklist in the app and takes a live photo of each zone from the marked spot. The camera opens inside the app; GPS, time and device are attached; a gallery photo is refused.
  3. The model locates the zone in the photo: this is the window, these are mannequins one to four, this is the display car with its price board. If it cannot find the zone, it says so and asks for a retake rather than guessing.
  4. Each check is compared against its reference: presence, count, position, colour and silhouette match, legible text on the signage, light box on or off, visible dust or clutter.
  5. Each verdict carries a confidence. High-confidence passes and fails are returned in seconds. Low-confidence checks are routed to a person, the VM coordinator or the area manager, with both images side by side.
  6. The score is computed from the weighted checks, and shown to the store immediately with the failed items and the reason for each in plain words, in the language the store chose.
  7. Failed items become a redo request in the same session. What is still failing at the end of the session becomes a ticket with an SLA, and whatever is open at the end of the day appears in the 19:30 digest to the store, the area manager and the regional head, with the photos.

What image AI can and cannot judge from a phone photo

The table is the honest version of the demo. "Yes" means a well-built model with a good reference answers reliably; "partly" means it catches the obvious cases and a person handles the rest; "no" means it should not be in the checklist as an AI item at all, and should stay a manager's question.

CheckCan AI judge it?HowHonest limit
Right outfit on each numbered mannequinYesColour and silhouette match against the docket imageClose shades (navy vs black), a mannequin turned away, a garment partly hidden by a prop
Mannequin and prop countYesObject count against the referenceReflections in the glass can double an object; take the photo at an angle
Signage present with the right creative and textYesImage match plus text readingGlare on the light box; small print at a distance
Facade and light box litMostlyBrightness of the light box relative to its surroundPhone exposure compensates for a dim box; pair with a time-of-day rule and a marked spot
Display car or demo unit present and pricedYesObject detection plus price board presenceWhether the amount is correct needs the price list, not the photo
Cleanliness of display cars, counters and floorPartlyVisible dust, clutter and marksFingerprints, faint scuffs and smell are beyond a photo
Team groomingPartlyUniform, badge, standard items presentNeatness is a judgement; keep it a manager's item
Trial rooms and stock room conditionPartlyLights on, door open, clutter visibleAnything outside the frame does not exist to the model
Fold quality, fabric handling, styling tasteNoNot a yes or no questionStays with the VM team on their visit
Whether the customer experience was goodNoNot in the photoMeasured from the conversation, not the window

The redo loop, and why seconds matter

The reason to score in seconds rather than overnight is that the person who can fix the window is standing in front of it at 09:50. A verdict at 09:51 that says "Mannequin 2: jacket does not match docket" gets a jacket changed and a retake at 09:55. The same verdict in a morning-after report gets a ticket, an area manager call, and a window that is wrong for a day of footfall. On the floor the loop sounds like this: "Photo dobara lo, mannequin do pe jacket galat hai." Most faults never become tickets.

The loop needs three rules to stay honest. The retake must be live, from the same spot, in the same session. A person can overrule the model in either direction, and the override is recorded with a name. And the verdict is evidence for a manager's conversation, not a disciplinary record on its own; a model that fails a window because a customer walked into the frame is wrong, and the store should be able to say so in one tap.

Honest limits every buyer should hear

  • Angle and distance change everything. A marked spot on the floor for each zone photo does more for accuracy than any model upgrade.
  • Night, glare and reflections in glass windows are the commonest causes of a wrong verdict. Expect more human review at opening time in winter and in mall windows with strong lighting opposite.
  • Customers, staff and trolleys in the frame block the model's view. Take the photo before opening or ask the model to flag occlusion rather than guess.
  • A new docket needs new reference images before the first audit against it. If the VM team publishes the docket without them, the first cycle is a manual audit.
  • Every model has a false-fail rate and a false-pass rate, and they move with the confidence threshold. Ask the vendor for both on your own photos, not a single accuracy figure from a slide; a single figure without your photos is a marketing number.
  • In the first weeks expect a meaningful share of checks to route to human review, falling as reference images and overrides accumulate. A working range from operations teams that have run this is a few checks per store per day early on, dropping to occasional exceptions within a couple of docket cycles; treat that as a range to test, not a promise.
  • Nothing outside the frame is checked. Stock accuracy, cash, keys and anything in a drawer stay as checklist items with a manager's yes or no.

Where BorentisOps and Drishti fit

Drishti is the visual merchandising AI inside BorentisOps, the store operations platform built for own showrooms and exclusive brand outlets. It is trained on the brand's own dockets and display standards for the facade, window, mannequins, display cars, demo units, accessories wall, billing counter and grooming, and it returns the verdict, the failed items and the reasons in seconds, in Hindi, English or Hinglish. Failed items go into the redo loop, then into tickets with SLA escalation, then into the 19:30 digest, with the operations score sitting next to the sales-conversation score from Borentis Floor on one store scorecard. It is priced per store and works offline, with photos queued and scored when connectivity returns. The product page is at /solutions/products/borentisops/, the buying guide for VM audit software at /learn/visual-merchandising-audit-software-with-ai/, and the audit app requirements at /learn/retail-store-audit-app-with-photo-proof/.

If you buy nothing, the seven steps still work with a docket as structured checks, a marked spot per zone and a VM coordinator reading photos on WhatsApp. It is slower, the redo happens the next day, and the coordinator becomes the bottleneck at fifty stores, which is the point at which the model earns its place.

Frequently asked questions

How accurate is AI visual merchandising verification?

There is no single honest figure. Accuracy depends on the reference images, the photo angle, the lighting and the confidence threshold, and every model has both a false-fail and a false-pass rate that move against each other. Test with your own photos, ten compliant and ten with known faults, and ask for both rates. Design the process so a person handles low-confidence checks and can overrule the model.

Does the AI need to be trained on my brand?

It needs your reference: the docket as structured checks with a reference image per zone, refreshed each docket cycle. A few reference photos per zone from the marked spot are usually enough to start; overrides from the VM team improve it over subsequent cycles. It does not need thousands of labelled images of your stores before the first audit.

Can photo verification work offline?

Capture can. The photo is taken live with GPS and time attached and stored on the phone, then uploaded and scored when connectivity returns, with the original capture time preserved. The verdict in seconds needs a connection, so a basement showroom will get its redo loop when it is back online rather than in the moment. BorentisOps is built offline first for this reason.

Is this the same as planogram compliance software?

No. Planogram compliance software reads a shelf in a store the brand does not own and counts positions against a plan. Visual merchandising verification for showrooms and exclusive brand outlets reads the brand's own window, mannequins, display cars and demo units against the brand's own docket. The methods overlap; the reference, the vocabulary and the owner of the fix do not.

Related reading

Where Borentis applies this

Borentis is the Agentic Operating System for Customer Interactions, built for Indian retail floors: consented one-tap capture on the advisor's phone, every conversation scored against your playbook with the evidence behind every number, leads created when a number is heard, and coaching from your own best conversations.