BORENTIS

Multimodal AI

Voice Intelligence vs Video Intelligence vs Multimodal Retail Intelligence

Voice intelligence, video intelligence and multimodal retail intelligence are three different purchases, often presented as one. Voice intelligence scores the consented conversation. Video intelligence counts and times the visit from camera events. Multimodal retail intelligence fuses the two at the event level. This comparison sets them side by side on what they measure, what they cost, how they fail and what they let a retailer decide, so the choice is made on evidence rather than on a vendor's slide.

Three-way matrix

DimensionVoice intelligenceVideo intelligenceMultimodal retail intelligence
Unit of analysisConversationVisit or trackJoined visit (conversation plus its surrounding events)
Primary outputPlaybook adherence, objections, offers, leadsFootfall, dwell, wait, group, walk-outs, hotspotsConversion by wait, group and path; unattended walk-outs next to capture rate
Depends on floor behaviourYes: advisor taps to recordNoPartly: joins exist only where a tap happened
Consent modelPer conversation, evidencedSignage; anonymous eventsBoth; the join adds no new personal data
ProcessingCloud transcription and scoring after consentEdge device in store; events leave, frames do notEdge for video, cloud for voice, fusion on events
Typical failureLow coverage; noisy audio; code-switching errorsTrack splits at crossings; occlusion; staff counted as visitorsAmbiguous joins: two advisors near one group; clock drift
Coaches an advisorStep by step, with transcript linesAssociate-level conversion at bestStep by step, with wait and group as context
Recovers a lost customerYes, if a number was sharedNoYes, plus knows how many were never spoken to
Cost driverSpeech minutesCameras, edge compute per storeBoth, plus the zone map and clock sync
India availabilityLive (Borentis Floor among others)Live (footfall and video analytics vendors)ShopperDNA on the roadmap; few joined products in market

What voice intelligence does well and where it stops

Voice intelligence is the only one of the three that knows why. It hears the need, the objection in the customer's words, the offer quoted right or wrong, the rival named, and the number shared for follow-up. It can coach an advisor from a transcript line and recover a walk-out from a lead. It is also the only one that depends on the floor: if the advisor does not tap, nothing is captured. It cannot count the customers who never reached a conversation, and it cannot see how long the ones who did had waited.

What video intelligence does well and where it stops

Video intelligence is objective and continuous. It counts every entry without anyone's cooperation, times dwell by zone, measures the wait before engagement and flags visitors who left unattended. It runs on an edge device from the store's existing cameras and emits events, not faces. It cannot hear, so it cannot explain a single walk-out, cannot tell a price objection from a delivery objection, and cannot tell a customer who was served badly from one who was served well and left to consult a spouse.

What multimodal adds and what it costs

  • Adds: the questions in the matrix that neither half can answer alone. Does a long wait change the objection? Do groups convert differently after the same pitch? What share of entries got no conversation at all?
  • Adds: an honest denominator. Capture rate becomes consented conversations divided by camera entries, not by an estimate.
  • Costs: a zone map per store, cameras covering the conversation zones, clock sync between phones and cameras, and a fusion layer that scores each join.
  • Costs: a new class of failure. A crossing splits one visual track into two. Two advisors near one shopper make it unclear whose tap belongs to which track. These are reported as weak joins, not hidden.
  • Does not cost: any new consent from the shopper. The video half stays anonymous; the voice half is already consented; the join uses time and zone.

Which to buy first

  1. If conversion is fine and the question is layout, queues or staffing, buy video intelligence and stop there.
  2. If conversion is the problem and the store already has advisors talking to most walk-ins, buy voice intelligence. It explains the loss and recovers part of it.
  3. If the problem is a mix of unattended walk-outs and weak conversations, buy voice first, because it changes behaviour on the floor within weeks, and plan the vision layer so the two join later on the same tap.
  4. If a vendor offers all three today as one product, ask how the join works, what confidence it reports and where the video is processed. The answers separate audio video intelligence from a demo.

The house position, plainly

The conversation is the record of the interaction. The camera is the record of its surroundings. Voice yields intent, objection, offer and outcome; vision yields footfall, wait, group, path and unattended walk-outs. The right join is on the advisor's tap, by time and zone, and never by identity. Borentis is voice first and live as Borentis Floor, capturing consented conversations in Hindi, English and Hinglish. ShopperDNA, the multimodal product, is on the roadmap.

Frequently asked questions

Is video intelligence the same as footfall counting?

Footfall counting is the simplest form of it. Video intelligence adds zone dwell, wait time, group size and walk-out detection from the same cameras. Neither hears anything.

Is multimodal AI just voice AI plus computer vision in one dashboard?

Two dashboards side by side is not multimodal. The word earns its meaning when the two are joined at the visit level, with a confidence score, so one record holds the wait, the conversation and the exit.

Which is cheaper to start with?

Voice, usually. It runs on the advisor's phone and needs no hardware in the store. Video needs an edge device per store and camera coverage of the right zones.

Can multimodal retail intelligence work without consent for audio?

No. The voice half is consented or it does not exist. Camera audio is not a substitute; it is unconsented and should not be processed.

Related reading

Where Borentis applies this

Borentis is the Agentic Operating System for Customer Interactions, built for Indian retail floors: consented one-tap capture on the advisor's phone, every conversation scored against your playbook with the evidence behind every number, leads created when a number is heard, and coaching from your own best conversations.