Executive summary
Accurate retail execution without adding work for the frontline.
A global consumer health company needed a more accurate and scalable way to verify execution across OTC pharmacies and HPC retail in China. Clobotics connected fine-grained recognition, channel rules, evidence checks, and human review within one continuously operated workflow.
The business context
In a low-growth market, every store matters
China’s pharmaceutical market entered 2026 under pressure. In the first quarter, all-channel pharmaceutical sales declined 5.4% year over year, while retail increased by only 0.6%, making it the only major terminal channel to remain in positive growth.
When the retail channel is protecting limited growth, details inside every pharmacy and store matter more: whether the right product is available, whether its price is correct, whether campaign materials are in place, and whether the shelf position meets the intended execution standard.
For this global consumer health company, the challenge was not simply collecting more store photos. It was converting those photos into accurate, explainable, and operationally fair decisions at scale.
Source: Sinohealth, “2026 Q1 Pharmaceutical Market,” citing CMH data.
The challenge
Pharmaceutical recognition is harder, and the cost of error is greater
A store visit can appear straightforward. A field representative arranges the products, places the required price tags and promotional materials, takes a photograph, and submits the task. A single recognition error can still break the workflow.
The system may identify the correct brand but select the wrong product variant because two packages use almost the same color, typography, and layout. The only meaningful difference may be a dosage or pack-size line printed in small text near the corner of the box. If the wrong SKU is recorded, the result cannot support a dependable execution decision.
The cost extends beyond one image. The representative may need to rearrange the display, adjust the angle, take another photograph, or raise a dispute. If task acceptance is connected to performance-based compensation, repeated errors can reduce trust in the system and make the process feel unfair. Headquarters may also receive a distorted record of what happened in the store.

Distinguish highly similar products
The program maintained 482 products: 85 OTC SKUs and 397 HPC SKUs. Brand-level recognition was not enough when variants differed only by small printed details.
Understand where each product was displayed
Presence was only the first question. The system also needed to interpret the fixture, shelf boundaries, levels, and relative position before applying the customer’s standard.
Make AI dependable for frontline work
High-volume cases needed fast automation, while obstructed, ambiguous, or disputed submissions required a clear path to human review.
Two channels, two definitions of good execution
OTC and HPC share a visual foundation, but not one ruleset
Both channels required SKU, facing, position, price, and POSM recognition. The meaning of compliant execution differed by channel.
Fine-grained variant decisions
Execution depended on recognizing the exact dosage or pack, matching the required price tag and POSM, and evaluating shelf level within the correct fixture.
- Operational sensitivity
- Task decisions could affect field-representative performance outcomes.
- Change pattern
- About five new POSM recognition requests per month.
Portfolio scale and display impact
Execution focused on presence, facings, display area, endcaps, floor displays, and whether priority products occupied high-visibility positions.
- Operational sensitivity
- Assortment growth and display impact drove the workload.
- Change pattern
- Approximately 5-20 new SKU recognition requests per month.
The underlying visual recognition capability could be reused, but the final execution decision had to reflect each channel’s task cards, display formats, and performance logic.
The solution
One workflow for recognition, verification, and action
Clobotics combined computer vision, multimodal AI, business rules, and human exception review in one retail execution workflow.
Collect store evidence
Field representatives captured the products, shelves, price tags, promotional materials, and displays required by each task.
Recognize the product and its context
The AI identified SKUs, facings, price tags, POSM, fixtures, shelf boundaries, levels, and spatial relationships within the same submission.
Apply channel-specific standards
Recognized evidence was evaluated against the customer’s OTC or HPC task requirements instead of one universal definition of compliance.
Focus people on the exceptions
Ambiguous or disputed submissions were routed to review, while each completed task returned a clear pass, correction, or further review outcome.
What the system verifies
From product recognition to complete visit verification
The customer needed more than a product count. A completed task also depended on price accuracy, the presence of required recommendation materials, and confidence that the submitted evidence belonged to the assigned visit.
Confirm the exact SKU
Variant, dosage or specification, presence, and facings prevent a brand-level match from being treated as the right product.
Interpret the physical context
Fixture type, shelf boundaries, level, display area, and relative position determine which execution standard applies.
Connect the label to the product
OCR value, price-tag type, and product association reduce incorrect matches to neighboring or promotional prices.
Verify content as well as presence
Material type, placement, required text, and configured identifiers confirm that the intended promotional material is in place.
Protect the visit record
Reused-image detection, screen-recapture indicators, and scene requirements improve confidence that evidence belongs to the assigned visit.
Return one clear next step
Every submission resolves to pass, correction required, or exception review for both the representative and headquarters.
Price recognition requires more than OCR
For clear and consistently formatted labels, OCR could read the visible price. The harder task was assigning that price to the correct product and interpreting it in context. A valid decision required a closed relationship between the product, its position, the relevant label, and the task requirement. Learn more about pricing and promotion verification.
POSM recognition verifies content as well as presence
The program covered 274 POSM types, including price tags, shelf talkers, backboards, large displays, and floor displays. Multimodal models helped verify not only that a material existed, but that it was the correct material and contained the required text or number combination.
Evidence integrity protects the visit record
The workflow checked for reused images and visual patterns associated with photographing a screen. It could also verify configured scene requirements. These controls made each visit record easier to trust and audit.
Continuous model operations
New products do not wait for a mature image library
The project was not a one-time deployment. New products, packaging updates, and campaign materials entered the recognition scope every month.
New items created a cold-start problem. Before launch, there were not enough store images to represent every angle, lighting condition, obstruction, and display format. Initial recognition accuracy for a newly introduced item was typically around 70%.
Clobotics used synthetic images for pre-training before launch, incorporated real submissions during the first days in stores, and returned human-reviewed errors to the training set during ongoing operations. Human attention stayed focused on the submissions that required it, while the model continued to improve from real operating evidence.
Whether the change involved new monthly POSM in the OTC channel or updated SKUs in HPC retail, models, rules, and human operations evolved together. For retail AI, deployment is only the beginning.
Business impact
More accurate execution without placing more work on the frontline
Product, price, POSM, placement, and evidence-integrity checks were connected within the same task workflow. Representatives spent less time on capture, validation, and reporting, while shorter visits and faster task processing increased field coverage without adding work to each visit.
Higher recognition accuracy and an explicit review path gave representatives greater confidence that completed work would be evaluated consistently. Difficult and disputed cases were directed to exception review, while headquarters received a clearer record of what field teams completed and why each task was accepted, corrected, or reviewed.
The operating lesson
Recognition and operations must improve together
This program showed that pharmaceutical retail AI must do more than recognize additional products. It must distinguish small packaging differences, interpret shelf position in context, apply different rules across OTC and HPC channels, and connect price, POSM, placement, and evidence integrity within a single visit.
The deeper change came from giving headquarters and frontline teams one shared evidence base. What the representative completed, why the system reached a decision, which submissions were unusual, and which cases required review could all be traced through the same workflow.
For Clobotics, this is what Physical AI means in retail: connecting real store conditions with models, business rules, and human operations so management becomes more precise, field execution becomes more efficient, and retail data remains aligned with what happened in the market. Explore the wider Clobotics retail platform and its approach to on-shelf availability.
Frequently asked questions
Why is pharmaceutical product recognition more difficult than standard CPG recognition?
Products within the same pharmaceutical brand or range can use nearly identical colors, typography, and package layouts. The difference between two SKUs may depend on a small dosage, specification, or pack-size line. Reliable execution therefore requires SKU-level recognition rather than a brand-level match.
How does Clobotics distinguish visually similar pharmaceutical SKUs?
Clobotics trains models on real store images and difficult examples involving unusual angles, complex lighting, and partial obstruction. Human-reviewed errors return to the training set so performance continues to improve during operation.
How are shelf positions evaluated in pharmacies and HPC retail?
The system identifies the fixture, shelf boundaries, levels, and relevant spatial relationships before applying the customer’s rules for that channel and display type. This prevents one shelf number from being treated as equally valuable across tall fixtures, endcaps, floor displays, and checkout environments.
What happens when the AI is uncertain or a result is disputed?
Highly similar products, severe obstruction, unusual scenes, and disputed outcomes can be routed to human exception review. In this program, AI accuracy exceeded 95%, while final accuracy surpassed 99% after review.
How quickly can the system learn a new product or promotional material?
Synthetic pre-training gives the model an initial understanding before enough store images exist. Real submissions and daily error feedback then improve performance after launch. In this program, recognition rose from approximately 70% at cold start to more than 90% within the first week.
Can one workflow support both OTC pharmacies and HPC retail?
Yes, but the final evaluation cannot rely on one generic ruleset. The visual foundation can be shared while task logic, display standards, POSM requirements, shelf-position rules, and review priorities are configured for each channel.
Make every store visit faster, fairer, and easier to verify.
Share your target channels, SKU portfolio, and current field workflow. Talk to the Clobotics retail team about a pharmaceutical or consumer health execution program.