How AI-powered anomaly detection transformed data quality across 2.5M+ retail products
One of the largest retail mart chains in the United States was struggling with growing data quality issues across its Retail Management System (RMS). Inconsistent supplier feeds, abbreviated product descriptions, incorrect product classifications, and supplier mismatches were affecting reporting accuracy, inventory management, and invoice reconciliation.
With over 2.5 million products, 100+ departments, and 6,000+ subclasses, manual data cleanup was no longer practical. The retailer needed a scalable solution that could not only clean existing data but also prevent new anomalies from entering the system in near real time.
Aximise built an AI-powered anomaly detection and data quality platform on Azure that continuously validates, enriches, and monitors retail master data without disrupting the existing RMS, which remained the system of record.
Problem Points
- Products assigned to incorrect departments and subclasses
- Incorrect product sizes, quantities, and Unit of Measure (UOM) values
- Supplier records not matching delivered quantities or packaging information
- Heavy use of abbreviations and shorthand in product descriptions
- No built-in anomaly detection or validation capabilities within the RMS
- Poor data quality is affecting reporting accuracy and supplier reconciliation
- Overpayments caused by invoice matching discrepancies
- Growing data volumes are making manual review impractical
- Need to clean existing RMS data while preventing future anomalies
Solutions Implemented
- Replicated RMS data into Azure using Qlik Replicate with Change Data Capture (CDC)
- Built a scalable Azure Data Lake architecture using bronze, silver, and gold data layers
- Developed an anomaly detection pipeline using Azure Databricks and PySpark
- Applied pattern-based validation and custom algorithms to identify suspicious records
- Used lightweight NLP techniques to detect abbreviated and inconsistent product descriptions
- Leveraged LLM-based expansion for flagged records to standardize incomplete product names
- Implemented embedding similarity validation to verify department and subclass assignments
- Applied business validation rules provided by retail domain experts
- Used statistical modelling and threshold-based analysis to identify attribute-level outliers
- Integrated anomaly detection outputs into supplier feed validation workflows
- Established continuous feedback loops to improve detection accuracy over time
Key Features Added
- Automated anomaly detection across product, supplier, customer, and order datasets
- Near real-time validation of incoming supplier feeds
- AI-powered expansion and standardization of abbreviated product descriptions
- Semantic validation using embedding similarity techniques
- Statistical outlier detection across item attributes
- Department and subclass classification validation
- Azure-based scalable processing architecture
- Continuous RMS monitoring without impacting production systems
- Data stewardship workflows for anomaly review and correction
- Curated gold-layer datasets for analytics and reporting
Business Impact
The solution transformed data quality management from a reactive cleanup exercise into a proactive, automated process. By combining business rules, statistical modelling, NLP, and selective AI enrichment, the retailer significantly improved data accuracy while maintaining low operational costs.
The platform validated more than 2.5 million items, identifying approximately 15% of records for cleanup and reducing anomalous data in production environments by over 95%. Supplier mismatches are now detected before invoices are processed, helping prevent costly overpayments. Multiple departments now rely on trusted, curated datasets for reporting and analytics, restoring confidence in RMS-driven business decisions.
