Root Cause
The incident was caused by an underlying infrastructure issue, not a defect in the IRIS or WMC applications themselves. A group of backend servers that support these systems became unstable at the same time, and a related configuration issue in a supporting data caching layer prevented the environment from automatically correcting itself as it normally would.
IRIS Impact
No impact to customers. IRIS users experienced intermittent issues accessing the IRIS 2.0 UI.
WMC Impact
Customers attempting to load an order page experienced the following error message:
”Something went wrong on our end. Try again in a few minutes, go back to a previous page, or refresh this page.”
Customers in the process of a purchase experienced endless spinning and processing, followed by an eventual timeout.
• 16 orders were lost in total across the affiliates.
• Approximately 660 unique customers were unable to load order pages at all.
Timeline:
08:55 (EST) Issue onset — error rates begin climbing based on retrospective error distribution analysis.
09:25 (EST) Impact worsens — error and timeout rates increase significantly across WMC order pages and checkout flow.
10:20 (EST) Issue resolved — nodes stabilized and Redis configuration corrected by IT.
Resolution
This was the first time the team had encountered this specific combination of infrastructure instability, which is why the environment's automated safeguards were unable to fully self-correct. IT stabilized the affected infrastructure and corrected the underlying configuration issue to prevent this from happening again.
Follow-up Actions
• Alerting and Monitoring fired as expected across teams allowing immediate investigation.
• Redis configuration has been corrected and additional safeguards are being placed.
• WMC team will follow up with each affiliate affected and provide customer details of any lost orders.