Eliminating Outages and High Costs During Peak Shopping Load
As a digital payments leader, Affirm requires continuous system reliability; outages immediately risk lost sales and damage merchant relationships. However, legacy observability tools frequently failed under heavy traffic spikes, leaving engineers blind during critical sales events like Black Friday. Detecting system destabilization often took hours and field type or limit conflicts caused issues with log ingestion, creating costly operational delays. At the same time, high telemetry volumes ingested low-value data noise, driving up storage costs while diluting data quality for internal developers. To protect revenue and maintain checkout performance across millions of consumers, Affirm needed an enterprise observability solution that could handle extreme traffic scale, unify metrics and logs and deliver real-time insights without failing during peak demand.
Real-Time Telemetry and Scalability with Cortex XCOR
Affirm selected Cortex XCOR from Palo Alto Networks to transform its cloud-native monitoring infrastructure. The engineering team embedded directly with Affirm to execute a smooth, three-month migration. During this transition, Affirm unified metrics and high-cardinality logs into a streamlined set of 55 dashboards on Cortex XCOR. By utilizing the solution’s aggregation tier, Affirm achieved a 91% data optimization rate, eliminating low-value noise and improving data quality for developers.
The transformation delivered immediate operational stability and financial efficiency. During Black Friday, Affirm saw significantly more traffic, requiring the team to scale up infrastructure by 4x. Cortex XCOR handled this scale — and the increased telemetry that came with it — with zero downtime. Affirm accelerated issue detection from hours to seconds while achieving significant cost savings despite sending higher overall data volumes.
“We've moved massive volumes of metrics, logs, and traces to Cortex XCOR, and they've delivered an exceptional product in terms of system availability, scalability, and cost controls."
— Arjun Nayak
Staff Observability Engineer, Affirm