Case Study: Abnormal Security cuts observability costs and halves incident detection time with Chronosphere
Key results
The challenge
As Abnormal Security's customer base grew rapidly since its 2018 launch, its homegrown Prometheus and Grafana monitoring system reached its limits, with 10-12 million active metrics on pace to reach 50 million. The team faced constant metric outages, a limited two-day retention period, dashboards that took over 30 minutes or failed to load, and an expensive, memory-intensive Amazon EC2 R5 instance, all while trying to keep up with over 300% year-over-year business growth and a 99.9% SLA target.
The solution
Abnormal chose Chronosphere's cloud native observability platform, having ruled out running Thanos in-house and other SaaS options. Chronosphere's control plane let Abnormal aggregate 98% of its metrics and offered flexible retention intervals, while native Prometheus ingestion meant no instrumentation changes were required.
The results, in context
Aggregating 98% of its metrics made Chronosphere 10x more cost-effective than the alternative SaaS and self-managed options Abnormal evaluated. Abnormal reduced MTTD and MTTR by 80% based on SLOs, loaded dashboards 8-10x faster, and improved reliability to greater than 99.9% uptime.