Case Study: ZipRecruiter gained 3x faster query performance and eliminated ELK downtime with Logz.io
Key results
The challenge
ZipRecruiter, an online employment marketplace, ran a self-managed ELK Stack to aggregate and analyze logs from roughly 100 services across a hybrid monolith-and-microservices architecture. As its Elasticsearch cluster grew to ingest about 2 TB of data per day, the SRE team battled high CPU and memory consumption, sluggish or crashing queries, and around 3-4 Elasticsearch-related incidents per month. Because the team relied on log data to debug production, this downtime was highly disruptive.
The solution
ZipRecruiter migrated to Logz.io, a managed log-management service built on the ELK Stack, which required only changing the Logstash output destination and no retraining on Kibana. The team now ships approximately 3 TB of log data per day into Logz.io, using Filebeat and Kafka to forward logs from around 100 services written in Java, Python, Scala, Perl, C++, and Go.
“Logz.io has had a major impact on our work, allowing us to focus more on what matters - building, deploying and monitoring our product.”
ABAlon BeckerSRE, ZipRecruiter
The results, in context
After migrating, ZipRecruiter reported a 3x improvement in query performance, with Kibana queries that previously crashed Elasticsearch now executing reliably and making troubleshooting faster. Eliminating ELK maintenance and downtime let the SRE team redirect time to other initiatives, including a migration to Kubernetes.