Case Study Deskcasestudydesk.com
Internet / Online Employment MarketplaceSourced

Case Study: ZipRecruiter gained 3x faster query performance and eliminated ELK downtime with Logz.io

ZipRecruiter Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
ZipRecruiter
Industry
Internet / Online Employment Marketplace
Challenge
ZipRecruiter's self-managed ELK Stack struggled at scale, with slow or crashing queries and recurring Elasticsearch downtime.
Headline result
3x faster query performance after replacing self-managed ELK

Key results

3x
Improvement in query performance
Kibana queries that previously crashed Elasticsearch
3 TB/day
Log data shipped into Logz.io
Up from ~2 TB/day on self-managed ELK
3-4/month
Elasticsearch incidents before Logz.io
Downtime disrupting production debugging

The challenge

ZipRecruiter, an online employment marketplace, ran a self-managed ELK Stack to aggregate and analyze logs from roughly 100 services across a hybrid monolith-and-microservices architecture. As its Elasticsearch cluster grew to ingest about 2 TB of data per day, the SRE team battled high CPU and memory consumption, sluggish or crashing queries, and around 3-4 Elasticsearch-related incidents per month. Because the team relied on log data to debug production, this downtime was highly disruptive.

The solution

ZipRecruiter migrated to Logz.io, a managed log-management service built on the ELK Stack, which required only changing the Logstash output destination and no retraining on Kibana. The team now ships approximately 3 TB of log data per day into Logz.io, using Filebeat and Kafka to forward logs from around 100 services written in Java, Python, Scala, Perl, C++, and Go.

Logz.io has had a major impact on our work, allowing us to focus more on what matters - building, deploying and monitoring our product.

AB
Alon Becker
SRE, ZipRecruiter

The results, in context

After migrating, ZipRecruiter reported a 3x improvement in query performance, with Kibana queries that previously crashed Elasticsearch now executing reliably and making troubleshooting faster. Eliminating ELK maintenance and downtime let the SRE team redirect time to other initiatives, including a migration to Kubernetes.

Products used

Logz.io Log Management