Case Study: Reducto cuts P90 latency 3x by moving 30+ models to Modal
Key results
The challenge
Reducto ran 30+ production models on Kubernetes that had to scale together despite varied usage patterns, producing unpredictable P90 latency during large bursty uploads of millions of pages. Traffic spikes strained infrastructure and threatened customer SLAs.
The solution
Reducto migrated to Modal, using GPU memory snapshotting, independent per-model scaling, and parameterized functions for customer-specific autoscaling pools. Granular regional compute control gave the team more placement flexibility.
“Modal gives us a lot of flexibility to do pretty complex stuff that we wouldn't get with an LLM inference service.”
RCRaunak ChowdhuriFounder, Reducto
The results, in context
Reducto reported a 3x reduction in P90 latency and an 83% reduction in cold boot times, from roughly 70 seconds to about 12 seconds. It scaled to 1,000+ GPUs in under an hour during load testing and reduced endpoint deployment from about 150 lines plus configuration to 2 lines of code.