Case Study: Suno launches faster by running inference on Modal instead of Kubernetes
Key results
The challenge
Suno's founders wanted to avoid running their own Kubernetes infrastructure, diverting engineers to it, and committing to long-term GPU reservations. They needed to scale inference and batch pre-processing efficiently through demand spikes.
The solution
Suno deployed batch pre-processing and model inference on Modal, using its auto-scaling compute and HTTP function endpoints. This let the team run workloads without managing infrastructure or configuration files.
“Modal reminded me of the difference between PyTorch and TensorFlow, where Torch catered more to the ML crowd and was okay deviating from some CS principles. That's the beauty of Modal.”
GKGeorg KucskoCo-founder and CTO, Suno
The results, in context
Suno estimated it shaved 4 months off its launch timeline by not having to hire additional infrastructure engineers. Its workloads auto-scaled past 1,000 GPUs during peak demand periods such as holidays, without pre-committed capacity.