Case Study Deskcasestudydesk.com
Generative AISourced

Case Study: Suno launches faster by running inference on Modal instead of Kubernetes

Suno Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Suno
Industry
Generative AI
Challenge
Suno wanted to scale inference without running its own Kubernetes or reserving GPUs long-term.
Headline result
Suno shaved 4 months off its launch and auto-scaled past 1,000 GPUs on demand

Key results

4 mo
Shaved off launch timeline
by avoiding added infrastructure hires
1,000+
GPUs at peak with auto-scaling

The challenge

Suno's founders wanted to avoid running their own Kubernetes infrastructure, diverting engineers to it, and committing to long-term GPU reservations. They needed to scale inference and batch pre-processing efficiently through demand spikes.

The solution

Suno deployed batch pre-processing and model inference on Modal, using its auto-scaling compute and HTTP function endpoints. This let the team run workloads without managing infrastructure or configuration files.

Modal reminded me of the difference between PyTorch and TensorFlow, where Torch catered more to the ML crowd and was okay deviating from some CS principles. That's the beauty of Modal.

GK
Georg Kucsko
Co-founder and CTO, Suno

The results, in context

Suno estimated it shaved 4 months off its launch timeline by not having to hire additional infrastructure engineers. Its workloads auto-scaled past 1,000 GPUs during peak demand periods such as holidays, without pre-committed capacity.

Products used

Modal Modal GPU computeModal Serverless inference