Case Study Deskcasestudydesk.com
Conversational AISourced

Case Study: Decagon cuts voice AI latency 65% with custom inference on Modal

Decagon Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Decagon
Industry
Conversational AI
Challenge
Decagon needed sub-second latency for real-time voice AI agents at scale.
Headline result
Decagon reduced Voice 2.0 latency 65% and lifted draft-model accept lengths 38%

Key results

65%
Latency reduction in Voice 2.0
38%
Higher draft-model accept lengths
vs open-source baselines
12%
Throughput improvement
SGLang runtime optimizations, up to
342ms
p90 latency
below sub-second threshold

The challenge

Decagon builds real-time voice AI agents that require sub-second latency, natural turn-taking, and consistent quality at scale. Small-parameter models struggled to generalize, with accuracy failing to transfer cleanly to new domains.

The solution

Working on Modal, Decagon combined supervised fine-tuning and reinforcement learning across multiple verticals with inference optimization. This included a custom EAGLE3 draft model trained via SpecForge and asynchronous scheduler re-engineering in SGLang.

The results, in context

Decagon reported a 65% latency reduction in Voice 2.0 and 38% higher accept lengths from its custom draft model versus open-source baselines. SGLang runtime optimizations added up to a 12% throughput improvement, and p90 latency reached 342ms, below the required sub-second threshold.

Products used

Modal Modal GPU computeModal Serverless inference