Case Study: Decagon cuts voice AI latency 65% with custom inference on Modal
Key results
The challenge
Decagon builds real-time voice AI agents that require sub-second latency, natural turn-taking, and consistent quality at scale. Small-parameter models struggled to generalize, with accuracy failing to transfer cleanly to new domains.
The solution
Working on Modal, Decagon combined supervised fine-tuning and reinforcement learning across multiple verticals with inference optimization. This included a custom EAGLE3 draft model trained via SpecForge and asynchronous scheduler re-engineering in SGLang.
The results, in context
Decagon reported a 65% latency reduction in Voice 2.0 and 38% higher accept lengths from its custom draft model versus open-source baselines. SGLang runtime optimizations added up to a 12% throughput improvement, and p90 latency reached 342ms, below the required sub-second threshold.