Case Study: Decagon cut cost per turn 6× and hit sub-400ms voice latency on Together AI
Key results
The challenge
Decagon builds conversational AI agents for enterprise customer support, including voice, where latency directly shapes the user experience. Closed-model APIs were costly per turn and struggled to consistently meet the sub-second latency bar that voice interactions require.
The solution
Decagon ran open models on Together AI using NVIDIA Blackwell GPUs with speculative decoding, high tensor parallelism, and prompt caching, deploying updated models on a weekly (sometimes daily) cadence.
“...optimizing our models with techniques like speculative decoding, and they’ve been a reliable production partner.”
MLMax LuHead of Research, Decagon
The results, in context
According to Together AI's published page, Decagon reduced p95 model latency per turn from seconds to under 400ms on inputs up to tens of thousands of tokens, and achieved roughly 6× lower cost per turn compared with closed models such as GPT-5 mini, while sustaining an average deflection rate above 80%.