Case Study Deskcasestudydesk.com
SoftwareSourced

Case Study: Decagon cut cost per turn 6× and hit sub-400ms voice latency on Together AI

Decagon Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Decagon
Industry
Software
Challenge
Voice agents need sub-second latency at sustainable cost
Headline result
Sub-second voice AI at roughly a sixth of closed-model cost per turn

Key results

cost reduction per turn
vs. closed models like GPT-5 mini
<400ms
p95 model latency per turn
inputs up to tens of thousands of tokens
80%+
average deflection rate
across concierge interactions

The challenge

Decagon builds conversational AI agents for enterprise customer support, including voice, where latency directly shapes the user experience. Closed-model APIs were costly per turn and struggled to consistently meet the sub-second latency bar that voice interactions require.

The solution

Decagon ran open models on Together AI using NVIDIA Blackwell GPUs with speculative decoding, high tensor parallelism, and prompt caching, deploying updated models on a weekly (sometimes daily) cadence.

...optimizing our models with techniques like speculative decoding, and they’ve been a reliable production partner.

ML
Max Lu
Head of Research, Decagon

The results, in context

According to Together AI's published page, Decagon reduced p95 model latency per turn from seconds to under 400ms on inputs up to tens of thousands of tokens, and achieved roughly 6× lower cost per turn compared with closed models such as GPT-5 mini, while sustaining an average deflection rate above 80%.

Products used

Together AI Dedicated EndpointsTogether AI NVIDIA Blackwell GPUsTogether AI Serverless Inference