SoftwareSourced
Case Study: Arcee AI cut time-to-first-token 95% moving from AWS to Together AI endpoints
Arcee AI Case StudySourced & dated by Case Study Desk
Key results
The challenge
Arcee AI serves many small language models routed by its Conductor system, but operating its own GPU fleet on AWS with load balancing, autoscaling, and node sharing required dedicated engineering and added latency.
The solution
Arcee AI migrated 7+ models to Together AI dedicated endpoints via a private Hugging Face repository handoff, offloading GPU fleet management to Together AI.
“Conductor will route to one of our models, which are 95% cheaper than GPT and Sonnet.”
MMMark McQuadeCEO, Arcee AI
The results, in context
Together AI's page reports Arcee AI improved time-to-first-token by up to 95%, reducing some models' latency from 485ms on AWS, and reached 41+ queries per second at 32 concurrent requests, while eliminating self-managed GPU overhead.
Products used
Together AI Dedicated EndpointsTogether AI Serverless Inference