Case Study Deskcasestudydesk.com
SoftwareSourced

Case Study: Arcee AI cut time-to-first-token 95% moving from AWS to Together AI endpoints

Arcee AI Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Arcee AI
Industry
Software
Challenge
Self-managing GPU fleet on AWS added cost and complexity
Headline result
Dedicated endpoints replace self-managed GPU infrastructure

Key results

95%
faster time-to-first-token
latency from 485ms on AWS
41+
queries per second
at 32 concurrent requests
7+
models deployed
on Together AI

The challenge

Arcee AI serves many small language models routed by its Conductor system, but operating its own GPU fleet on AWS with load balancing, autoscaling, and node sharing required dedicated engineering and added latency.

The solution

Arcee AI migrated 7+ models to Together AI dedicated endpoints via a private Hugging Face repository handoff, offloading GPU fleet management to Together AI.

Conductor will route to one of our models, which are 95% cheaper than GPT and Sonnet.

MM
Mark McQuade
CEO, Arcee AI

The results, in context

Together AI's page reports Arcee AI improved time-to-first-token by up to 95%, reducing some models' latency from 485ms on AWS, and reached 41+ queries per second at 32 concurrent requests, while eliminating self-managed GPU overhead.

Products used

Together AI Dedicated EndpointsTogether AI Serverless Inference