Case Study Deskcasestudydesk.com
Healthcare AISourced

Case Study: Sully.ai cuts inference cost 90% moving to open-source on Baseten

Sully.ai Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Sully.ai
Industry
Healthcare AI
Challenge
Closed-source inference costs and latency did not fit clinical scale
Headline result
90% lower inference cost and 65% lower median latency

Key results

90%
Inference cost savings
65%
Lower median latency
30M
Minutes added to workforce
21x
Return on agent spend (ROAS)

The challenge

Sully.ai builds autonomous clinical and administrative agents that operate in live healthcare environments where latency and accuracy are critical. Running on proprietary closed-source models, Sully hit elevated and inconsistent response times as request volume grew, along with token- and request-based costs that scaled linearly with usage. Those costs became a bottleneck to pricing its agents affordably enough to serve all types of health systems.

The solution

Sully moved its real-time inference workloads from closed-source models to open-source models served on Baseten, gaining more control over performance and deployment characteristics. The shift let the team tune latency and cost for its clinical-grade, real-time use cases.

The results, in context

Moving to open-source models on Baseten delivered 90% inference cost savings and 65% lower median latency, with a 21x return on agent spend. Sully reported adding roughly 30 million minutes back to the healthcare workforce through its agents.

Products used

Baseten Open-source model deploymentBaseten Baseten Inference Stack