Case Study: Sully.ai cuts inference cost 90% moving to open-source on Baseten
Key results
The challenge
Sully.ai builds autonomous clinical and administrative agents that operate in live healthcare environments where latency and accuracy are critical. Running on proprietary closed-source models, Sully hit elevated and inconsistent response times as request volume grew, along with token- and request-based costs that scaled linearly with usage. Those costs became a bottleneck to pricing its agents affordably enough to serve all types of health systems.
The solution
Sully moved its real-time inference workloads from closed-source models to open-source models served on Baseten, gaining more control over performance and deployment characteristics. The shift let the team tune latency and cost for its clinical-grade, real-time use cases.
The results, in context
Moving to open-source models on Baseten delivered 90% inference cost savings and 65% lower median latency, with a 21x return on agent spend. Sully reported adding roughly 30 million minutes back to the healthcare workforce through its agents.