Case Study: Latent Health delivers pharmaceutical search with 99.999% uptime on Baseten
Key results
The challenge
Latent Health provides clinical question answering to large US health systems using an ensemble of OCR, fine-tuned LLMs, and custom embedding and reranker models over millions of documents daily. Deploying these compound workflows directly on a large cloud provider left the team managing complex infrastructure, struggling to orchestrate many models, and unable to autoscale pipeline stages independently to meet aggressive latency targets.
The solution
Latent adopted Baseten Chains to deploy its compound AI system as independently scaling Chainlets, running LLMs on GPUs while OCR steps ran on CPUs for cost efficiency. Baseten engineers made runtime optimizations with the Baseten Inference Stack, including TensorRT-LLM, plus cross-cluster autoscaling and optimized cold starts.
The results, in context
Latent delivers pharmaceutical search with 99.999% uptime and 600 ms P90 end-to-end latency, alongside 6x improved GPU utilization. Decoupled autoscaling let it optimize both latency and cost as document volume grew.