Case Study: Writer serves custom 70B LLMs with 60% higher throughput on Baseten
Key results
The challenge
Writer builds domain-specific Palmyra LLMs for enterprises, including 70-billion-parameter models such as Palmyra-Med-70B and Palmyra-Fin-70B for compliance-heavy industries. Serving 70B models requires multiple high-end A100 or H100 GPUs and cutting-edge optimization, detail-oriented inference work that pulled focus from the team's core competence of model training.
The solution
Writer worked with Baseten's model performance and forward-deployed engineers to build TensorRT-LLM model-specific engines for each LLM, compiling specialized CUDA instructions tuned to real-world sequence shapes and batch sizes and using in-flight batching for production serving.
“Inference for custom-built LLMs could be a major headache. Thanks to Baseten, we're getting cost-effective high-performance model serving without any extra burden on our internal engineering teams. Instead, we get to focus our expertise on creating the best possible domain-specific LLMs for our customers.”
WAWaseem AlshikhCTO and Co-Founder, Writer
The results, in context
In a benchmark running the LLMs in FP16 on four NVIDIA A100 GPUs, Writer saw 60% higher tokens per second, 23% lower time to first token, and 35% lower cost per million tokens. Writer surpassed its performance requirements ahead of launching the new models.