Case Study: How Handshake saves 50% on LLM GPU costs with Anyscale
Key results
The challenge
Handshake's small ML infrastructure team needed to deploy and fine-tune large language models and deep learning recommenders at scale without reserving expensive A100 GPU capacity, and without requiring significant ML Ops experience from its data scientists.
The solution
Handshake used the Anyscale platform to host large (Mixtral 8x22B) and small (Llama-3 8B) LLM endpoints, deploy graph neural networks and two-tower recommenders, and scale training across multiple GPUs and machines using Ray features such as fractional GPUs.
“We can scale with cheaper GPUs, operating at a fraction of the cost by not having to reserve A100-GPU capacity and taking advantage of Ray features like fractional GPUs. Anyscale also handles dependency management automatically as we move things around.”
KGKyle GallatinSenior ML Infra Engineer, Handshake
The results, in context
On Anyscale, Handshake saved 50% on cloud costs (more than 50% on LLM GPUs) and gained 5x faster iteration for AI workloads and 10x scalability for LLM GPUs by using cheaper commodity GPUs. After deploying graph neural networks and two-tower recommenders on Anyscale, engagement on jobs increased by 90% year over year.