Case Study: Zapier lifts AI feature accuracy from 50% to 90%+ with eval-driven development on Braintrust
Key results
The challenge
As one of the earliest adopters of generative AI, Zapier needed a systematic way to take AI features from early prototypes to production quality while managing the risk of shipping unproven models. Early versions could launch with sub-50% accuracy, and the team needed confidence to iterate safely.
The solution
Zapier adopted an eval-driven process on Braintrust: prototyping with flagship models, shipping a limited V1, collecting explicit and implicit user feedback, and building golden datasets from real customer examples. Changes were then tested against these evals before shipping, letting the team iterate and later swap in cheaper models with confidence.
The results, in context
Zapier reported improving AI feature accuracy from around 50% to 90%+ within roughly two to three months of production-quality iteration. The evaluation framework let the team expand availability and capabilities as accuracy increased.