Case Study Deskcasestudydesk.com
SoftwareSourced

Case Study: Zapier lifts AI feature accuracy from 50% to 90%+ with eval-driven development on Braintrust

Zapier Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Zapier
Industry
Software
Challenge
Take AI features from sub-50% prototypes to production quality without shipping unproven models.
Headline result
Zapier reported improving an AI feature's accuracy from about 50% to 90%+ within roughly two to three months using an eval-driven process on Braintrust.

Key results

50% → 90%+
AI feature accuracy improvement
2-3 months
To reach production-quality accuracy

The challenge

As one of the earliest adopters of generative AI, Zapier needed a systematic way to take AI features from early prototypes to production quality while managing the risk of shipping unproven models. Early versions could launch with sub-50% accuracy, and the team needed confidence to iterate safely.

The solution

Zapier adopted an eval-driven process on Braintrust: prototyping with flagship models, shipping a limited V1, collecting explicit and implicit user feedback, and building golden datasets from real customer examples. Changes were then tested against these evals before shipping, letting the team iterate and later swap in cheaper models with confidence.

The results, in context

Zapier reported improving AI feature accuracy from around 50% to 90%+ within roughly two to three months of production-quality iteration. The evaluation framework let the team expand availability and capabilities as accuracy increased.

Products used

Braintrust Braintrust (evals, datasets)