Case Study: Notion deploys new frontier models in under 24 hours with evals on Braintrust
Key results
The challenge
Notion set out to give customers access to the latest frontier models as quickly as possible, ideally within hours of release, which required rigorous evaluation at scale. The AI team needed to keep roughly 70 engineers aligned on a shared evaluation practice and to surface narrow, needle-in-a-haystack failures affecting specific segments such as multilingual workspaces across very large LLM traces.
The solution
Notion adopted Braintrust as its evaluation framework, pairing regression evals that catch breakage with frontier evals that measure improvements from new models. The team deployed custom evaluation code for Notion-specific data structures and used Brainstore for performant search across large trace volumes, enabling systematic iteration from evaluation to production.
“I first started working with Braintrust on my first day at Notion. I sat down in Braintrust and looked at some of the worst experiences our customers had and tried to understand how we can be better.”
SSSarah SachsAI Modeling Lead, Notion
The results, in context
Notion reported deploying new frontier models in under 24 hours of release while keeping about 70 engineers aligned on evals. The company stated that roughly 80% of what its AI team does is based on evaluating feedback and traces in Braintrust.