Case Study: Coursera delivers 45x more feedback with AI grading evaluated on Braintrust
Key results
The challenge
Coursera's AI grading and feedback initiative relied on fragmented, offline evaluation using spreadsheets and manual human labeling, with teams writing separate scripts and lacking a shared way to collaborate. This made it difficult to validate AI features with confidence before shipping them to learners.
The solution
Coursera built a structured evaluation process on Braintrust: defining success criteria up front, curating datasets that combine real-world and synthetic edge-case examples, applying hybrid scorers that pair deterministic checks with LLM-as-a-judge, and iterating with online monitoring and offline batch testing.
The results, in context
Coursera reported that learners now receive grades within about one minute of submission and roughly 45x more feedback through AI grading, alongside a 90% learner satisfaction rating. The company also reported a 16.7% increase in course completions within a day of peer review.