Case Study: Gretel achieves 10x experimentation velocity for synthetic-data models with Weights & Biases
Key results
The challenge
Gretel builds synthetic-data models and needed to fine-tune them faster across domains, including a Text-to-SQL model. Its experimentation cycles were slow, and its datasets originally comprised over a trillion tokens, limiting how quickly the team could iterate.
The solution
Gretel adopted Weights & Biases experiment tracking, logging, and evaluation tools to identify gaps, make adjustments, and launch new experiment batches every few days rather than queuing a handful at a time.
“W&B's logging and evaluation tools let us quickly identify gaps, make adjustments, and launch new experiment batches every few days. Instead of queuing up 5-10 experiments, we could now run 50-100 experiments in each compute block.”
AWAlex WatsonCo-Founder and CPO, Gretel
The results, in context
Gretel reported 10x experimentation velocity, scaling from 3-5 experiments per week to 250 experiments in just 45 days before its first model launch, and from 5-10 to 50-100 experiments per compute block. Its Text-to-SQL model achieved a 62% improvement in overall performance and a 35% enhancement in SQL task-specific correctness, while training data was reduced from a trillion to a billion tokens — a 1000x reduction.