Case Study: Parallel Web Systems gets 3x throughput and 3x cost savings on Baseten
Key results
The challenge
Parallel powers web research for AI-native teams, running structured agent pipelines where a single job can involve many model calls and tens of thousands of multi-turn requests per hour with bursty traffic. Two weeks before its Search API launch it was serving thousands of multi-hop searches on closed-source models, whose costs would be unsustainable at its projected long-term scale.
The solution
Parallel launched on Baseten Model APIs with pay-per-token pricing to absorb unpredictable demand, then moved to a dedicated deployment of an open-source model. Baseten's model performance team applied speculative decoding and KV cache-aware routing to optimize latency and throughput.
“Our workloads can be very spiky. We need to be able to manage the peaks, but we also don't want to pay for the valleys. Pay-per-token with Model APIs enabled us to easily make the transition to open-source until our scale merited dedicated deployments optimized for throughput and flexible autoscaling.”
MLMatt LeeEngineering Lead, Parallel Web Systems
The results, in context
Baseten's runtime cut latency by 50% and increased throughput 3x through KV cache-aware routing and speculative decoding, delivering 3x cost savings versus the equivalent closed-source model. Within a few months of the initial deployment, Parallel scaled traffic on Baseten roughly 100x.