Case Study: Zed serves 2x faster code completions with 3.6x throughput on Baseten
Key results
The challenge
Zed Industries builds a high-performance, Rust-based code editor whose Edit Prediction feature, powered by its open-source Zeta model, was one week from launch. Its existing inference provider was not meeting Zed's targets of P90 under 500 ms and P50 under 200 ms, and offered limited compute capacity, single-region deployment, and little visibility into what drove model performance.
The solution
Baseten paired Zed with forward deployed and model performance engineers who tried over 75 optimizations within days, including TensorRT-LLM with KV caching and custom-tuned speculative decoding, lookahead decoding, multi-cloud capacity management, and geo-aware routing. Because the platform is OpenAI-compatible, Zed moved all traffic over within a day with no code changes.
“I want the best possible experience for our users, but also for our company. Baseten has hands down provided both. We really appreciate the level of commitment and support from your entire team.”
NSNathan SoboCo-founder, Zed Industries
The results, in context
Baseten exceeded Zed's initial performance goals with 45% lower p90 latency, 3.6x higher throughput, and 100% uptime across unlimited autoscaling. With additional post-launch improvements, Zed now delivers over 2x faster Edit Prediction with Zeta compared to its previous inference provider.