Case Study Deskcasestudydesk.com
Developer ToolsSourced

Case Study: Zed serves 2x faster code completions with 3.6x throughput on Baseten

Zed Industries Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Zed Industries
Industry
Developer Tools
Challenge
One week to launch Edit Prediction with latency targets unmet
Headline result
3.6x throughput and 45% lower p90 latency

Key results

45%
Lower p90 latency
3.6x
Higher throughput
100%
Uptime
2x+
Faster Edit Prediction
vs. previous provider

The challenge

Zed Industries builds a high-performance, Rust-based code editor whose Edit Prediction feature, powered by its open-source Zeta model, was one week from launch. Its existing inference provider was not meeting Zed's targets of P90 under 500 ms and P50 under 200 ms, and offered limited compute capacity, single-region deployment, and little visibility into what drove model performance.

The solution

Baseten paired Zed with forward deployed and model performance engineers who tried over 75 optimizations within days, including TensorRT-LLM with KV caching and custom-tuned speculative decoding, lookahead decoding, multi-cloud capacity management, and geo-aware routing. Because the platform is OpenAI-compatible, Zed moved all traffic over within a day with no code changes.

I want the best possible experience for our users, but also for our company. Baseten has hands down provided both. We really appreciate the level of commitment and support from your entire team.

NS
Nathan Sobo
Co-founder, Zed Industries

The results, in context

Baseten exceeded Zed's initial performance goals with 45% lower p90 latency, 3.6x higher throughput, and 100% uptime across unlimited autoscaling. With additional post-launch improvements, Zed now delivers over 2x faster Edit Prediction with Zeta compared to its previous inference provider.

Products used

Baseten TensorRT-LLMBaseten Speculative decodingBaseten Multi-cloud capacity managementBaseten Geo-aware routing