Case Study Deskcasestudydesk.com
Productivity SoftwareSourced

Case Study: Notion cut AI response latency from ~2 seconds to 350ms with Fireworks AI

Notion Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Notion
Industry
Productivity Software
Challenge
AI feature latency too high to launch at scale
Headline result
Fine-tuning on Fireworks delivered 4x faster AI responses at scale

Key results

350ms
AI response latency
down from ~2 seconds
4x
Faster responses
after fine-tuning

The challenge

Notion set out to deliver AI features to a very large user base without the response times undermining the experience. Baseline model latency of roughly two seconds was too slow for interactive use.

The solution

Notion fine-tuned models served on Fireworks AI to reduce response latency while preserving quality, enabling AI features to be rolled out across its user base.

By fine-tuning models, we reduced latency from about 2 seconds to 350 milliseconds, significantly improving performance and enabling us to launch AI features at scale.

SS
Sarah Sachs
Head of AI Engineering, Notion

The results, in context

After fine-tuning, Notion reduced AI response latency from about 2 seconds to 350 milliseconds, roughly a 4x improvement, which the company says let it launch AI features at scale. Notion states it serves more than 100 million users.

Products used

Fireworks AI Fireworks AIFireworks AI Model fine-tuningFireworks AI Serverless inference