How a French unicorn cut root-cause time to under 5 minutes and on-call stress by 80%
The team that owns Brevo's entire transactional email pipeline finds root causes in 2 to 5 minutes, onboards new on-call engineers in a third of the time, and stays on task when a page fires.

Time to root cause
from 10-15 min
MTTR, high-sev incident
from 15 min
On-call onboarding
from 15 days · 3x faster
Engineering time saved / year
~5 hrs per week
About Brevo
Brevo is an all-in-one customer relationship platform used by businesses worldwide to run their email, SMS and marketing automation. Its transactional email service carries the messages recipients are actively waiting on, which makes reliability a direct part of the product promise. The team in this story owns that entire pipeline.
| Sector | Customer relationship management |
| Team | Transactional email platform |
| Scale | 15 to 20+ microservices, 3 engineers on call |
Engineers on the transactional team
Alerts per month
Incidents per month, all severities
The Challenge
A hard day used to start in the middle of the night
Transactional messages carry information the recipient is already waiting on: confirmations, codes, receipts. A delay is a direct hit to customer experience, so every incident in this pipeline is business-critical, and a small team absorbs roughly 200 alerts a month.
A page would fire, sleep would break, and the engineer had to open a laptop and start digging with no head start. When several alerts landed at once, that meant triaging a pile of unknowns under pressure, alone, at 2am. Roughly half of every investigation went to finding the signal rather than fixing the problem, and onboarding a new on-call engineer took three weeks.
“If an engineer is bombarded with multiple alerts at the same time, it's helpful that an AI agent is already investigating all the alerts in parallel, so the engineer doesn't need to panic.”
Piyush Singh, Lead Engineer 2, Brevo
The Redis Cascade
A queuing Redis instance backing the transactional pipeline degraded, and multiple downstream services began failing at once. With so many services affected simultaneously, the symptoms were everywhere and the true source stayed hidden.
What ewake did
It cut through the noise of multiple failing services and pinpointed the specific affected Redis instance, tracing the impact on sendmail-notifications back to a release in a different service. Where an engineer would have ruled out services one by one, ewake correlated the signals and surfaced the true origin immediately.
At Stake
of customers impacted, with business-critical messages sitting in a backlog
Time to Root Cause
against 15 to 30 minutes and extra engineers pulled in to triage in parallel
Results
Before ewake, and with ewake
| Metric | Before | With ewake |
|---|---|---|
| Time to root cause (typical high-sev alert) | 10-15 min | 2-5 min |
| Time to fully mitigate (MTTR, high-sev incident) | 15 min | 10 min |
| Time to onboard a new on-call joiner | 15 days | 5 days |
of incident hypotheses pinpoint the exact root cause
narrow to the right service or recent change
context-switching into external tools mid-investigation
Hard outcomes
- Root-cause time down to 2-5 min on high-sev alerts
- MTTR cut by a third, 15 → 10 min
- On-call onboarding 3x faster, 15 → 5 days
- ~260 engineering hours reclaimed per year
Soft outcomes
- On-call stress reduced by ~80%. A page is now a check, not a scramble
- Newer engineers reason about incidents faster, with ewake acting as an on-call buddy that walks a joiner through every alert
- Parallel analysis of many simultaneous alerts, each returning a concrete recommended next step
“It changed the way we used to see an alert. Now we know ewake is sitting there just to investigate the alert and provide the details, so we can continue our work without stopping in between.”
Aayush Agrawal, Senior Software Engineer, Brevo
See ewake on your own stack
The fastest way to understand ewake is to point it at something real.

