02 / RELIABILITY

Reliability practices

SLOs you would defend in a review, alerts worth waking up for, and incident write-ups that change the system instead of the roster. Most teams we meet do not need more monitoring. They need less noise.

What this actually involves

01
SLOs tied to what users feel

Not uptime for its own sake. We start from what a bad experience actually looks like for your customer and build the target back from there.

02
Cutting alerts, not adding them

We go through the alert list and delete anything that has fired without anyone acting on it. Most on-call fatigue is a pruning problem, not a tooling one.

03
Incident review that produces changes

A review that ends without an owned action item did not happen, as far as the system is concerned. We run reviews that end with something that gets built.

04
On-call your engineers can sustain

Rotation size, escalation paths and what counts as a page. Small changes here are usually why retention on a platform team gets better or worse.

Signs this is worth a call

On-call engineers are burning out and the rotation keeps getting rewritten instead of the system.

The same incident keeps recurring under a slightly different name.

Leadership is asking for an SLA and nobody is confident enough in the current numbers to commit to one.

Other practices
Process audit & improvement → AI deployment & integration → Systems architecture & deployment →

Start with the two-week diagnostic