A policy is a series of steps with delays between them. Each step notifies more loudly than the last, and the whole chain stops the instant anyone acknowledges — from an email link, an SMS, or by pressing 4 on the phone call. No dashboard login required at 3am.
Severity
Critical
Time to ack
31m
Policy
prod-critical
maya@ — email + Slack
maya@ · +254 7•• ••• 412
press 4 to acknowledge
voice:+254 7•• ••• 412
escalation stops the moment someone acknowledges — no login required
Email, then SMS, then a voice call — and it ends the moment it is acknowledged.
Steps address the person who is on call at that moment, fixed addresses, or both. The on-call member is resolved live from the schedule, so a policy written a year ago still pages whoever is genuinely covering tonight.
A schedule is built from layers. Each layer has its own members, its own rotation length, and optionally its own coverage window — so a business-hours layer and a nights-and-weekends layer can sit on the same schedule and hand off correctly between them. The topmost layer covering the current moment wins.
The bug most on-call tools ship
When something fires, Atlas opens an incident and attaches the paging workflow to it. The Incidents page is therefore the single list of things going wrong, each row carrying its own live state — Paging · step 2, or Acked · maya@, with an acknowledge button inline.
There is deliberately no second list of active pages elsewhere in the product. Two views of the same problem is how a problem gets acknowledged in one place and worked in neither.
The fastest way to make monitoring useless is to page someone about something that resolved itself. Several rules exist purely to prevent that:
Anomaly detection
PagerDuty and Opsgenie charge per responder, per month — the bill grows every time you add someone to the rotation. Here the on-call scheduling, the multi-step escalation, the SMS and the voice calls are part of the plan, not a per-seat line item. Put the whole team on a rotation without watching a meter.
And it is the same tool that already holds your metrics, uptime and incidents, so the thing that pages you is the thing that saw the problem — no integration to wire up, nothing to keep in sync.