Separate paging symptoms from diagnostic cause metrics.
Why this matters Good paging is a promise to interrupt a human only when human action is needed soon. Google SRE alerting practice favors symptom-based alerts because symptoms describe user harm: elevated error rate, missed latency objective, or availability budget burn. Cause metrics such as CPU, heap, and queue depth are still essential, but they belong in dashboards, tickets, or diagnostic alerts unless they reliably predict imminent user harm. The mechanism is attention protection. On-call attention is a scarce reliability resource. False pages train responders to distrust the system, while cause-only pages create noise without a user-impact decision. Symptom alerts…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in