Keeping the system healthy in production — observability, on-call, incidents, backups, and capacity.
Observability
Verify a service emits the signals needed to diagnose an unfamiliar failure without shipping new code.
On-Call Handover
Verify nothing is dropped when the pager changes hands between two on-call shifts.
Incident Management
Verify an incident is detected, coordinated, communicated, and closed out without improvisation.
Postmortem
Verify an incident review is blameless, accurate, and produces action items that actually get done.
Backup and Recovery
Verify backups exist, are protected from your own mistakes, and have been restored successfully at least once.
Capacity Planning
Verify a service has measured headroom, credible demand forecasts, and limits that fail safely under load.