IT·Checklists
    • Web Application Security Review
    • Cloud Security Posture
    • Security Code Review
    • Secrets Management
    • Security Incident Response
    • Penetration Test Readiness
    • Production Readiness Review
    • CI/CD Pipeline Review
    • Container Image Hardening
    • Kubernetes Deployment
    • Infrastructure as Code Review
    • Release Day Runbook
    • Cloud Landing Zone Setup
    • Cloud Cost Optimisation
    • Cloud Migration Cutover
    • Serverless Readiness
    • Code Review
    • REST API Design
    • Database Schema Migration
    • New Service Bootstrap
    • Open Source Release
    • Frontend Performance
    • Web Accessibility (WCAG)
    • Observability
    • On-Call Handover
    • Incident Management
    • Postmortem
    • Backup and Recovery
    • Capacity Planning
    • Data Pipeline Review
    • Data Quality
    • ML Model Deployment
    • Data Warehouse Migration
    • Network Change
    • DNS Migration
    • TLS Certificate Management
    • Load Balancer Configuration
    • GDPR Readiness
    • ISO/IEC 27001 ISMS
    • SOC 2 Audit Readiness
    • Vendor Security Assessment
    • Employee IT Onboarding
    • Employee IT Offboarding
    • Change Management
    • Disaster Recovery Drill
    • IT Asset Management
    IT·Checklists
    • GitHub
  • to navigate
  • to select
  • to close
    • Home
    • SRE & Operations
    On this page
    monitor_heart

    SRE & Operations

    Keeping the system healthy in production — observability, on-call, incidents, backups, and capacity.

    insights

    Observability

    Verify a service emits the signals needed to diagnose an unfamiliar failure without shipping new code.

    phone_in_talk

    On-Call Handover

    Verify nothing is dropped when the pager changes hands between two on-call shifts.

    report

    Incident Management

    Verify an incident is detected, coordinated, communicated, and closed out without improvisation.

    history_edu

    Postmortem

    Verify an incident review is blameless, accurate, and produces action items that actually get done.

    backup

    Backup and Recovery

    Verify backups exist, are protected from your own mistakes, and have been restored successfully at least once.

    trending_up

    Capacity Planning

    Verify a service has measured headroom, credible demand forecasts, and limits that fail safely under load.


    © 2026 IT Checklists · Built with Hugo and Lotus Docs