---
title: "Runbook Automation"
source: https://docs.autohand.ai/guides/sre/runbook-automation
---

# Runbook Automation

Define automated remediation workflows with triggers, actions, safeguards, and infrastructure-as-code integration.

## Overview

Runbooks encode your operational knowledge into executable workflows. Rather than relying on engineers to remember steps during high-pressure incidents, automated runbooks execute consistently and immediately.

A runbook consists of:

-   **Triggers**: Conditions that activate the runbook
-   **Actions**: Steps to execute for remediation
-   **Safeguards**: Controls that prevent harmful outcomes
-   **Rollback**: Recovery steps if actions fail

## Runbook structure

Runbooks are defined in YAML and stored in your repository:

``` yaml
# runbooks/disk-pressure.yaml
name: Disk Pressure Response
description: Handle disk space alerts by cleaning up and scaling storage
version: 1.2.0

trigger:
  metric: disk.used_percent
  operator: ">"
  threshold: 85
  duration: 10m
  labels:
    env: production

actions:
  - name: analyze_usage
    type: command
    command: du -sh /var/log/* | sort -rh | head -10
    timeout: 30s

  - name: cleanup_logs
    type: command
    command: find /var/log -name "*.log" -mtime +7 -delete
    condition: "{{ disk.used_percent > 90 }}"

  - name: expand_volume
    type: terraform
    module: storage/ebs
    variables:
      volume_size: "{{ current_size * 1.5 }}"
    condition: "{{ disk.used_percent > 95 }}"
    approval: required

safeguards:
  max_disk_reduction: 20%
  require_approval_above: 50GB
  dry_run_first: true
  cooldown: 1h

rollback:
  - name: restore_logs
    type: command
    command: "echo 'Manual intervention required'"

notifications:
  slack:
    channel: "#sre-alerts"
    on: [start, complete, failure]
```

## Trigger types

Runbooks can be activated by multiple trigger types:

### Metric triggers

``` yaml
# Trigger on metric threshold
trigger:
  type: metric
  source: prometheus
  query: "sum(rate(http_errors_total[5m]))"
  operator: ">"
  threshold: 100
  duration: 5m
```

### Alert triggers

``` yaml
# Trigger on PagerDuty alert
trigger:
  type: alert
  source: pagerduty
  service: payment-api
  priority: P1
```

### Log pattern triggers

``` yaml
# Trigger on log pattern
trigger:
  type: log
  pattern: "OutOfMemoryError"
  source: kubernetes
  namespace: production
  count: 3
  window: 5m
```

### Scheduled triggers

``` yaml
# Scheduled maintenance runbook
trigger:
  type: schedule
  cron: "0 2 * * 0"  # Sundays at 2am
  timezone: UTC
```

## Action types

Actions define what the runbook executes:

### Command actions

``` yaml
actions:
  - name: restart_service
    type: command
    command: systemctl restart nginx
    timeout: 60s
    retries: 2
```

### Kubernetes actions

``` yaml
actions:
  - name: scale_deployment
    type: kubernetes
    action: scale
    resource: deployment/api-server
    namespace: production
    replicas: "{{ current_replicas + 2 }}"

  - name: rollback_deployment
    type: kubernetes
    action: rollback
    resource: deployment/api-server
    namespace: production
    revision: previous
```

### Terraform actions

``` yaml
actions:
  - name: scale_rds
    type: terraform
    workspace: production
    module: database/rds
    action: apply
    variables:
      instance_class: db.r5.xlarge
    approval: required
```

### API actions

``` yaml
actions:
  - name: toggle_feature_flag
    type: api
    method: POST
    url: "https://api.launchdarkly.com/api/v2/flags/my-project/circuit-breaker"
    headers:
      Authorization: "Bearer {{ secrets.LAUNCHDARKLY_TOKEN }}"
    body:
      patch:
        - op: replace
          path: /environments/production/on
          value: false
```

## Terraform integration

Terraform provides a declarative, version-controlled model of infrastructure. When remediation requires persistent changes, Terraform ensures they are safe and auditable.

### Benefits

-   **Immutable changes**: Generate or modify Terraform plans through a secure pipeline
-   **Drift detection**: Identify when production deviates from desired state
-   **Audit trail**: Every change is logged with context

### Configuration

``` yaml
# .autohand/terraform.yaml
terraform:
  backend: terraform-cloud
  organization: my-company

  workspaces:
    production:
      auto_apply: false
      require_approval: true

    staging:
      auto_apply: true
      require_approval: false

  plan_review:
    enabled: true
    max_resources_changed: 10
    block_on_destroy: true

  notifications:
    slack: "#infra-changes"
```

### Plan review workflow

``` bash
# Runbook generates a plan
autohand sre terraform plan --runbook disk-pressure

# Output:
# Terraform Plan Summary:
# + aws_ebs_volume.data: 100GB -> 150GB
#
# Resources: 0 to add, 1 to change, 0 to destroy
#
# Awaiting approval in Slack #infra-changes
```

## Safeguards

Safeguards prevent runbooks from causing more harm than the original incident:

``` yaml
safeguards:
  # Require human approval for risky actions
  require_approval:
    - type: terraform
      resource_count: "> 5"
    - type: kubernetes
      action: delete

  # Prevent actions during sensitive periods
  blackout_windows:
    - cron: "0 9-17 * * 1-5"  # Business hours
      reason: "No auto-scaling during peak traffic"

  # Limit blast radius
  max_affected_resources: 10
  max_downtime: 5m

  # Rate limiting
  cooldown_period: 30m
  max_executions_per_hour: 3

  # Always dry-run first
  dry_run_first: true

  # Circuit breaker
  abort_on_failure_rate: 50%
```

## Human-in-the-loop

For high-risk actions, configure approval workflows:

``` yaml
approval:
  required_for:
    - terraform_apply
    - kubernetes_delete
    - database_migration

  channels:
    slack:
      channel: "#sre-approvals"
      mention: "@sre-oncall"

    pagerduty:
      escalation_policy: infrastructure

  timeout: 15m
  default_action: abort

  approvers:
    - team: sre
      min_approvals: 1
    - team: platform
      min_approvals: 1
```

When approval is required, the runbook pauses and sends a notification. Engineers can approve, deny, or modify the proposed action.

## Testing runbooks

Validate runbooks before production deployment:

``` bash
# Validate runbook syntax
autohand sre runbook validate runbooks/disk-pressure.yaml

# Dry-run against staging
autohand sre runbook test runbooks/disk-pressure.yaml \
  --environment staging \
  --dry-run

# Simulate trigger condition
autohand sre runbook simulate runbooks/disk-pressure.yaml \
  --trigger "disk.used_percent=92"

# View execution plan
autohand sre runbook plan runbooks/disk-pressure.yaml
```

## Runbook library

Common runbook patterns to get started:

| Runbook | Trigger | Action |
|---|---|---|
| Pod restart loop | CrashLoopBackOff count > 5 | Rollback deployment, notify team |
| High memory pressure | Memory > 90% for 5m | Analyze top processes, scale pods |
| Certificate expiry | Cert expires in < 7 days | Trigger renewal workflow |
| Database connection pool | Available connections < 10% | Identify slow queries, scale RDS |
| API error rate spike | 5xx rate > 1% | Enable circuit breaker, page oncall |

``` bash
# List available runbook templates
autohand sre runbook templates

# Create from template
autohand sre runbook create --template high-memory \
  --name my-api-memory \
  --namespace production
```