🚨 Runbook: Incident Response agent

> Mode: NEXUS-Micro | Duration: Minutes to hours | Agents: 3-8

by msitarzewskiΒ·MIT licenseΒ·β˜… 154,023 Stars on the repoΒ·GitHub β†—

Files of 🚨 Runbook: Incident Response

msitarzewski/main1 file
scenario-incident-response.md
Show the full text218 lines

🚨 Runbook: Incident Response

Mode: NEXUS-Micro | Duration: Minutes to hours | Agents: 3-8


Scenario

Something is broken in production. Users are affected. Speed of response matters, but so does doing it right. This runbook covers detection through post-mortem.

Severity Classification

Level Definition Examples Response Time
P0 β€” Critical Service completely down, data loss, security breach Database corruption, DDoS attack, auth system failure Immediate (all hands)
P1 β€” High Major feature broken, significant performance degradation Payment processing down, 50%+ error rate, 10x latency < 1 hour
P2 β€” Medium Minor feature broken, workaround available Search not working, non-critical API errors < 4 hours
P3 β€” Low Cosmetic issue, minor inconvenience Styling bug, typo, minor UI glitch Next sprint

Response Teams by Severity

P0 β€” Critical Response Team
Agent Role Action
Infrastructure Maintainer Incident commander Assess scope, coordinate response
DevOps Automator Deployment/rollback Execute rollback if needed
Backend Architect Root cause investigation Diagnose system issues
Frontend Developer UI-side investigation Diagnose client-side issues
Support Responder User communication Status page updates, user notifications
Executive Summary Generator Stakeholder communication Real-time executive updates
P1 β€” High Response Team
Agent Role
Infrastructure Maintainer Incident commander
DevOps Automator Deployment support
Relevant Developer Agent Fix implementation
Support Responder User communication
P2 β€” Medium Response
Agent Role
Relevant Developer Agent Fix implementation
Evidence Collector Verify fix
P3 β€” Low Response
Agent Role
Sprint Prioritizer Add to backlog

Incident Response Sequence

Step 1: Detection & Triage (0-5 minutes)
TRIGGER: Alert from monitoring / User report / Agent detection

Infrastructure Maintainer:
1. Acknowledge alert
2. Assess scope and impact
   - How many users affected?
   - Which services are impacted?
   - Is data at risk?
3. Classify severity (P0/P1/P2/P3)
4. Activate appropriate response team
5. Create incident channel/thread

Output: Incident classification + response team activated
Step 2: Investigation (5-30 minutes)
PARALLEL INVESTIGATION:

Infrastructure Maintainer:
β”œβ”€β”€ Check system metrics (CPU, memory, network, disk)
β”œβ”€β”€ Review error logs
β”œβ”€β”€ Check recent deployments
└── Verify external dependencies

Backend Architect (if P0/P1):
β”œβ”€β”€ Check database health
β”œβ”€β”€ Review API error rates
β”œβ”€β”€ Check service communication
└── Identify failing component

DevOps Automator:
β”œβ”€β”€ Review recent deployment history
β”œβ”€β”€ Check CI/CD pipeline status
β”œβ”€β”€ Prepare rollback if needed
└── Verify infrastructure state

Output: Root cause identified (or narrowed to component)
Step 3: Mitigation (15-60 minutes)
DECISION TREE:

IF caused by recent deployment:
  β†’ DevOps Automator: Execute rollback
  β†’ Infrastructure Maintainer: Verify recovery
  β†’ Evidence Collector: Confirm fix

IF caused by infrastructure issue:
  β†’ Infrastructure Maintainer: Scale/restart/failover
  β†’ DevOps Automator: Support infrastructure changes
  β†’ Verify recovery

IF caused by code bug:
  β†’ Relevant Developer Agent: Implement hotfix
  β†’ Evidence Collector: Verify fix
  β†’ DevOps Automator: Deploy hotfix
  β†’ Infrastructure Maintainer: Monitor recovery

IF caused by external dependency:
  β†’ Infrastructure Maintainer: Activate fallback/cache
  β†’ Support Responder: Communicate to users
  β†’ Monitor for external recovery

THROUGHOUT:
  β†’ Support Responder: Update status page every 15 minutes
  β†’ Executive Summary Generator: Brief stakeholders (P0 only)
Step 4: Resolution Verification (Post-fix)
Evidence Collector:
1. Verify the fix resolves the issue
2. Screenshot evidence of working state
3. Confirm no new issues introduced

Infrastructure Maintainer:
1. Verify all metrics returning to normal
2. Confirm no cascading failures
3. Monitor for 30 minutes post-fix

API Tester (if API-related):
1. Run regression on affected endpoints
2. Verify response times normalized
3. Confirm error rates at baseline

Output: Incident resolved confirmation
Step 5: Post-Mortem (Within 48 hours)
Workflow Optimizer leads post-mortem:

1. Timeline reconstruction
   - When was the issue introduced?
   - When was it detected?
   - When was it resolved?
   - Total user impact duration

2. Root cause analysis
   - What failed?
   - Why did it fail?
   - Why wasn't it caught earlier?
   - 5 Whys analysis

3. Impact assessment
   - Users affected
   - Revenue impact
   - Reputation impact
   - Data impact

4. Prevention measures
   - What monitoring would have caught this sooner?
   - What testing would have prevented this?
   - What process changes are needed?
   - What infrastructure changes are needed?

5. Action items
   - [Action] β†’ [Owner] β†’ [Deadline]
   - [Action] β†’ [Owner] β†’ [Deadline]
   - [Action] β†’ [Owner] β†’ [Deadline]

Output: Post-Mortem Report β†’ Sprint Prioritizer adds prevention tasks to backlog

Communication Templates

Status Page Update (Support Responder)
[TIMESTAMP] β€” [SERVICE NAME] Incident

Status: [Investigating / Identified / Monitoring / Resolved]
Impact: [Description of user impact]
Current action: [What we're doing about it]
Next update: [When to expect the next update]
Executive Update (Executive Summary Generator β€” P0 only)
INCIDENT BRIEF β€” [TIMESTAMP]

SITUATION: [Service] is [down/degraded] affecting [N users/% of traffic]
CAUSE: [Known/Under investigation] β€” [Brief description if known]
ACTION: [What's being done] β€” ETA [time estimate]
IMPACT: [Business impact β€” revenue, users, reputation]
NEXT UPDATE: [Timestamp]

Escalation Matrix

Condition Escalate To Action
P0 not resolved in 30 min Studio Producer Additional resources, vendor escalation
P1 not resolved in 2 hours Project Shepherd Resource reallocation
Data breach suspected Legal Compliance Checker Regulatory notification assessment
User data affected Legal Compliance Checker + Executive Summary Generator GDPR/CCPA notification
Revenue impact > $X Finance Tracker + Studio Producer Business impact assessment
1# 🚨 Runbook: Incident Response
2 
3> **Mode**: NEXUS-Micro | **Duration**: Minutes to hours | **Agents**: 3-8
4 
5---
6 
7## Scenario
8 
9Something is broken in production. Users are affected. Speed of response matters, but so does doing it right. This runbook covers detection through post-mortem.
10 
11## Severity Classification
12 
13| Level | Definition | Examples | Response Time |
14|-------|-----------|----------|--------------|
15| **P0 β€” Critical** | Service completely down, data loss, security breach | Database corruption, DDoS attack, auth system failure | Immediate (all hands) |
16| **P1 β€” High** | Major feature broken, significant performance degradation | Payment processing down, 50%+ error rate, 10x latency | < 1 hour |
17| **P2 β€” Medium** | Minor feature broken, workaround available | Search not working, non-critical API errors | < 4 hours |
18| **P3 β€” Low** | Cosmetic issue, minor inconvenience | Styling bug, typo, minor UI glitch | Next sprint |
19 
20## Response Teams by Severity
21 
22### P0 β€” Critical Response Team
23| Agent | Role | Action |
24|-------|------|--------|
25| **Infrastructure Maintainer** | Incident commander | Assess scope, coordinate response |
26| **DevOps Automator** | Deployment/rollback | Execute rollback if needed |
27| **Backend Architect** | Root cause investigation | Diagnose system issues |
28| **Frontend Developer** | UI-side investigation | Diagnose client-side issues |
29| **Support Responder** | User communication | Status page updates, user notifications |
30| **Executive Summary Generator** | Stakeholder communication | Real-time executive updates |
31 
32### P1 β€” High Response Team
33| Agent | Role |
34|-------|------|
35| **Infrastructure Maintainer** | Incident commander |
36| **DevOps Automator** | Deployment support |
37| **Relevant Developer Agent** | Fix implementation |
38| **Support Responder** | User communication |
39 
40### P2 β€” Medium Response
41| Agent | Role |
42|-------|------|
43| **Relevant Developer Agent** | Fix implementation |
44| **Evidence Collector** | Verify fix |
45 
46### P3 β€” Low Response
47| Agent | Role |
48|-------|------|
49| **Sprint Prioritizer** | Add to backlog |
50 
51## Incident Response Sequence
52 
53### Step 1: Detection & Triage (0-5 minutes)
54 
55```
56TRIGGER: Alert from monitoring / User report / Agent detection
57 
58Infrastructure Maintainer:
591. Acknowledge alert
602. Assess scope and impact
61 - How many users affected?
62 - Which services are impacted?
63 - Is data at risk?
643. Classify severity (P0/P1/P2/P3)
654. Activate appropriate response team
665. Create incident channel/thread
67 
68Output: Incident classification + response team activated
69```
70 
71### Step 2: Investigation (5-30 minutes)
72 
73```
74PARALLEL INVESTIGATION:
75 
76Infrastructure Maintainer:
77β”œβ”€β”€ Check system metrics (CPU, memory, network, disk)
78β”œβ”€β”€ Review error logs
79β”œβ”€β”€ Check recent deployments
80└── Verify external dependencies
81 
82Backend Architect (if P0/P1):
83β”œβ”€β”€ Check database health
84β”œβ”€β”€ Review API error rates
85β”œβ”€β”€ Check service communication
86└── Identify failing component
87 
88DevOps Automator:
89β”œβ”€β”€ Review recent deployment history
90β”œβ”€β”€ Check CI/CD pipeline status
91β”œβ”€β”€ Prepare rollback if needed
92└── Verify infrastructure state
93 
94Output: Root cause identified (or narrowed to component)
95```
96 
97### Step 3: Mitigation (15-60 minutes)
98 
99```
100DECISION TREE:
101 
102IF caused by recent deployment:
103 β†’ DevOps Automator: Execute rollback
104 β†’ Infrastructure Maintainer: Verify recovery
105 β†’ Evidence Collector: Confirm fix
106 
107IF caused by infrastructure issue:
108 β†’ Infrastructure Maintainer: Scale/restart/failover
109 β†’ DevOps Automator: Support infrastructure changes
110 β†’ Verify recovery
111 
112IF caused by code bug:
113 β†’ Relevant Developer Agent: Implement hotfix
114 β†’ Evidence Collector: Verify fix
115 β†’ DevOps Automator: Deploy hotfix
116 β†’ Infrastructure Maintainer: Monitor recovery
117 
118IF caused by external dependency:
119 β†’ Infrastructure Maintainer: Activate fallback/cache
120 β†’ Support Responder: Communicate to users
121 β†’ Monitor for external recovery
122 
123THROUGHOUT:
124 β†’ Support Responder: Update status page every 15 minutes
125 β†’ Executive Summary Generator: Brief stakeholders (P0 only)
126```
127 
128### Step 4: Resolution Verification (Post-fix)
129 
130```
131Evidence Collector:
1321. Verify the fix resolves the issue
1332. Screenshot evidence of working state
1343. Confirm no new issues introduced
135 
136Infrastructure Maintainer:
1371. Verify all metrics returning to normal
1382. Confirm no cascading failures
1393. Monitor for 30 minutes post-fix
140 
141API Tester (if API-related):
1421. Run regression on affected endpoints
1432. Verify response times normalized
1443. Confirm error rates at baseline
145 
146Output: Incident resolved confirmation
147```
148 
149### Step 5: Post-Mortem (Within 48 hours)
150 
151```
152Workflow Optimizer leads post-mortem:
153 
1541. Timeline reconstruction
155 - When was the issue introduced?
156 - When was it detected?
157 - When was it resolved?
158 - Total user impact duration
159 
1602. Root cause analysis
161 - What failed?
162 - Why did it fail?
163 - Why wasn't it caught earlier?
164 - 5 Whys analysis
165 
1663. Impact assessment
167 - Users affected
168 - Revenue impact
169 - Reputation impact
170 - Data impact
171 
1724. Prevention measures
173 - What monitoring would have caught this sooner?
174 - What testing would have prevented this?
175 - What process changes are needed?
176 - What infrastructure changes are needed?
177 
1785. Action items
179 - [Action] β†’ [Owner] β†’ [Deadline]
180 - [Action] β†’ [Owner] β†’ [Deadline]
181 - [Action] β†’ [Owner] β†’ [Deadline]
182 
183Output: Post-Mortem Report β†’ Sprint Prioritizer adds prevention tasks to backlog
184```
185 
186## Communication Templates
187 
188### Status Page Update (Support Responder)
189```
190[TIMESTAMP] β€” [SERVICE NAME] Incident
191 
192Status: [Investigating / Identified / Monitoring / Resolved]
193Impact: [Description of user impact]
194Current action: [What we're doing about it]
195Next update: [When to expect the next update]
196```
197 
198### Executive Update (Executive Summary Generator β€” P0 only)
199```
200INCIDENT BRIEF β€” [TIMESTAMP]
201 
202SITUATION: [Service] is [down/degraded] affecting [N users/% of traffic]
203CAUSE: [Known/Under investigation] β€” [Brief description if known]
204ACTION: [What's being done] β€” ETA [time estimate]
205IMPACT: [Business impact β€” revenue, users, reputation]
206NEXT UPDATE: [Timestamp]
207```
208 
209## Escalation Matrix
210 
211| Condition | Escalate To | Action |
212|-----------|------------|--------|
213| P0 not resolved in 30 min | Studio Producer | Additional resources, vendor escalation |
214| P1 not resolved in 2 hours | Project Shepherd | Resource reallocation |
215| Data breach suspected | Legal Compliance Checker | Regulatory notification assessment |
216| User data affected | Legal Compliance Checker + Executive Summary Generator | GDPR/CCPA notification |
217| Revenue impact > $X | Finance Tracker + Studio Producer | Business impact assessment |
218 

Discussion

Alternatives