| 1 | # π¨ Runbook: Incident Response |
| 2 | |
| 3 | > **Mode**: NEXUS-Micro | **Duration**: Minutes to hours | **Agents**: 3-8 |
| 4 | |
| 5 | --- |
| 6 | |
| 7 | ## Scenario |
| 8 | |
| 9 | Something is broken in production. Users are affected. Speed of response matters, but so does doing it right. This runbook covers detection through post-mortem. |
| 10 | |
| 11 | ## Severity Classification |
| 12 | |
| 13 | | Level | Definition | Examples | Response Time | |
| 14 | |-------|-----------|----------|--------------| |
| 15 | | **P0 β Critical** | Service completely down, data loss, security breach | Database corruption, DDoS attack, auth system failure | Immediate (all hands) | |
| 16 | | **P1 β High** | Major feature broken, significant performance degradation | Payment processing down, 50%+ error rate, 10x latency | < 1 hour | |
| 17 | | **P2 β Medium** | Minor feature broken, workaround available | Search not working, non-critical API errors | < 4 hours | |
| 18 | | **P3 β Low** | Cosmetic issue, minor inconvenience | Styling bug, typo, minor UI glitch | Next sprint | |
| 19 | |
| 20 | ## Response Teams by Severity |
| 21 | |
| 22 | ### P0 β Critical Response Team |
| 23 | | Agent | Role | Action | |
| 24 | |-------|------|--------| |
| 25 | | **Infrastructure Maintainer** | Incident commander | Assess scope, coordinate response | |
| 26 | | **DevOps Automator** | Deployment/rollback | Execute rollback if needed | |
| 27 | | **Backend Architect** | Root cause investigation | Diagnose system issues | |
| 28 | | **Frontend Developer** | UI-side investigation | Diagnose client-side issues | |
| 29 | | **Support Responder** | User communication | Status page updates, user notifications | |
| 30 | | **Executive Summary Generator** | Stakeholder communication | Real-time executive updates | |
| 31 | |
| 32 | ### P1 β High Response Team |
| 33 | | Agent | Role | |
| 34 | |-------|------| |
| 35 | | **Infrastructure Maintainer** | Incident commander | |
| 36 | | **DevOps Automator** | Deployment support | |
| 37 | | **Relevant Developer Agent** | Fix implementation | |
| 38 | | **Support Responder** | User communication | |
| 39 | |
| 40 | ### P2 β Medium Response |
| 41 | | Agent | Role | |
| 42 | |-------|------| |
| 43 | | **Relevant Developer Agent** | Fix implementation | |
| 44 | | **Evidence Collector** | Verify fix | |
| 45 | |
| 46 | ### P3 β Low Response |
| 47 | | Agent | Role | |
| 48 | |-------|------| |
| 49 | | **Sprint Prioritizer** | Add to backlog | |
| 50 | |
| 51 | ## Incident Response Sequence |
| 52 | |
| 53 | ### Step 1: Detection & Triage (0-5 minutes) |
| 54 | |
| 55 | ``` |
| 56 | TRIGGER: Alert from monitoring / User report / Agent detection |
| 57 | |
| 58 | Infrastructure Maintainer: |
| 59 | 1. Acknowledge alert |
| 60 | 2. Assess scope and impact |
| 61 | - How many users affected? |
| 62 | - Which services are impacted? |
| 63 | - Is data at risk? |
| 64 | 3. Classify severity (P0/P1/P2/P3) |
| 65 | 4. Activate appropriate response team |
| 66 | 5. Create incident channel/thread |
| 67 | |
| 68 | Output: Incident classification + response team activated |
| 69 | ``` |
| 70 | |
| 71 | ### Step 2: Investigation (5-30 minutes) |
| 72 | |
| 73 | ``` |
| 74 | PARALLEL INVESTIGATION: |
| 75 | |
| 76 | Infrastructure Maintainer: |
| 77 | βββ Check system metrics (CPU, memory, network, disk) |
| 78 | βββ Review error logs |
| 79 | βββ Check recent deployments |
| 80 | βββ Verify external dependencies |
| 81 | |
| 82 | Backend Architect (if P0/P1): |
| 83 | βββ Check database health |
| 84 | βββ Review API error rates |
| 85 | βββ Check service communication |
| 86 | βββ Identify failing component |
| 87 | |
| 88 | DevOps Automator: |
| 89 | βββ Review recent deployment history |
| 90 | βββ Check CI/CD pipeline status |
| 91 | βββ Prepare rollback if needed |
| 92 | βββ Verify infrastructure state |
| 93 | |
| 94 | Output: Root cause identified (or narrowed to component) |
| 95 | ``` |
| 96 | |
| 97 | ### Step 3: Mitigation (15-60 minutes) |
| 98 | |
| 99 | ``` |
| 100 | DECISION TREE: |
| 101 | |
| 102 | IF caused by recent deployment: |
| 103 | β DevOps Automator: Execute rollback |
| 104 | β Infrastructure Maintainer: Verify recovery |
| 105 | β Evidence Collector: Confirm fix |
| 106 | |
| 107 | IF caused by infrastructure issue: |
| 108 | β Infrastructure Maintainer: Scale/restart/failover |
| 109 | β DevOps Automator: Support infrastructure changes |
| 110 | β Verify recovery |
| 111 | |
| 112 | IF caused by code bug: |
| 113 | β Relevant Developer Agent: Implement hotfix |
| 114 | β Evidence Collector: Verify fix |
| 115 | β DevOps Automator: Deploy hotfix |
| 116 | β Infrastructure Maintainer: Monitor recovery |
| 117 | |
| 118 | IF caused by external dependency: |
| 119 | β Infrastructure Maintainer: Activate fallback/cache |
| 120 | β Support Responder: Communicate to users |
| 121 | β Monitor for external recovery |
| 122 | |
| 123 | THROUGHOUT: |
| 124 | β Support Responder: Update status page every 15 minutes |
| 125 | β Executive Summary Generator: Brief stakeholders (P0 only) |
| 126 | ``` |
| 127 | |
| 128 | ### Step 4: Resolution Verification (Post-fix) |
| 129 | |
| 130 | ``` |
| 131 | Evidence Collector: |
| 132 | 1. Verify the fix resolves the issue |
| 133 | 2. Screenshot evidence of working state |
| 134 | 3. Confirm no new issues introduced |
| 135 | |
| 136 | Infrastructure Maintainer: |
| 137 | 1. Verify all metrics returning to normal |
| 138 | 2. Confirm no cascading failures |
| 139 | 3. Monitor for 30 minutes post-fix |
| 140 | |
| 141 | API Tester (if API-related): |
| 142 | 1. Run regression on affected endpoints |
| 143 | 2. Verify response times normalized |
| 144 | 3. Confirm error rates at baseline |
| 145 | |
| 146 | Output: Incident resolved confirmation |
| 147 | ``` |
| 148 | |
| 149 | ### Step 5: Post-Mortem (Within 48 hours) |
| 150 | |
| 151 | ``` |
| 152 | Workflow Optimizer leads post-mortem: |
| 153 | |
| 154 | 1. Timeline reconstruction |
| 155 | - When was the issue introduced? |
| 156 | - When was it detected? |
| 157 | - When was it resolved? |
| 158 | - Total user impact duration |
| 159 | |
| 160 | 2. Root cause analysis |
| 161 | - What failed? |
| 162 | - Why did it fail? |
| 163 | - Why wasn't it caught earlier? |
| 164 | - 5 Whys analysis |
| 165 | |
| 166 | 3. Impact assessment |
| 167 | - Users affected |
| 168 | - Revenue impact |
| 169 | - Reputation impact |
| 170 | - Data impact |
| 171 | |
| 172 | 4. Prevention measures |
| 173 | - What monitoring would have caught this sooner? |
| 174 | - What testing would have prevented this? |
| 175 | - What process changes are needed? |
| 176 | - What infrastructure changes are needed? |
| 177 | |
| 178 | 5. Action items |
| 179 | - [Action] β [Owner] β [Deadline] |
| 180 | - [Action] β [Owner] β [Deadline] |
| 181 | - [Action] β [Owner] β [Deadline] |
| 182 | |
| 183 | Output: Post-Mortem Report β Sprint Prioritizer adds prevention tasks to backlog |
| 184 | ``` |
| 185 | |
| 186 | ## Communication Templates |
| 187 | |
| 188 | ### Status Page Update (Support Responder) |
| 189 | ``` |
| 190 | [TIMESTAMP] β [SERVICE NAME] Incident |
| 191 | |
| 192 | Status: [Investigating / Identified / Monitoring / Resolved] |
| 193 | Impact: [Description of user impact] |
| 194 | Current action: [What we're doing about it] |
| 195 | Next update: [When to expect the next update] |
| 196 | ``` |
| 197 | |
| 198 | ### Executive Update (Executive Summary Generator β P0 only) |
| 199 | ``` |
| 200 | INCIDENT BRIEF β [TIMESTAMP] |
| 201 | |
| 202 | SITUATION: [Service] is [down/degraded] affecting [N users/% of traffic] |
| 203 | CAUSE: [Known/Under investigation] β [Brief description if known] |
| 204 | ACTION: [What's being done] β ETA [time estimate] |
| 205 | IMPACT: [Business impact β revenue, users, reputation] |
| 206 | NEXT UPDATE: [Timestamp] |
| 207 | ``` |
| 208 | |
| 209 | ## Escalation Matrix |
| 210 | |
| 211 | | Condition | Escalate To | Action | |
| 212 | |-----------|------------|--------| |
| 213 | | P0 not resolved in 30 min | Studio Producer | Additional resources, vendor escalation | |
| 214 | | P1 not resolved in 2 hours | Project Shepherd | Resource reallocation | |
| 215 | | Data breach suspected | Legal Compliance Checker | Regulatory notification assessment | |
| 216 | | User data affected | Legal Compliance Checker + Executive Summary Generator | GDPR/CCPA notification | |
| 217 | | Revenue impact > $X | Finance Tracker + Studio Producer | Business impact assessment | |
| 218 | |