Runbook Writer Skill

Write an operational runbook for a service, incident type, or deployment procedure.

Runbook Writer Skill — The Skill Playground: pick the Executive Update skill, fill in a few notes, hit run, and watch a structured executive… (from the mohitagw15856/pm-claude-skills README)

From the mohitagw15856/pm-claude-skills README — shows the whole collection, not only this skill. · view on GitHub

How to use it

Claude Code
  1. Run the line below. It pulls the whole folder into ~/.claude/skills/runbook-writer, including the files SKILL.md points to.
  2. Describe your job in plain words. Claude Code follows the skill from there.
Claude Code — installs the whole folder, not just SKILL.md
npx degit mohitagw15856/pm-claude-skills/skills/runbook-writer#main ~/.claude/skills/runbook-writer

For one project only, change the path to .claude/skills/runbook-writer. This skill also uses Node.js — copying SKILL.md alone won't be enough. See the folder on GitHub.

Claude (web or desktop app)
  1. On this page open ⋯ → Download .md.
  2. Save it as SKILL.md in a folder, zip the folder, then Customize → Skills → + → Create skill → Upload a skill.
  3. Pick the file and Save. Claude shows the name and description and runs a security scan.
  4. Check the skill is switched on.
  5. Start a new chat and describe your job in plain words. The AI follows the skill from there.
ChatGPT or another app
  1. ChatGPT: make a Project and paste it into Instructions.
  2. Neither? Paste it at the top of a new chat — it works for that chat.
Not working?
  • Check which app you pasted it into — the steps above name the right one.
  • Some skills need the paid tier of Claude or ChatGPT.
Step-by-step guide with screenshots · Ask in the forum

Paste into Claude, ChatGPT or Cursor.

Source of Runbook Writer Skill

Show the full text174 lines
namedescription
runbook-writerWrite an operational runbook for a service, incident type, or deployment procedure. Use when asked to write a runbook, create an ops guide, document an operational procedure, or prepare an incident response playbook. Produces a runbook with overview, prerequisites, step-by-step procedures, rollback steps, troubleshooting table, and escalation paths.

Runbook Writer Skill

Produces operational runbooks for services, incident types, and deployment procedures — structured so an on-call engineer who's never touched the system can follow them under pressure.

Required Inputs

Ask for these if not provided:

  • What the runbook is for (e.g. deploying the payment service, responding to a database failover, rotating API keys)
  • Runbook type (Deployment / Incident Response / Maintenance / Disaster Recovery)
  • System/service name and what it does (brief description)
  • Audience (new on-call engineers / experienced SREs / DevOps team)
  • Tech stack (where relevant — e.g. Kubernetes, AWS RDS, Node.js)
  • Monitoring tools (e.g. Grafana, Datadog, CloudWatch, Splunk — used to name specific dashboards and alert links in the steps)
  • Key environment details (e.g. Kubernetes cluster name, AWS account/region, relevant namespaces or resource names — paste what's relevant for exact commands)

Output Format


Runbook: [Runbook Title] Service: [Service Name] Type: [Deployment / Incident Response / Maintenance / DR] Last Updated: [Insert today's date in YYYY-MM-DD format] Owner: [Team or person] Severity: [P1 / P2 / P3 — if incident-type]


Overview

What this runbook covers: [1–2 sentences on the scenario this runbook handles]

When to use this runbook:

  • [Specific trigger condition 1 — e.g. PagerDuty alert: high-error-rate-payment-service]
  • [Specific trigger condition 2 — e.g. Deploy needed after PR merged to main]

Estimated time to complete: [X minutes / X–Y minutes depending on outcome]

Impact if not completed correctly: [e.g. Payment processing degraded / Data loss risk / Users locked out]


Prerequisites

Access required:

  • [System/tool access — e.g. AWS Console: production-account]
  • [Credential — e.g. vault read secret/payment-service]
  • [VPN / bastion access if needed]

Tools required:

  • [Tool name and version — e.g. kubectl v1.28+]
  • [CLI or dashboard name]

Before you start:

  • [Prerequisite check — e.g. Verify current deployment is healthy in Grafana]
  • [Prerequisite action — e.g. Announce in #ops-live that you're starting]

Procedure

Number every step. Use exact commands. Do not paraphrase tool names or flags.

Step 1: [Action name] [What you're doing and why — one sentence]

# Exact command
[command here]

Expected output: [what should appear if this worked] If this fails: [Exact error message to look for] → [What to do, or see Troubleshooting]

Step 2: [Action name] [Same structure as Step 1]

Step 3: Verify Always include a verification step after the main procedure:

[verification command]

Expected state: [What a healthy system looks like after this runbook completes]


Rollback

How to undo this procedure if something went wrong:

Step R1: [Rollback action]

[rollback command]

Verify rollback: [command to confirm rollback succeeded]


Troubleshooting
Symptom Likely Cause Resolution
[Error message or observable symptom] [Why this happens] [Exact fix or next step]
[Another symptom] [Cause] [Resolution]

Escalation

If this runbook does not resolve the issue:

Condition Who to Contact How
[e.g. DB unavailable after 10 min] [DBA on-call] [PagerDuty policy: db-oncall]
[e.g. Payment provider unresponsive] [Vendor contact] [Contact in 1Password: vendor-escalation]

Always update the incident timeline in [tool] before escalating.


Post-Procedure Checklist

After completing the runbook:

  • Announce completion in #ops-live with outcome
  • Update the incident ticket / deploy log
  • Verify alerts have resolved in monitoring dashboard
  • If this revealed a gap in this runbook — update it now (link to edit process)

Deeper Materials

This skill ships with support files — use them when they are available:

  • references/3am-usability.md — The 3AM Test: Runbooks for Degraded Humans. Apply it while producing the output; it carries the calibration and judgment calls the method summary above compresses.
  • templates/runbook.md — a fill-in version of the deliverable with the quality gates inline. Offer it when the user wants to work the document themselves rather than have it generated.

Scoring Rubric (0–40)

Score any output of this skill before handing it over; 32+ is ship-quality.

Dimension 0 5 10
Command exactness Steps are vague actions ("run the deploy script") with no commands Most steps have commands, but some paraphrase tool names or omit flags Every step has an exact, copy-pasteable command with correct flags and placeholders clearly marked
Success & failure signalling No expected outputs; reader cannot tell whether a step worked Expected output on the happy path only; failure paths say "investigate" Every step states expected output and a named failure path or troubleshooting link
Rollback completeness Rollback missing or left as a placeholder Rollback commands exist but are partial and have no verification Rollback is complete, independently testable, and each rollback step has its own verify command
Cold-reader usability Assumes system knowledge; unexplained jargon; escalation cells like "[Team name]" Followable by a team member, but tribal-knowledge gaps would stall an outsider An engineer who has never touched the system can execute it under pressure; every escalation row has a real contact or an explicit [FILL IN] flag

Quality Checks

  • Every step has an exact command (no "run the deploy script")
  • Expected output is specified for each step so engineer knows if it worked
  • Failure path is explicit for each step (not "if it fails, investigate")
  • Rollback procedure is complete and independently testable
  • Escalation table has no cells containing only "[Team name]" — every row must either have a real contact or be explicitly flagged as [FILL IN: on-call rotation link]
  • Rollback section contains at least one concrete command (not left as "[rollback command]" placeholder)
  • Runbook can be followed by someone who has never touched this system

Usage Examples

  • "Write a runbook for [service] deployment"
  • "Create an incident response runbook for [alert type]"
  • "I need a runbook for [procedure]"
  • "Document the operational procedure for [X]"
  • "Write an ops playbook for [scenario]"

Anti-Patterns

  • Do not write steps as vague actions like "run the deploy script" — every step must include the exact command
  • Do not leave the rollback section as a placeholder — a runbook without a tested rollback procedure is incomplete and dangerous
  • Do not omit expected output for each step — without it, the on-call engineer cannot tell if the step succeeded
  • Do not write escalation contacts as "[Team name]" — every escalation row must have a real contact or an explicit flag to fill in
  • Do not assume the reader knows the system — write for someone who has never touched it before
1---
2name: runbook-writer
3description: "Write an operational runbook for a service, incident type, or deployment procedure. Use when asked to write a runbook, create an ops guide, document an operational procedure, or prepare an incident response playbook. Produces a runbook with overview, prerequisites, step-by-step procedures, rollback steps, troubleshooting table, and escalation paths."
4---
5 
6# Runbook Writer Skill
7 
8Produces operational runbooks for services, incident types, and deployment procedures — structured so an on-call engineer who's never touched the system can follow them under pressure.
9 
10## Required Inputs
11 
12Ask for these if not provided:
13- **What the runbook is for** (e.g. deploying the payment service, responding to a database failover, rotating API keys)
14- **Runbook type** (Deployment / Incident Response / Maintenance / Disaster Recovery)
15- **System/service name and what it does** (brief description)
16- **Audience** (new on-call engineers / experienced SREs / DevOps team)
17- **Tech stack** (where relevant — e.g. Kubernetes, AWS RDS, Node.js)
18- **Monitoring tools** (e.g. Grafana, Datadog, CloudWatch, Splunk — used to name specific dashboards and alert links in the steps)
19- **Key environment details** (e.g. Kubernetes cluster name, AWS account/region, relevant namespaces or resource names — paste what's relevant for exact commands)
20 
21## Output Format
22 
23---
24**Runbook:** [Runbook Title]
25**Service:** [Service Name]
26**Type:** [Deployment / Incident Response / Maintenance / DR]
27**Last Updated:** [Insert today's date in YYYY-MM-DD format]
28**Owner:** [Team or person]
29**Severity:** [P1 / P2 / P3 — if incident-type]
30 
31---
32 
33### Overview
34**What this runbook covers:**
35[1–2 sentences on the scenario this runbook handles]
36 
37**When to use this runbook:**
38- [Specific trigger condition 1 — e.g. PagerDuty alert: `high-error-rate-payment-service`]
39- [Specific trigger condition 2 — e.g. Deploy needed after PR merged to `main`]
40 
41**Estimated time to complete:** [X minutes / X–Y minutes depending on outcome]
42 
43**Impact if not completed correctly:** [e.g. Payment processing degraded / Data loss risk / Users locked out]
44 
45---
46 
47### Prerequisites
48 
49**Access required:**
50- [ ] [System/tool access — e.g. AWS Console: `production-account`]
51- [ ] [Credential — e.g. `vault read secret/payment-service`]
52- [ ] [VPN / bastion access if needed]
53 
54**Tools required:**
55- [ ] [Tool name and version — e.g. `kubectl` v1.28+]
56- [ ] [CLI or dashboard name]
57 
58**Before you start:**
59- [ ] [Prerequisite check — e.g. Verify current deployment is healthy in Grafana]
60- [ ] [Prerequisite action — e.g. Announce in `#ops-live` that you're starting]
61 
62---
63 
64### Procedure
65 
66Number every step. Use exact commands. Do not paraphrase tool names or flags.
67 
68**Step 1: [Action name]**
69[What you're doing and why — one sentence]
70```bash
71# Exact command
72[command here]
73```
74**Expected output:** `[what should appear if this worked]`
75**If this fails:** [Exact error message to look for] → [What to do, or see Troubleshooting]
76 
77**Step 2: [Action name]**
78[Same structure as Step 1]
79 
80**Step 3: Verify**
81Always include a verification step after the main procedure:
82```bash
83[verification command]
84```
85**Expected state:** [What a healthy system looks like after this runbook completes]
86 
87---
88 
89### Rollback
90 
91How to undo this procedure if something went wrong:
92 
93**Step R1: [Rollback action]**
94```bash
95[rollback command]
96```
97**Verify rollback:** `[command to confirm rollback succeeded]`
98 
99---
100 
101### Troubleshooting
102 
103| Symptom | Likely Cause | Resolution |
104|---|---|---|
105| [Error message or observable symptom] | [Why this happens] | [Exact fix or next step] |
106| [Another symptom] | [Cause] | [Resolution] |
107 
108---
109 
110### Escalation
111 
112If this runbook does not resolve the issue:
113 
114| Condition | Who to Contact | How |
115|---|---|---|
116| [e.g. DB unavailable after 10 min] | [DBA on-call] | [PagerDuty policy: `db-oncall`] |
117| [e.g. Payment provider unresponsive] | [Vendor contact] | [Contact in 1Password: `vendor-escalation`] |
118 
119**Always update the incident timeline in [tool] before escalating.**
120 
121---
122 
123### Post-Procedure Checklist
124 
125After completing the runbook:
126- [ ] Announce completion in `#ops-live` with outcome
127- [ ] Update the incident ticket / deploy log
128- [ ] Verify alerts have resolved in monitoring dashboard
129- [ ] If this revealed a gap in this runbook — update it now (link to edit process)
130 
131---
132 
133## Deeper Materials
134 
135This skill ships with support files — use them when they are available:
136 
137- **`references/3am-usability.md`** — The 3AM Test: Runbooks for Degraded Humans. Apply it while producing the output; it carries the calibration and judgment calls the method summary above compresses.
138- **`templates/runbook.md`** — a fill-in version of the deliverable with the quality gates inline. Offer it when the user wants to work the document themselves rather than have it generated.
139 
140## Scoring Rubric (0–40)
141 
142Score any output of this skill before handing it over; 32+ is ship-quality.
143 
144| Dimension | 0 | 5 | 10 |
145|---|---|---|---|
146| **Command exactness** | Steps are vague actions ("run the deploy script") with no commands | Most steps have commands, but some paraphrase tool names or omit flags | Every step has an exact, copy-pasteable command with correct flags and placeholders clearly marked |
147| **Success & failure signalling** | No expected outputs; reader cannot tell whether a step worked | Expected output on the happy path only; failure paths say "investigate" | Every step states expected output and a named failure path or troubleshooting link |
148| **Rollback completeness** | Rollback missing or left as a placeholder | Rollback commands exist but are partial and have no verification | Rollback is complete, independently testable, and each rollback step has its own verify command |
149| **Cold-reader usability** | Assumes system knowledge; unexplained jargon; escalation cells like "[Team name]" | Followable by a team member, but tribal-knowledge gaps would stall an outsider | An engineer who has never touched the system can execute it under pressure; every escalation row has a real contact or an explicit [FILL IN] flag |
150 
151## Quality Checks
152- [ ] Every step has an exact command (no "run the deploy script")
153- [ ] Expected output is specified for each step so engineer knows if it worked
154- [ ] Failure path is explicit for each step (not "if it fails, investigate")
155- [ ] Rollback procedure is complete and independently testable
156- [ ] Escalation table has no cells containing only "[Team name]" — every row must either have a real contact or be explicitly flagged as [FILL IN: on-call rotation link]
157- [ ] Rollback section contains at least one concrete command (not left as "[rollback command]" placeholder)
158- [ ] Runbook can be followed by someone who has never touched this system
159 
160## Usage Examples
161- "Write a runbook for [service] deployment"
162- "Create an incident response runbook for [alert type]"
163- "I need a runbook for [procedure]"
164- "Document the operational procedure for [X]"
165- "Write an ops playbook for [scenario]"
166 
167## Anti-Patterns
168 
169- [ ] Do not write steps as vague actions like "run the deploy script" — every step must include the exact command
170- [ ] Do not leave the rollback section as a placeholder — a runbook without a tested rollback procedure is incomplete and dangerous
171- [ ] Do not omit expected output for each step — without it, the on-call engineer cannot tell if the step succeeded
172- [ ] Do not write escalation contacts as "[Team name]" — every escalation row must have a real contact or an explicit flag to fill in
173- [ ] Do not assume the reader knows the system — write for someone who has never touched it before
174 

Discussion

Alternatives

Also in Process docsSee all 58 in Operations →
On call handoff patternsMaster on-call shift handoffs with context transfer, escalation procedures, and documentation. Use this skill when transitioning on-call responsibilities between engineers and ensuring the incoming responder has full situational awareness, when writing a shift summary that captures active incidents, ongoing investigations, and recent changes, when handing off mid-incident so a fresh engineer can take over the incident commander role without losing context, when onboarding a new engineer to the on-call rotation for the first time, or when auditing and improving the quality of existing handoff processes across teams.Infrastructure & ops · MITHandoffCompact the current conversation into a handoff document for another agent to pick up. Save to a user-configured location (OS temp, home folder, or per-project .handoff/), redact secrets before write, suggest skills for the next session, and auto-load the latest handoff on the next SessionStart. First-run setup asks where to save so the project folder never gets cluttered. Use when the user says 'hand this off', 'handoff doc', 'summarize this for a new session', 'compact this conversation', 'I'm ending this session', 'pick this up later', or any variation signaling intent to pass work to a fresh agent. Also trigger on implicit signals: the user announcing they're switching machines, ending the day mid-task, or context is growing long without a natural stopping point.Business & ops · MITFinancial Due Diligence SkillGenerate a financial due diligence checklist and analysis framework for any investment, acquisition, or partnership. Use when asked for a due diligence checklist, M&A financial review, investment analysis framework, or vendor financial assessment. Produces a document request list, key analytical questions, red flags checklist, and a summarised financial health assessment.Business & ops · MITOffsite Planner SkillPlan a team offsite that earns its cost — the purpose split (connection vs. decisions vs. planning, weighted on purpose), the agenda that alternates work and air, the logistics runbook, and the follow-through that makes Monday different from before. Use when asked plan our team offsite, design two days for the team, make this offsite not a waste, or what do we actually do at the offsite. Produces the purpose weighting, the day designs, the logistics checklist, and the commitments-capture that survives re-entry.Business & ops · MIT