Shipping and launch

Prepares production launches.

How to use it

Claude Code
  1. Run the line below. It pulls the whole folder into ~/.claude/skills/shipping-and-launch.
  2. Describe your job in plain words. Claude Code follows the skill from there.
Claude Code — installs the whole folder, not just SKILL.md
npx degit addyosmani/agent-skills/skills/shipping-and-launch#main ~/.claude/skills/shipping-and-launch

For one project only, change the path to .claude/skills/shipping-and-launch.

Claude (web or desktop app)
  1. On this page open ⋯ → Download .md.
  2. Save it as SKILL.md in a folder, zip the folder, then Customize → Skills → + → Create skill → Upload a skill.
  3. Pick the file and Save. Claude shows the name and description and runs a security scan.
  4. Check the skill is switched on.
  5. Start a new chat and describe your job in plain words. The AI follows the skill from there.
ChatGPT or another app
  1. ChatGPT: make a Project and paste it into Instructions.
  2. Neither? Paste it at the top of a new chat — it works for that chat.
Not working?
  • Check which app you pasted it into — the steps above name the right one.
  • Some skills need the paid tier of Claude or ChatGPT.
Step-by-step guide with screenshots · Ask in the forum

Paste into Claude, ChatGPT or Cursor.

Source of Shipping and launch

Show the full text331 lines
namedescription
shipping-and-launchPrepares production launches. Use when preparing to deploy to production, or when asking what needs to be in place before shipping. Use when you need a pre-launch checklist, when setting up monitoring, when planning a staged rollout, or when you need a rollback strategy.

Shipping and Launch

Overview

Ship with confidence. The goal is not just to deploy — it's to deploy safely, with monitoring in place, a rollback plan ready, and a clear understanding of what success looks like. Every launch should be reversible, observable, and incremental.

When to Use

  • Deploying a feature to production for the first time
  • Releasing a significant change to users
  • Migrating data or infrastructure
  • Opening a beta or early access program
  • Any deployment that carries risk (all of them)

The Pre-Launch Checklist

Code Quality
  • All tests pass (unit, integration, e2e)
  • Build succeeds with no warnings
  • Lint and type checking pass
  • Code reviewed and approved
  • No TODO comments that should be resolved before launch
  • No console.log debugging statements in production code
  • Error handling covers expected failure modes
Security
  • No secrets in code or version control
  • The ecosystem's dependency audit (npm audit, pip-audit, cargo audit, ...) shows no critical or high vulnerabilities
  • Input validation on all user-facing endpoints
  • Authentication and authorization checks in place
  • Security headers configured (CSP, HSTS, etc.)
  • Rate limiting on authentication endpoints
  • CORS configured to specific origins (not wildcard)
Performance
  • Core Web Vitals within "Good" thresholds
  • No N+1 queries in critical paths
  • Images optimized (compression, responsive sizes, lazy loading)
  • Bundle size within budget
  • Database queries have appropriate indexes
  • Caching configured for static assets and repeated queries
Accessibility
  • Keyboard navigation works for all interactive elements
  • Screen reader can convey page content and structure
  • Color contrast meets WCAG 2.1 AA (4.5:1 for text)
  • Focus management correct for modals and dynamic content
  • Error messages are descriptive and associated with form fields
  • No accessibility warnings in axe-core or Lighthouse
Infrastructure
  • Environment variables set in production
  • Database migrations applied (or ready to apply)
  • DNS and SSL configured
  • CDN configured for static assets
  • Logging and error reporting configured
  • Health check endpoint exists and responds
Documentation
  • README updated with any new setup requirements
  • API documentation current
  • ADRs written for any architectural decisions
  • Changelog updated
  • User-facing documentation updated (if applicable)

Feature Flag Strategy

Ship behind feature flags to decouple deployment from release:

// Feature flag check
const flags = await getFeatureFlags(userId);

if (flags.taskSharing) {
  // New feature: task sharing
  return <TaskSharingPanel task={task} />;
}

// Default: existing behavior
return null;

Feature flag lifecycle:

1. DEPLOY with flag OFF     → Code is in production but inactive
2. ENABLE for team/beta     → Internal testing in production environment
3. GRADUAL ROLLOUT          → 5% → 25% → 50% → 100% of users
4. MONITOR at each stage    → Watch error rates, performance, user feedback
5. CLEAN UP                 → Remove flag and dead code path after full rollout

Rules:

  • Every feature flag has an owner and an expiration date
  • Clean up flags within 2 weeks of full rollout
  • Don't nest feature flags (creates exponential combinations)
  • Test both flag states (on and off) in CI

Staged Rollout

The Rollout Sequence
1. DEPLOY to staging
   └── Full test suite in staging environment
   └── Manual smoke test of critical flows

2. DEPLOY to production (feature flag OFF)
   └── Verify deployment succeeded (health check)
   └── Check error monitoring (no new errors)

3. ENABLE for team (flag ON for internal users)
   └── Team uses the feature in production
   └── 24-hour monitoring window

4. CANARY rollout (flag ON for 5% of users)
   └── Monitor error rates, latency, user behavior
   └── Compare metrics: canary vs. baseline
   └── 24-48 hour monitoring window
   └── Advance only if all thresholds pass (see table below)

5. GRADUAL increase (25% -> 50% -> 100%)
   └── Same monitoring at each step
   └── Ability to roll back to previous percentage at any point

6. FULL rollout (flag ON for all users)
   └── Monitor for 1 week
   └── Clean up feature flag
Rollout Decision Thresholds

Use these thresholds to decide whether to advance, hold, or roll back at each stage:

Metric Advance (green) Hold and investigate (yellow) Roll back (red)
Error rate Within 10% of baseline 10-100% above baseline >2x baseline
P95 latency Within 20% of baseline 20-50% above baseline >50% above baseline
Client JS errors No new error types New errors at <0.1% of sessions New errors at >0.1% of sessions
Business metrics Neutral or positive Decline <5% (may be noise) Decline >5%
When to Roll Back

Roll back immediately if:

  • Error rate increases by more than 2x baseline
  • P95 latency increases by more than 50%
  • User-reported issues spike
  • Data integrity issues detected
  • Security vulnerability discovered

Monitoring and Observability

What to Monitor
Application metrics:
├── Error rate (total and by endpoint)
├── Response time (p50, p95, p99)
├── Request volume
├── Active users
└── Key business metrics (conversion, engagement)

Infrastructure metrics:
├── CPU and memory utilization
├── Database connection pool usage
├── Disk space
├── Network latency
└── Queue depth (if applicable)

Client metrics:
├── Core Web Vitals (LCP, INP, CLS)
├── JavaScript errors
├── API error rates from client perspective
└── Page load time
Error Reporting
// Set up error boundary with reporting
class ErrorBoundary extends React.Component {
  componentDidCatch(error: Error, info: React.ErrorInfo) {
    // Report to error tracking service
    reportError(error, {
      componentStack: info.componentStack,
      userId: getCurrentUser()?.id,
      page: window.location.pathname,
    });
  }

  render() {
    if (this.state.hasError) {
      return <ErrorFallback onRetry={() => this.setState({ hasError: false })} />;
    }
    return this.props.children;
  }
}

// Server-side error reporting
app.use((err: Error, req: Request, res: Response, next: NextFunction) => {
  reportError(err, {
    method: req.method,
    url: req.url,
    userId: req.user?.id,
  });

  // Don't expose internals to users
  res.status(500).json({
    error: { code: 'INTERNAL_ERROR', message: 'Something went wrong' },
  });
});
Post-Launch Verification

In the first hour after launch:

1. Check health endpoint returns 200
2. Check error monitoring dashboard (no new error types)
3. Check latency dashboard (no regression)
4. Test the critical user flow manually
5. Verify logs are flowing and readable
6. Confirm rollback mechanism works (dry run if possible)

Error Budget Release Gate

Your service's error budget — the fraction of requests or time your SLO allows to fail — determines whether it's safe to ship. Use it as an objective gate — not a negotiation:

Budget remaining > 20%  →  Ship normally; monitor closely
Budget remaining 0–20%  →  Slow rollouts only; no high-risk changes
Budget exhausted        →  Freeze feature work; focus entirely on reliability
Budget resets           →  Resume normal pace; bake in the fix that recovered it

A high burn rate during a canary (consuming budget faster than the baseline pace) is a hold signal in the rollout thresholds table above — treat it the same as an elevated error rate.

Rollback Strategy

Every deployment needs a rollback plan before it happens:

## Rollback Plan for [Feature/Release]

### Trigger Conditions
- Error rate > 2x baseline
- P95 latency > [X]ms
- User reports of [specific issue]

### Rollback Steps
1. Disable feature flag (if applicable)
   OR
1. Deploy previous version: `git revert <commit> && git push`
2. Verify rollback: health check, error monitoring
3. Communicate: notify team of rollback

### Database Considerations
- Migration [X] has a rollback: `npx prisma migrate rollback`
- Data inserted by new feature: [preserved / cleaned up]

### Time to Rollback
- Feature flag: < 1 minute
- Redeploy previous version: < 5 minutes
- Database rollback: < 15 minutes

See Also

  • For the project-wide Definition of Done that every change must clear before this checklist, see ../../references/definition-of-done.md
  • For security pre-launch checks, see ../../references/security-checklist.md
  • For performance pre-launch checklist, see ../../references/performance-checklist.md
  • For accessibility verification before launch, see ../../references/accessibility-checklist.md
  • For the alerting rules and SLO-tied thresholds, see observability-and-instrumentation

Common Rationalizations

Rationalization Reality
"It works in staging, it'll work in production" Production has different data, traffic patterns, and edge cases. Monitor after deploy.
"We don't need feature flags for this" Every feature benefits from a kill switch. Even "simple" changes can break things.
"Monitoring is overhead" Not having monitoring means you discover problems from user complaints instead of dashboards.
"We'll add monitoring later" Add it before launch. You can't debug what you can't see.
"Rolling back is admitting failure" Rolling back is responsible engineering. Shipping a broken feature is the failure.
"The error rate looks fine, let's keep shipping" Check the burn rate, not just the current error rate. Consuming budget faster than baseline is a hold signal even when individual thresholds are green.

Red Flags

  • Deploying without a rollback plan
  • No monitoring or error reporting in production
  • Big-bang releases (everything at once, no staging)
  • Feature flags with no expiration or owner
  • No one monitoring the deploy for the first hour
  • Production environment configuration done by memory, not code
  • "It's Friday afternoon, let's ship it"
  • Error budget exhausted but feature work continues unchanged

Verification

Before deploying:

  • Pre-launch checklist completed (all sections green)
  • Feature flag configured (if applicable)
  • Rollback plan documented
  • Monitoring dashboards set up
  • Team notified of deployment

After deploying:

  • Health check returns 200
  • Error rate is normal
  • Latency is normal
  • Critical user flow works
  • Logs are flowing
  • Rollback tested or verified ready

For every shipped service:

  • Error budget policy in place: know what action to take when budget drops below 20% and when it's exhausted
1---
2name: shipping-and-launch
3description: Prepares production launches. Use when preparing to deploy to production, or when asking what needs to be in place before shipping. Use when you need a pre-launch checklist, when setting up monitoring, when planning a staged rollout, or when you need a rollback strategy.
4---
5 
6# Shipping and Launch
7 
8## Overview
9 
10Ship with confidence. The goal is not just to deploy — it's to deploy safely, with monitoring in place, a rollback plan ready, and a clear understanding of what success looks like. Every launch should be reversible, observable, and incremental.
11 
12## When to Use
13 
14- Deploying a feature to production for the first time
15- Releasing a significant change to users
16- Migrating data or infrastructure
17- Opening a beta or early access program
18- Any deployment that carries risk (all of them)
19 
20## The Pre-Launch Checklist
21 
22### Code Quality
23 
24- [ ] All tests pass (unit, integration, e2e)
25- [ ] Build succeeds with no warnings
26- [ ] Lint and type checking pass
27- [ ] Code reviewed and approved
28- [ ] No TODO comments that should be resolved before launch
29- [ ] No `console.log` debugging statements in production code
30- [ ] Error handling covers expected failure modes
31 
32### Security
33 
34- [ ] No secrets in code or version control
35- [ ] The ecosystem's dependency audit (`npm audit`, `pip-audit`, `cargo audit`, ...) shows no critical or high vulnerabilities
36- [ ] Input validation on all user-facing endpoints
37- [ ] Authentication and authorization checks in place
38- [ ] Security headers configured (CSP, HSTS, etc.)
39- [ ] Rate limiting on authentication endpoints
40- [ ] CORS configured to specific origins (not wildcard)
41 
42### Performance
43 
44- [ ] Core Web Vitals within "Good" thresholds
45- [ ] No N+1 queries in critical paths
46- [ ] Images optimized (compression, responsive sizes, lazy loading)
47- [ ] Bundle size within budget
48- [ ] Database queries have appropriate indexes
49- [ ] Caching configured for static assets and repeated queries
50 
51### Accessibility
52 
53- [ ] Keyboard navigation works for all interactive elements
54- [ ] Screen reader can convey page content and structure
55- [ ] Color contrast meets WCAG 2.1 AA (4.5:1 for text)
56- [ ] Focus management correct for modals and dynamic content
57- [ ] Error messages are descriptive and associated with form fields
58- [ ] No accessibility warnings in axe-core or Lighthouse
59 
60### Infrastructure
61 
62- [ ] Environment variables set in production
63- [ ] Database migrations applied (or ready to apply)
64- [ ] DNS and SSL configured
65- [ ] CDN configured for static assets
66- [ ] Logging and error reporting configured
67- [ ] Health check endpoint exists and responds
68 
69### Documentation
70 
71- [ ] README updated with any new setup requirements
72- [ ] API documentation current
73- [ ] ADRs written for any architectural decisions
74- [ ] Changelog updated
75- [ ] User-facing documentation updated (if applicable)
76 
77## Feature Flag Strategy
78 
79Ship behind feature flags to decouple deployment from release:
80 
81```typescript
82// Feature flag check
83const flags = await getFeatureFlags(userId);
84 
85if (flags.taskSharing) {
86 // New feature: task sharing
87 return <TaskSharingPanel task={task} />;
88}
89 
90// Default: existing behavior
91return null;
92```
93 
94**Feature flag lifecycle:**
95 
96```
971. DEPLOY with flag OFF → Code is in production but inactive
982. ENABLE for team/beta → Internal testing in production environment
993. GRADUAL ROLLOUT → 5% → 25% → 50% → 100% of users
1004. MONITOR at each stage → Watch error rates, performance, user feedback
1015. CLEAN UP → Remove flag and dead code path after full rollout
102```
103 
104**Rules:**
105- Every feature flag has an owner and an expiration date
106- Clean up flags within 2 weeks of full rollout
107- Don't nest feature flags (creates exponential combinations)
108- Test both flag states (on and off) in CI
109 
110## Staged Rollout
111 
112### The Rollout Sequence
113 
114```
1151. DEPLOY to staging
116 └── Full test suite in staging environment
117 └── Manual smoke test of critical flows
118 
1192. DEPLOY to production (feature flag OFF)
120 └── Verify deployment succeeded (health check)
121 └── Check error monitoring (no new errors)
122 
1233. ENABLE for team (flag ON for internal users)
124 └── Team uses the feature in production
125 └── 24-hour monitoring window
126 
1274. CANARY rollout (flag ON for 5% of users)
128 └── Monitor error rates, latency, user behavior
129 └── Compare metrics: canary vs. baseline
130 └── 24-48 hour monitoring window
131 └── Advance only if all thresholds pass (see table below)
132 
1335. GRADUAL increase (25% -> 50% -> 100%)
134 └── Same monitoring at each step
135 └── Ability to roll back to previous percentage at any point
136 
1376. FULL rollout (flag ON for all users)
138 └── Monitor for 1 week
139 └── Clean up feature flag
140```
141 
142### Rollout Decision Thresholds
143 
144Use these thresholds to decide whether to advance, hold, or roll back at each stage:
145 
146| Metric | Advance (green) | Hold and investigate (yellow) | Roll back (red) |
147|--------|-----------------|-------------------------------|-----------------|
148| Error rate | Within 10% of baseline | 10-100% above baseline | >2x baseline |
149| P95 latency | Within 20% of baseline | 20-50% above baseline | >50% above baseline |
150| Client JS errors | No new error types | New errors at <0.1% of sessions | New errors at >0.1% of sessions |
151| Business metrics | Neutral or positive | Decline <5% (may be noise) | Decline >5% |
152 
153### When to Roll Back
154 
155Roll back immediately if:
156- Error rate increases by more than 2x baseline
157- P95 latency increases by more than 50%
158- User-reported issues spike
159- Data integrity issues detected
160- Security vulnerability discovered
161 
162## Monitoring and Observability
163 
164### What to Monitor
165 
166```
167Application metrics:
168├── Error rate (total and by endpoint)
169├── Response time (p50, p95, p99)
170├── Request volume
171├── Active users
172└── Key business metrics (conversion, engagement)
173 
174Infrastructure metrics:
175├── CPU and memory utilization
176├── Database connection pool usage
177├── Disk space
178├── Network latency
179└── Queue depth (if applicable)
180 
181Client metrics:
182├── Core Web Vitals (LCP, INP, CLS)
183├── JavaScript errors
184├── API error rates from client perspective
185└── Page load time
186```
187 
188### Error Reporting
189 
190```typescript
191// Set up error boundary with reporting
192class ErrorBoundary extends React.Component {
193 componentDidCatch(error: Error, info: React.ErrorInfo) {
194 // Report to error tracking service
195 reportError(error, {
196 componentStack: info.componentStack,
197 userId: getCurrentUser()?.id,
198 page: window.location.pathname,
199 });
200 }
201 
202 render() {
203 if (this.state.hasError) {
204 return <ErrorFallback onRetry={() => this.setState({ hasError: false })} />;
205 }
206 return this.props.children;
207 }
208}
209 
210// Server-side error reporting
211app.use((err: Error, req: Request, res: Response, next: NextFunction) => {
212 reportError(err, {
213 method: req.method,
214 url: req.url,
215 userId: req.user?.id,
216 });
217 
218 // Don't expose internals to users
219 res.status(500).json({
220 error: { code: 'INTERNAL_ERROR', message: 'Something went wrong' },
221 });
222});
223```
224 
225### Post-Launch Verification
226 
227In the first hour after launch:
228 
229```
2301. Check health endpoint returns 200
2312. Check error monitoring dashboard (no new error types)
2323. Check latency dashboard (no regression)
2334. Test the critical user flow manually
2345. Verify logs are flowing and readable
2356. Confirm rollback mechanism works (dry run if possible)
236```
237 
238## Error Budget Release Gate
239 
240Your service's error budget — the fraction of requests or time your SLO allows to fail — determines whether it's safe to ship. Use it as an objective gate — not a negotiation:
241 
242```
243Budget remaining > 20% → Ship normally; monitor closely
244Budget remaining 0–20% → Slow rollouts only; no high-risk changes
245Budget exhausted → Freeze feature work; focus entirely on reliability
246Budget resets → Resume normal pace; bake in the fix that recovered it
247```
248 
249A high burn rate during a canary (consuming budget faster than the baseline pace) is a **hold** signal in the rollout thresholds table above — treat it the same as an elevated error rate.
250 
251## Rollback Strategy
252 
253Every deployment needs a rollback plan before it happens:
254 
255```markdown
256## Rollback Plan for [Feature/Release]
257 
258### Trigger Conditions
259- Error rate > 2x baseline
260- P95 latency > [X]ms
261- User reports of [specific issue]
262 
263### Rollback Steps
2641. Disable feature flag (if applicable)
265 OR
2661. Deploy previous version: `git revert <commit> && git push`
2672. Verify rollback: health check, error monitoring
2683. Communicate: notify team of rollback
269 
270### Database Considerations
271- Migration [X] has a rollback: `npx prisma migrate rollback`
272- Data inserted by new feature: [preserved / cleaned up]
273 
274### Time to Rollback
275- Feature flag: < 1 minute
276- Redeploy previous version: < 5 minutes
277- Database rollback: < 15 minutes
278```
279## See Also
280 
281- For the project-wide Definition of Done that every change must clear before this checklist, see `../../references/definition-of-done.md`
282- For security pre-launch checks, see `../../references/security-checklist.md`
283- For performance pre-launch checklist, see `../../references/performance-checklist.md`
284- For accessibility verification before launch, see `../../references/accessibility-checklist.md`
285- For the alerting rules and SLO-tied thresholds, see `observability-and-instrumentation`
286 
287## Common Rationalizations
288 
289| Rationalization | Reality |
290|---|---|
291| "It works in staging, it'll work in production" | Production has different data, traffic patterns, and edge cases. Monitor after deploy. |
292| "We don't need feature flags for this" | Every feature benefits from a kill switch. Even "simple" changes can break things. |
293| "Monitoring is overhead" | Not having monitoring means you discover problems from user complaints instead of dashboards. |
294| "We'll add monitoring later" | Add it before launch. You can't debug what you can't see. |
295| "Rolling back is admitting failure" | Rolling back is responsible engineering. Shipping a broken feature is the failure. |
296| "The error rate looks fine, let's keep shipping" | Check the burn rate, not just the current error rate. Consuming budget faster than baseline is a hold signal even when individual thresholds are green. |
297 
298## Red Flags
299 
300- Deploying without a rollback plan
301- No monitoring or error reporting in production
302- Big-bang releases (everything at once, no staging)
303- Feature flags with no expiration or owner
304- No one monitoring the deploy for the first hour
305- Production environment configuration done by memory, not code
306- "It's Friday afternoon, let's ship it"
307- Error budget exhausted but feature work continues unchanged
308 
309## Verification
310 
311Before deploying:
312 
313- [ ] Pre-launch checklist completed (all sections green)
314- [ ] Feature flag configured (if applicable)
315- [ ] Rollback plan documented
316- [ ] Monitoring dashboards set up
317- [ ] Team notified of deployment
318 
319After deploying:
320 
321- [ ] Health check returns 200
322- [ ] Error rate is normal
323- [ ] Latency is normal
324- [ ] Critical user flow works
325- [ ] Logs are flowing
326- [ ] Rollback tested or verified ready
327 
328For every shipped service:
329 
330- [ ] Error budget policy in place: know what action to take when budget drops below 20% and when it's exhausted
331 

Discussion

Alternatives

Also in Launch planningSee all 277 in Product →
AI Product Launch PlaybookLaunch your AI product to global attention — the playbook behind Manus, Devin, and AFFiNE's breakout launches. Covers AI-specific GTM strategy, hype cycle management, waitlist tactics, and multi-market rollout for maximum day-one impact.Business & ops · MITLaunch StrategyWhen the user wants to plan a product launch, feature announcement, or release strategy. Also use when the user mentions 'launch,' 'Product Hunt,' 'feature release,' 'announcement,' 'go-to-market,' 'beta launch,' 'early access,' 'waitlist,' 'product update,' 'how do I launch this,' 'launch checklist,' 'GTM plan,' or 'we're about to ship.' Use this whenever someone is preparing to release something publicly. For ongoing marketing after launch, see marketing-ideas. For the offer being launched (bonuses, guarantees, scarcity, naming), see offers.Marketing · MITPacsomaticOperator toolkit for nf-core/pacsomatic matched tumor-normal workflows from BAM inputs. Use this skill when the user needs to validate run inputs, generate pacsomatic-compliant samplesheets, prepare reproducible Nextflow launch artifacts, run locally or submit to schedulers (LSF/Slurm/PBS/SGE), and triage execution failures. Triggers on requests to run pacsomatic, prepare launch commands/scripts, perform dry-run checks, or troubleshoot pipeline startup and scheduler submission errors.Science · MITagent-launcher — Domain OrchestratorUse when a user wants to build, launch, grade, or schedule a Claude Managed Agent (CMA) in their own Anthropic account — "build me an agent", "launch this as a managed agent", "run this on a schedule", "grade my agent against a rubric", "set up a nightly worker". Reads the per-session goal (./my-agent/goal.json), routes deterministically to one of five phase sub-skills (interview → stage-launch → grade-iterate → run-without-you → wrap-up) via goal_router.py, and compiles the goal+phase into an execution shape (single-pass workflow / bounded grade→iterate loop / recurring cron deployment loop) via loop_compiler.py. Forks context so heavy intake (build sheets, payloads, eval cases) stays out of the parent thread. All launches are emitted as BYOK curl the user runs with their own key; no tool makes API calls. Inspired by anthropics/launch-your-agent (Apache-2.0). Distinct from engineering/agent-harness (generic domain loop) and engineering/write-a-skill (authors Claude Code skills, not CMAs).Business & ops · MIT