DevOps & Infrastructure skill
Guides Docker, CI/CD pipelines, deployment strategies, infrastructure as code, and observability setup.
by CloudAI-X·MIT license·★ 1,416 Stars on the repo·GitHub ↗
Use now
npx degit CloudAI-X/claude-workflow-v2/skills/devops-infrastructure#main ~/.claude/skills/devops-infrastructureChecked ·commit main
Files of DevOps & Infrastructure
SKILL.md
Show the full text481 lines
DevOps & Infrastructure
When to Load
- Trigger: Docker, CI/CD pipelines, deployment configuration, monitoring, infrastructure as code
- Skip: Application logic only with no infrastructure or deployment concerns
DevOps Workflow
Copy this checklist and track progress:
DevOps Setup Progress:
- [ ] Step 1: Containerize application (Dockerfile)
- [ ] Step 2: Set up CI/CD pipeline
- [ ] Step 3: Define deployment strategy
- [ ] Step 4: Configure monitoring & alerting
- [ ] Step 5: Set up environment management
- [ ] Step 6: Document runbooks
- [ ] Step 7: Validate against anti-patterns checklist
Docker Best Practices
Multi-Stage Build
# WRONG: Single stage, bloated image
FROM node:22
WORKDIR /app
COPY . .
RUN npm install
RUN npm run build
CMD ["node", "dist/index.js"]
# Result: 1.2GB image with devDependencies and source code
# CORRECT: Multi-stage build
FROM node:22-alpine AS builder
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY . .
RUN npm run build
RUN npm prune --omit=dev
FROM node:22-alpine AS runner
WORKDIR /app
ENV NODE_ENV=production
RUN addgroup -g 1001 appgroup && adduser -u 1001 -G appgroup -s /bin/sh -D appuser
COPY --from=builder /app/dist ./dist
COPY --from=builder /app/node_modules ./node_modules
COPY --from=builder /app/package.json ./
USER appuser
EXPOSE 3000
CMD ["node", "dist/index.js"]
# Result: ~150MB image, no devDependencies, non-root user
Python Multi-Stage
FROM python:3.12-slim AS builder
WORKDIR /app
RUN pip install uv
COPY pyproject.toml uv.lock ./
RUN uv sync --frozen --no-dev --no-install-project
COPY . .
RUN uv sync --frozen --no-dev
FROM python:3.12-slim AS runner
WORKDIR /app
RUN useradd -r -s /bin/false appuser
COPY --from=builder /app/.venv /app/.venv
COPY --from=builder /app/src ./src
ENV PATH="/app/.venv/bin:$PATH"
USER appuser
CMD ["python", "-m", "src.main"]
Layer Caching
# WRONG: Cache busted on every code change
COPY . .
RUN npm ci
# CORRECT: Dependencies cached separately
COPY package*.json ./
RUN npm ci # cached unless package.json changes
COPY . . # only source code changes bust this layer
.dockerignore
node_modules
.git
.env
*.md
.vscode
coverage
dist
__pycache__
.pytest_cache
*.pyc
Security
# Always pin versions
FROM node:22-alpine # NOT node:latest
# Don't run as root
USER appuser
# Read-only filesystem where possible
# docker run --read-only --tmpfs /tmp myapp
# Scan images
# docker scout cves myimage:latest
# trivy image myimage:latest
CI/CD Pipeline Design
GitHub Actions Structure
name: CI/CD
on:
push:
branches: [main]
pull_request:
branches: [main]
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: "npm"
- run: npm ci
- run: npm run lint
test:
runs-on: ubuntu-latest
needs: lint
services:
postgres:
image: postgres:16
env:
POSTGRES_DB: testdb
POSTGRES_PASSWORD: postgres
ports: ["5432:5432"]
options: >-
--health-cmd pg_isready
--health-interval 10s
--health-timeout 5s
--health-retries 5
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: "npm"
- run: npm ci
- run: npm test
build:
runs-on: ubuntu-latest
needs: test
permissions:
contents: read
packages: write
steps:
- uses: actions/checkout@v4
- uses: docker/setup-buildx-action@v3
- uses: docker/login-action@v3
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- id: meta
uses: docker/metadata-action@v5
with:
images: ghcr.io/${{ github.repository }}
tags: type=sha
- uses: docker/build-push-action@v5
with:
push: ${{ github.event_name == 'push' }}
tags: ${{ steps.meta.outputs.tags }}
cache-from: type=gha
cache-to: type=gha,mode=max
deploy:
runs-on: ubuntu-latest
needs: build
if: github.ref == 'refs/heads/main'
environment: production
steps:
- run: echo "Deploy to production"
Caching Strategies
# Node modules
- uses: actions/setup-node@v4
with:
cache: "npm"
# Python with uv
- name: Cache uv
uses: actions/cache@v4
with:
path: ~/.cache/uv
key: uv-${{ runner.os }}-${{ hashFiles('uv.lock') }}
# Docker layer caching
- uses: docker/build-push-action@v5
with:
cache-from: type=gha
cache-to: type=gha,mode=max
Deployment Strategies
Blue-Green Deployment
1. Run two identical environments: Blue (live) and Green (idle)
2. Deploy new version to Green
3. Run smoke tests on Green
4. Switch load balancer to Green
5. Green is now live, Blue is idle
6. Rollback: switch back to Blue
Pros: Instant rollback, zero downtime
Cons: 2x infrastructure cost during deploy
Canary Deployment
1. Deploy new version to small subset (5% of traffic)
2. Monitor error rates and latency
3. Gradually increase: 5% -> 25% -> 50% -> 100%
4. Rollback: route all traffic back to old version
Pros: Limited blast radius, real-world testing
Cons: More complex routing, longer rollout
Rolling Deployment
1. Replace instances one at a time
2. Each new instance passes health checks before next starts
3. Continue until all instances updated
Pros: No extra infrastructure, gradual rollout
Cons: Mixed versions during deploy, slower rollback
Feature Flags
// Simple feature flag implementation
const features = {
NEW_CHECKOUT: process.env.FF_NEW_CHECKOUT === "true",
DARK_MODE: process.env.FF_DARK_MODE === "true",
};
function getCheckoutFlow(user: User) {
if (features.NEW_CHECKOUT && user.betaGroup) {
return newCheckoutFlow(user);
}
return legacyCheckoutFlow(user);
}
// Use a proper service for production: LaunchDarkly, Unleash, Flagsmith
Infrastructure as Code
Terraform Basics
# main.tf
terraform {
required_version = ">= 1.5"
backend "s3" {
bucket = "myapp-terraform-state"
key = "prod/terraform.tfstate"
region = "us-east-1"
use_lockfile = true
}
}
resource "aws_instance" "web" {
ami = var.ami_id
instance_type = var.instance_type
tags = {
Name = "web-${var.environment}"
Environment = var.environment
ManagedBy = "terraform"
}
}
# variables.tf
variable "ami_id" {
type = string
}
variable "environment" {
type = string
default = "dev"
}
variable "instance_type" {
type = string
default = "t3.micro"
}
Terraform Rules
1. Always use remote state (S3, GCS, Terraform Cloud)
2. Lock state files to prevent concurrent modifications
3. Use variables and modules for reusability
4. Tag all resources with environment and ManagedBy
5. Run `terraform plan` before `terraform apply`
6. Never edit infrastructure manually (all changes via code)
7. Use workspaces or separate state files per environment
Monitoring & Observability
The Three Pillars
METRICS: Numeric measurements over time
- Request rate, error rate, latency (RED method)
- CPU, memory, disk, network (USE method)
- Business metrics (signups, purchases)
Tools: Prometheus, Datadog, CloudWatch
LOGS: Discrete events with context
- Structured JSON format
- Correlation IDs across services
- Log levels: DEBUG, INFO, WARN, ERROR
Tools: ELK Stack, Loki, CloudWatch Logs
TRACES: Request flow across services
- Distributed tracing with span context
- Latency breakdown per service
- Dependency mapping
Tools: Jaeger, Zipkin, Datadog APM
Health Check Endpoint
// Express health check
app.get("/health", async (req, res) => {
const checks = {
uptime: process.uptime(),
timestamp: Date.now(),
database: "unknown",
redis: "unknown",
};
try {
await db.query("SELECT 1");
checks.database = "healthy";
} catch (e) {
checks.database = "unhealthy";
}
try {
await redis.ping();
checks.redis = "healthy";
} catch (e) {
checks.redis = "unhealthy";
}
const isHealthy = checks.database === "healthy";
res.status(isHealthy ? 200 : 503).json(checks);
});
Alerting Rules
Good alerts:
- Error rate > 1% for 5 minutes (actionable)
- P99 latency > 2s for 10 minutes (meaningful)
- Disk usage > 80% (preventive)
Bad alerts:
- CPU spike for 30 seconds (too noisy)
- Any single 500 error (too sensitive)
- "Something might be wrong" (not actionable)
Alert fatigue is real. Every alert should require human action.
Environment Management
Dev/Staging/Prod Parity
# docker-compose.yml for local development
services:
app:
build: .
env_file: .env
ports: ["3000:3000"]
depends_on:
postgres:
condition: service_healthy
postgres:
image: postgres:16
environment:
POSTGRES_DB: myapp
POSTGRES_PASSWORD: postgres
healthcheck:
test: ["CMD-SHELL", "pg_isready"]
interval: 5s
volumes:
- pgdata:/var/lib/postgresql/data
redis:
image: redis:7-alpine
ports: ["6379:6379"]
volumes:
pgdata:
Environment Variables
# .env.example (committed to git, no real values)
DATABASE_URL=postgresql://user:placeholder@localhost:5432/myapp
REDIS_URL=redis://localhost:6379
LOG_LEVEL=debug
API_KEY=your-key-here
# .env (never committed, listed in .gitignore)
# Contains real values for local development
Common Anti-Patterns Summary
AVOID DO INSTEAD
-------------------------------------------------------------------
FROM node:latest Pin versions (node:22-alpine)
Running as root in container Create and use non-root user
No .dockerignore Exclude .git, node_modules, .env
Single CI job does everything Separate lint, test, build, deploy stages
Manual deployment Automated pipeline with approvals
No health checks Liveness + readiness probes
Alerts on every error Alert on error RATE thresholds
Same config in all environments Per-environment configuration
No rollback plan Test rollback before every deploy
Logs as unstructured strings Structured JSON logs with correlation IDs
| 1 | |
| 2 | name devops-infrastructure |
| 3 | description Guides Docker, CI/CD pipelines, deployment strategies, infrastructure as code, and observability setup. Use when writing Dockerfiles, configuring GitHub Actions, planning deployments, setting up monitoring, or when asked about containers, pipelines, Terraform, or production infrastructure. |
| 4 | |
| 5 | |
| 6 | # DevOps & Infrastructure |
| 7 | |
| 8 | ### When to Load |
| 9 | |
| 10 | **Trigger**: Docker, CI/CD pipelines, deployment configuration, monitoring, infrastructure as code |
| 11 | **Skip**: Application logic only with no infrastructure or deployment concerns |
| 12 | |
| 13 | ## DevOps Workflow |
| 14 | |
| 15 | Copy this checklist and track progress: |
| 16 | |
| 17 | |
| 18 | DevOps Setup Progress: |
| 19 | - [ ] Step 1: Containerize application (Dockerfile) |
| 20 | - [ ] Step 2: Set up CI/CD pipeline |
| 21 | - [ ] Step 3: Define deployment strategy |
| 22 | - [ ] Step 4: Configure monitoring & alerting |
| 23 | - [ ] Step 5: Set up environment management |
| 24 | - [ ] Step 6: Document runbooks |
| 25 | - [ ] Step 7: Validate against anti-patterns checklist |
| 26 | |
| 27 | |
| 28 | ## Docker Best Practices |
| 29 | |
| 30 | ### Multi-Stage Build |
| 31 | |
| 32 | |
| 33 | # WRONG: Single stage, bloated image |
| 34 | FROM node:22 |
| 35 | WORKDIR /app |
| 36 | COPY . . |
| 37 | RUN npm install |
| 38 | RUN npm run build |
| 39 | CMD ["node", "dist/index.js"] |
| 40 | # Result: 1.2GB image with devDependencies and source code |
| 41 | |
| 42 | # CORRECT: Multi-stage build |
| 43 | FROM node:22-alpine AS builder |
| 44 | WORKDIR /app |
| 45 | COPY package*.json ./ |
| 46 | RUN npm ci |
| 47 | COPY . . |
| 48 | RUN npm run build |
| 49 | RUN npm prune --omit=dev |
| 50 | |
| 51 | FROM node:22-alpine AS runner |
| 52 | WORKDIR /app |
| 53 | ENV NODE_ENV=production |
| 54 | RUN addgroup -g 1001 appgroup && adduser -u 1001 -G appgroup -s /bin/sh -D appuser |
| 55 | COPY --from=builder /app/dist ./dist |
| 56 | COPY --from=builder /app/node_modules ./node_modules |
| 57 | COPY --from=builder /app/package.json ./ |
| 58 | USER appuser |
| 59 | EXPOSE 3000 |
| 60 | CMD ["node", "dist/index.js"] |
| 61 | # Result: ~150MB image, no devDependencies, non-root user |
| 62 | |
| 63 | |
| 64 | ### Python Multi-Stage |
| 65 | |
| 66 | |
| 67 | FROM python:3.12-slim AS builder |
| 68 | WORKDIR /app |
| 69 | RUN pip install uv |
| 70 | COPY pyproject.toml uv.lock ./ |
| 71 | RUN uv sync --frozen --no-dev --no-install-project |
| 72 | COPY . . |
| 73 | RUN uv sync --frozen --no-dev |
| 74 | |
| 75 | FROM python:3.12-slim AS runner |
| 76 | WORKDIR /app |
| 77 | RUN useradd -r -s /bin/false appuser |
| 78 | COPY --from=builder /app/.venv /app/.venv |
| 79 | COPY --from=builder /app/src ./src |
| 80 | ENV PATH="/app/.venv/bin:$PATH" |
| 81 | USER appuser |
| 82 | CMD ["python", "-m", "src.main"] |
| 83 | |
| 84 | |
| 85 | ### Layer Caching |
| 86 | |
| 87 | |
| 88 | # WRONG: Cache busted on every code change |
| 89 | COPY . . |
| 90 | RUN npm ci |
| 91 | |
| 92 | # CORRECT: Dependencies cached separately |
| 93 | COPY package*.json ./ |
| 94 | RUN npm ci # cached unless package.json changes |
| 95 | COPY . . # only source code changes bust this layer |
| 96 | |
| 97 | |
| 98 | ### .dockerignore |
| 99 | |
| 100 | |
| 101 | node_modules |
| 102 | .git |
| 103 | .env |
| 104 | *.md |
| 105 | .vscode |
| 106 | coverage |
| 107 | dist |
| 108 | __pycache__ |
| 109 | .pytest_cache |
| 110 | *.pyc |
| 111 | |
| 112 | |
| 113 | ### Security |
| 114 | |
| 115 | |
| 116 | # Always pin versions |
| 117 | FROM node:22-alpine # NOT node:latest |
| 118 | |
| 119 | # Don't run as root |
| 120 | USER appuser |
| 121 | |
| 122 | # Read-only filesystem where possible |
| 123 | # docker run --read-only --tmpfs /tmp myapp |
| 124 | |
| 125 | # Scan images |
| 126 | # docker scout cves myimage:latest |
| 127 | # trivy image myimage:latest |
| 128 | |
| 129 | |
| 130 | ## CI/CD Pipeline Design |
| 131 | |
| 132 | ### GitHub Actions Structure |
| 133 | |
| 134 | |
| 135 | name: CI/CD |
| 136 | on: |
| 137 | push: |
| 138 | branches: [main] |
| 139 | pull_request: |
| 140 | branches: [main] |
| 141 | |
| 142 | jobs: |
| 143 | lint: |
| 144 | runs-on: ubuntu-latest |
| 145 | steps: |
| 146 | - uses: actions/checkout@v4 |
| 147 | - uses: actions/setup-node@v4 |
| 148 | with: |
| 149 | node-version: 22 |
| 150 | cache: "npm" |
| 151 | - run: npm ci |
| 152 | - run: npm run lint |
| 153 | |
| 154 | test: |
| 155 | runs-on: ubuntu-latest |
| 156 | needs: lint |
| 157 | services: |
| 158 | postgres: |
| 159 | image: postgres:16 |
| 160 | env: |
| 161 | POSTGRES_DB: testdb |
| 162 | POSTGRES_PASSWORD: postgres |
| 163 | ports: ["5432:5432"] |
| 164 | options: >- |
| 165 | --health-cmd pg_isready |
| 166 | --health-interval 10s |
| 167 | --health-timeout 5s |
| 168 | --health-retries 5 |
| 169 | steps: |
| 170 | - uses: actions/checkout@v4 |
| 171 | - uses: actions/setup-node@v4 |
| 172 | with: |
| 173 | node-version: 22 |
| 174 | cache: "npm" |
| 175 | - run: npm ci |
| 176 | - run: npm test |
| 177 | |
| 178 | build: |
| 179 | runs-on: ubuntu-latest |
| 180 | needs: test |
| 181 | permissions: |
| 182 | contents: read |
| 183 | packages: write |
| 184 | steps: |
| 185 | - uses: actions/checkout@v4 |
| 186 | - uses: docker/setup-buildx-action@v3 |
| 187 | - uses: docker/login-action@v3 |
| 188 | with: |
| 189 | registry: ghcr.io |
| 190 | username: ${{ github.actor }} |
| 191 | password: ${{ secrets.GITHUB_TOKEN }} |
| 192 | - id: meta |
| 193 | uses: docker/metadata-action@v5 |
| 194 | with: |
| 195 | images: ghcr.io/${{ github.repository }} |
| 196 | tags: type=sha |
| 197 | - uses: docker/build-push-action@v5 |
| 198 | with: |
| 199 | push: ${{ github.event_name == 'push' }} |
| 200 | tags: ${{ steps.meta.outputs.tags }} |
| 201 | cache-from: type=gha |
| 202 | cache-to: type=gha,mode=max |
| 203 | |
| 204 | deploy: |
| 205 | runs-on: ubuntu-latest |
| 206 | needs: build |
| 207 | if: github.ref == 'refs/heads/main' |
| 208 | environment: production |
| 209 | steps: |
| 210 | - run: echo "Deploy to production" |
| 211 | |
| 212 | |
| 213 | ### Caching Strategies |
| 214 | |
| 215 | |
| 216 | # Node modules |
| 217 | - uses: actions/setup-node@v4 |
| 218 | with: |
| 219 | cache: "npm" |
| 220 | |
| 221 | # Python with uv |
| 222 | - name: Cache uv |
| 223 | uses: actions/cache@v4 |
| 224 | with: |
| 225 | path: ~/.cache/uv |
| 226 | key: uv-${{ runner.os }}-${{ hashFiles('uv.lock') }} |
| 227 | |
| 228 | # Docker layer caching |
| 229 | - uses: docker/build-push-action@v5 |
| 230 | with: |
| 231 | cache-from: type=gha |
| 232 | cache-to: type=gha,mode=max |
| 233 | |
| 234 | |
| 235 | ## Deployment Strategies |
| 236 | |
| 237 | ### Blue-Green Deployment |
| 238 | |
| 239 | |
| 240 | 1. Run two identical environments: Blue (live) and Green (idle) |
| 241 | 2. Deploy new version to Green |
| 242 | 3. Run smoke tests on Green |
| 243 | 4. Switch load balancer to Green |
| 244 | 5. Green is now live, Blue is idle |
| 245 | 6. Rollback: switch back to Blue |
| 246 | |
| 247 | Pros: Instant rollback, zero downtime |
| 248 | Cons: 2x infrastructure cost during deploy |
| 249 | |
| 250 | |
| 251 | ### Canary Deployment |
| 252 | |
| 253 | |
| 254 | 1. Deploy new version to small subset (5% of traffic) |
| 255 | 2. Monitor error rates and latency |
| 256 | 3. Gradually increase: 5% -> 25% -> 50% -> 100% |
| 257 | 4. Rollback: route all traffic back to old version |
| 258 | |
| 259 | Pros: Limited blast radius, real-world testing |
| 260 | Cons: More complex routing, longer rollout |
| 261 | |
| 262 | |
| 263 | ### Rolling Deployment |
| 264 | |
| 265 | |
| 266 | 1. Replace instances one at a time |
| 267 | 2. Each new instance passes health checks before next starts |
| 268 | 3. Continue until all instances updated |
| 269 | |
| 270 | Pros: No extra infrastructure, gradual rollout |
| 271 | Cons: Mixed versions during deploy, slower rollback |
| 272 | |
| 273 | |
| 274 | ### Feature Flags |
| 275 | |
| 276 | |
| 277 | // Simple feature flag implementation |
| 278 | const features = { |
| 279 | NEW_CHECKOUT: process.env.FF_NEW_CHECKOUT === "true", |
| 280 | DARK_MODE: process.env.FF_DARK_MODE === "true", |
| 281 | }; |
| 282 | |
| 283 | function getCheckoutFlow(user: User) { |
| 284 | if (features.NEW_CHECKOUT && user.betaGroup) { |
| 285 | return newCheckoutFlow(user); |
| 286 | } |
| 287 | return legacyCheckoutFlow(user); |
| 288 | } |
| 289 | |
| 290 | // Use a proper service for production: LaunchDarkly, Unleash, Flagsmith |
| 291 | |
| 292 | |
| 293 | ## Infrastructure as Code |
| 294 | |
| 295 | ### Terraform Basics |
| 296 | |
| 297 | |
| 298 | # main.tf |
| 299 | terraform { |
| 300 | required_version = ">= 1.5" |
| 301 | backend "s3" { |
| 302 | bucket = "myapp-terraform-state" |
| 303 | key = "prod/terraform.tfstate" |
| 304 | region = "us-east-1" |
| 305 | use_lockfile = true |
| 306 | } |
| 307 | } |
| 308 | |
| 309 | resource "aws_instance" "web" { |
| 310 | ami = var.ami_id |
| 311 | instance_type = var.instance_type |
| 312 | tags = { |
| 313 | Name = "web-${var.environment}" |
| 314 | Environment = var.environment |
| 315 | ManagedBy = "terraform" |
| 316 | } |
| 317 | } |
| 318 | |
| 319 | # variables.tf |
| 320 | variable "ami_id" { |
| 321 | type = string |
| 322 | } |
| 323 | |
| 324 | variable "environment" { |
| 325 | type = string |
| 326 | default = "dev" |
| 327 | } |
| 328 | |
| 329 | variable "instance_type" { |
| 330 | type = string |
| 331 | default = "t3.micro" |
| 332 | } |
| 333 | |
| 334 | |
| 335 | ### Terraform Rules |
| 336 | |
| 337 | |
| 338 | 1. Always use remote state (S3, GCS, Terraform Cloud) |
| 339 | 2. Lock state files to prevent concurrent modifications |
| 340 | 3. Use variables and modules for reusability |
| 341 | 4. Tag all resources with environment and ManagedBy |
| 342 | 5. Run `terraform plan` before `terraform apply` |
| 343 | 6. Never edit infrastructure manually (all changes via code) |
| 344 | 7. Use workspaces or separate state files per environment |
| 345 | |
| 346 | |
| 347 | ## Monitoring & Observability |
| 348 | |
| 349 | ### The Three Pillars |
| 350 | |
| 351 | |
| 352 | METRICS: Numeric measurements over time |
| 353 | - Request rate, error rate, latency (RED method) |
| 354 | - CPU, memory, disk, network (USE method) |
| 355 | - Business metrics (signups, purchases) |
| 356 | Tools: Prometheus, Datadog, CloudWatch |
| 357 | |
| 358 | LOGS: Discrete events with context |
| 359 | - Structured JSON format |
| 360 | - Correlation IDs across services |
| 361 | - Log levels: DEBUG, INFO, WARN, ERROR |
| 362 | Tools: ELK Stack, Loki, CloudWatch Logs |
| 363 | |
| 364 | TRACES: Request flow across services |
| 365 | - Distributed tracing with span context |
| 366 | - Latency breakdown per service |
| 367 | - Dependency mapping |
| 368 | Tools: Jaeger, Zipkin, Datadog APM |
| 369 | |
| 370 | |
| 371 | ### Health Check Endpoint |
| 372 | |
| 373 | |
| 374 | // Express health check |
| 375 | app.get("/health", async (req, res) => { |
| 376 | const checks = { |
| 377 | uptime: process.uptime(), |
| 378 | timestamp: Date.now(), |
| 379 | database: "unknown", |
| 380 | redis: "unknown", |
| 381 | }; |
| 382 | |
| 383 | try { |
| 384 | await db.query("SELECT 1"); |
| 385 | checks.database = "healthy"; |
| 386 | } catch (e) { |
| 387 | checks.database = "unhealthy"; |
| 388 | } |
| 389 | |
| 390 | try { |
| 391 | await redis.ping(); |
| 392 | checks.redis = "healthy"; |
| 393 | } catch (e) { |
| 394 | checks.redis = "unhealthy"; |
| 395 | } |
| 396 | |
| 397 | const isHealthy = checks.database === "healthy"; |
| 398 | res.status(isHealthy ? 200 : 503).json(checks); |
| 399 | }); |
| 400 | |
| 401 | |
| 402 | ### Alerting Rules |
| 403 | |
| 404 | |
| 405 | Good alerts: |
| 406 | - Error rate > 1% for 5 minutes (actionable) |
| 407 | - P99 latency > 2s for 10 minutes (meaningful) |
| 408 | - Disk usage > 80% (preventive) |
| 409 | |
| 410 | Bad alerts: |
| 411 | - CPU spike for 30 seconds (too noisy) |
| 412 | - Any single 500 error (too sensitive) |
| 413 | - "Something might be wrong" (not actionable) |
| 414 | |
| 415 | Alert fatigue is real. Every alert should require human action. |
| 416 | |
| 417 | |
| 418 | ## Environment Management |
| 419 | |
| 420 | ### Dev/Staging/Prod Parity |
| 421 | |
| 422 | |
| 423 | # docker-compose.yml for local development |
| 424 | services: |
| 425 | app: |
| 426 | build: . |
| 427 | env_file: .env |
| 428 | ports: ["3000:3000"] |
| 429 | depends_on: |
| 430 | postgres: |
| 431 | condition: service_healthy |
| 432 | |
| 433 | postgres: |
| 434 | image: postgres:16 |
| 435 | environment: |
| 436 | POSTGRES_DB: myapp |
| 437 | POSTGRES_PASSWORD: postgres |
| 438 | healthcheck: |
| 439 | test: ["CMD-SHELL", "pg_isready"] |
| 440 | interval: 5s |
| 441 | volumes: |
| 442 | - pgdata:/var/lib/postgresql/data |
| 443 | |
| 444 | redis: |
| 445 | image: redis:7-alpine |
| 446 | ports: ["6379:6379"] |
| 447 | |
| 448 | volumes: |
| 449 | pgdata: |
| 450 | |
| 451 | |
| 452 | ### Environment Variables |
| 453 | |
| 454 | |
| 455 | # .env.example (committed to git, no real values) |
| 456 | DATABASE_URL=postgresql://user:placeholder@localhost:5432/myapp |
| 457 | REDIS_URL=redis://localhost:6379 |
| 458 | LOG_LEVEL=debug |
| 459 | API_KEY=your-key-here |
| 460 | |
| 461 | # .env (never committed, listed in .gitignore) |
| 462 | # Contains real values for local development |
| 463 | |
| 464 | |
| 465 | ## Common Anti-Patterns Summary |
| 466 | |
| 467 | |
| 468 | AVOID DO INSTEAD |
| 469 | ------------------------------------------------------------------- |
| 470 | FROM node:latest Pin versions (node:22-alpine) |
| 471 | Running as root in container Create and use non-root user |
| 472 | No .dockerignore Exclude .git, node_modules, .env |
| 473 | Single CI job does everything Separate lint, test, build, deploy stages |
| 474 | Manual deployment Automated pipeline with approvals |
| 475 | No health checks Liveness + readiness probes |
| 476 | Alerts on every error Alert on error RATE thresholds |
| 477 | Same config in all environments Per-environment configuration |
| 478 | No rollback plan Test rollback before every deploy |
| 479 | Logs as unstructured strings Structured JSON logs with correlation IDs |
| 480 | |
| 481 |
Discussion
Alternatives
Azure app onboardEnd-to-end orchestrator: from a business idea, app idea, or existing app to running Azure deployment with cost estimates and pre-deploy approval. Analyzes your app, auto-detects the right Azure services, scaffolds infrastructure code, and deploys — tailored to your app, not a template. Handles moving existing apps to Azure without rewriting or with minimal changes. WHEN: bring your app to Azure, plan my app, cost to run, is my code ready to deploy, deploy my app to the cloud, deploy all my services, what Azure services do I need, plan my Azure deployment, deploy my new app to Azure, one-click deploy, I have an app and want it on Azure, migrate my app to Azure, help me get started, build an app, no code yet, starter project. DO NOT USE FOR: use azd for deployment(use azure-deploy), optimizing existing costs (use cost-optimization), code readiness checks only (use azure-app-onboard-prereq).Azure App Onboard Prereq — Repository EvaluationAssess whether source code is ready to deploy to Azure — the check BEFORE infrastructure work. Evaluates build health, app completeness, dependencies and local services, stack compatibility, and deployment feasibility. Answers questions about what your app needs before it can be deployed — frameworks, dependencies, and configuration. Checks whether dependencies are compatible and identifies deployment blockers and unsupported frameworks. WHEN: "evaluate my repo", "is my app ready to deploy", "what does my app need to deploy", "what do I need before deploying", "does my app need", "can I ship this to Azure", "scan my repo for issues", "is this app deployable", "check if my app is ready for Azure", "do I need a Dockerfile", "what's blocking my deployment", "are there any blockers", "are my dependencies compatible", "does Azure support my framework", "what needs to change before deploying", "check my app configuration".
Docker MCP gatewayDocker's own CLI plugin: run any server from the Docker MCP Catalog in its own container, behind one connection, with secrets kept out of env vars.Azure cloud migrateAssess and migrate cross-cloud workloads to Azure with reports and code conversion. Supports Lambda→Functions, Beanstalk/Heroku/App Engine→App Service, Fargate/Kubernetes/Cloud Run/Spring Boot→Container Apps. WHEN: migrate Lambda to Functions, AWS to Azure, migrate Beanstalk, migrate Heroku, migrate App Engine, Cloud Run migration, Fargate to ACA, ECS/Kubernetes/GKE/EKS to Container Apps, Spring Boot to Container Apps, cross-cloud migration.
Browse more free Claude skills or everything in Development.