Azure Reliability Assessment & Configuration skill

Assess and improve the reliability posture of PaaS Applications (Azure Functions and Azure App Service).

by microsoft·MIT license·★ 1,532 Stars on the repo·GitHub ↗

Use now

Files of Azure Reliability Assessment & Configuration

microsoft/main1 file shown
SKILL.md
Show the full text388 lines

Azure Reliability Assessment & Configuration

Quick Reference

Property Details
Best for Reliability posture assessment, zone redundancy enablement, multi-region failover setup
Primary capabilities Reliability assessment table, Zone Redundancy Configuration, Multi-Region IaC Generation
Supported services Azure Functions, App Service (Container Apps planned for a future version)
MCP tools Azure Resource Graph queries, Azure CLI commands

When to Use This Skill

Activate this skill when user wants to:

  • "Assess my Function app's reliability"
  • "Assess my Web app's reliability"
  • "Check the reliability of my resource group" (App Service and Functions resources only)
  • "Is my app zone redundant?" (App Service and Functions resources only)
  • "Is my app service plan zone redundant?"
  • "Make my app zone redundant" (App Service and Functions resources only)
  • "Make my app service plan zone redundant"
  • "Set up multi-region failover for my app" (App Service and Functions resources only)
  • "Check my reliability posture"
  • "Find single points of failure" (App Service and Functions resources only)
  • "Enable high availability for my app" (App Service and Functions resources only)
  • "Check disaster recovery readiness"
  • "Improve my app's resilience" (App Service and Functions resources only)

Scope note: This skill currently covers Azure Functions and Azure App Service only. If the user asks about Azure Container Apps reliability, acknowledge that support is planned but not yet available, and only proceed with the parts that apply to App Service and Functions resources in scope.

Prerequisites

  • Authentication: user is logged in to Azure via az login
  • Permissions: Reader access on target subscription/resource group (for assessment)
  • Permissions: Contributor access (for configuration changes)
  • Azure Resource Graph extension: az extension add --name resource-graph

MCP Tools

Tool Purpose
mcp_azure_mcp_extension_cli_generate Generate az CLI commands for resource queries and configuration
mcp_azure_mcp_subscription_list List available subscriptions
mcp_azure_mcp_group_list List resource groups

Primary query method: Azure Resource Graph via az graph query (requires az extension add --name resource-graph).

Assessment Workflow

Phase 1: Discover Resources
  1. Identify scope — Ask user for resource group, subscription, or app name
  2. Query Azure Resource Graph to discover all resources in scope
  3. Classify resources by service type (Functions, Storage, etc.). If Container Apps is found, note it but do not deep-dive.

Important: Always scope queries to the user's specified resource group or subscription. Add these filters to every Resource Graph query:

  • Resource group: | where resourceGroup =~ '<rg-name>'
  • Subscription: Use --subscriptions <sub-id> flag on az graph query
  • App name: | where name =~ '<app-name>'
Phase 2: Assess Reliability

Two-step assessment: platform-level discovery first, then per-service deep dive.

Step 1 — Platform discovery (find what's there). Use these to enumerate resources in scope and detect cross-cutting reliability gaps:

Platform check Reference
Zone redundancy — discovery references/zone-redundancy-checks.md
Storage redundancy (cross-service) references/storage-redundancy-checks.md
Multi-region & global load balancers references/multi-region-checks.md
Front Door / Traffic Manager / App Insights probes references/health-probe-checks.md

Step 2 — Per-service deep dive. For each compute resource discovered in Step 1, load the matching service reference. The service reference is the single source of truth for that service's plan/SKU rules, assessment queries, CLI commands, IaC patches (Bicep + Terraform + AVM), and reporting hints.

This skill version ships only the Azure Functions and App Service per-service references. Other compute services are listed below explicitly so the dispatch logic is unambiguous: if a resource matches an unsupported row, do not attempt to load a reference, fabricate CLI commands, or generate IaC patches for it.

Service detected Reference
Azure Functions (microsoft.web/serverfarms with kind contains 'functionapp') references/services/functions/reliability.md
Azure App Service (non-Functions sites: microsoft.web/sites without kind contains 'functionapp', microsoft.web/serverfarms without kind contains 'functionapp') references/services/app-service/reliability.md
Azure Container Apps (microsoft.app/containerapps, microsoft.app/managedenvironments) ⚪ Not yet shipped — planned for a future version

Handling unsupported services: If a resource matches an unsupported row above, surface it in the discovery summary, mark it as ⚪ not assessed (planned) in the Phase 3 table, and skip the per-service remediation steps for it. Do not attempt to fabricate CLI commands or IaC patches for those services.

Phase 3: Generate Reliability Checklist

Present findings as a feature-pivoted table: one row per reliability feature (Zone redundancy on compute, Zone-redundant storage, Health probes, Multi-region failover), with a single status indicator and the specific resources that are relevant to that feature. This avoids the noise of one-row-per-resource with mostly n/a cells. Do not assign numeric scores or grades.

🔍 Reliability Assessment — {scope}
─────────────────────────────────────────────────────────────────────────────────────────────
Reliability Feature              Status      Resources
─────────────────────────────────────────────────────────────────────────────────────────────
Zone redundancy — compute        🔴 OFF      • plan-web-ii5trxva2ark4 (P1v3)
                                              • plan-ii5trxva2ark4 (FC1)

Zone-redundant storage           🔴 GRS      • stii5trxva2ark4 (defaulted; no SKU set in IaC)

Health probes                    🔴 OFF      • func-api-ii5trxva2ark4 — needs code change (FC1)
                                              • app-web-ii5trxva2ark4 — no health check path

Multi-region failover            🔴 OFF      • Single region (eastus) only — Front Door not configured
─────────────────────────────────────────────────────────────────────────────────────────────

Want me to fix the 🔴 items? I'll do the quick wins first (App
plan zone redundancy + health checks on supported plans), then ask before
storage migration and multi-region setup. (yes/no)

Rules for the table:

  • Four feature rows, in this order: Zone redundancy — compute · Zone-redundant storage · Health probes · Multi-region failover. Omit a row entirely only if no resource in scope could ever apply to it.
  • Status column is one symbol + one short word, no other characters:
    • 🟢 ON — feature is fully enabled across all relevant resources in scope
    • 🟡 PARTIAL — some resources have it, some don't (or partial config like liveness-only)
    • 🔴 OFF — feature is missing on all relevant resources
    • For storage, replace OFF with the current SKU when relevant (🔴 LRS, 🔴 GRS, 🟢 ZRS, 🟢 GZRS). When no SKU is set in IaC, label as 🔴 GRS (ARM/AVM default) and note that in the resource line.
  • Resources column lists only what's relevant to that feature, one bullet per resource:
    • For "needs fixing" resources, include a short inline reason ((FC1), (defaulted; no SKU set), liveness only, needs code change (FC1)).
    • For resources that are already ON for that feature, mention them on the same row with — already ON so the user sees credit for what's right.
  • Do not include n/a, —, or empty cells. If a feature doesn't apply to any resource in scope, drop the row.
  • Do not include numeric scores, grades, or point totals.
  • End the assessment with a single yes/no question that kicks off the staged remediation flow. Do not enumerate the per-resource fix list here — the user will see it after they say yes (Configuration Workflow Step 1).

UX Note: If the assessment finds the app already has all core reliability features (zone redundancy, ZRS/GZRS storage, health probes), skip the fix-it question and jump straight to Configuration Workflow Step 3 (Multi-region follow-up). Do NOT start any multi-region work without explicit consent.

Configuration Workflow

When user wants to fix findings from the assessment:

⛔ ALWAYS confirm with user before executing changes. Show what will change, any cost implications, and any destructive actions (e.g., environment recreation).

Step 1: Present Fix Plan + Choose Path

After assessment, if user says "fix it" / "improve my reliability" / "enable zone redundancy":

  1. List each fixable finding with the specific action
  2. Flag any cost implications or breaking changes
  3. Ask user which path they want:
I'll start with the quick wins (no downtime, fast):

1. ✏️  Enable zone redundancy on plan-ii5trxva2ark4 (Flex Consumption — no cost change)
2. ✏️  Set health check path to /api/health on func-api-ii5trxva2ark4

Then, separately, I'll ask if you want to upgrade storage:

3. 🕒  Upgrade stii5trxva2ark4 from LRS → ZRS (small cost increase, migration takes hours)
   — Required for full zone redundancy, but I'll confirm with you before starting.

How would you like to apply these changes?

  A) Fix now — Run az CLI commands against your live resources (immediate, one-time)
  B) Patch my IaC — Update your Bicep/Terraform files so changes persist across deploys

(If you use azd or Terraform, option B is recommended so `azd up` won't overwrite changes.)
Path A: Fix Now (CLI)

Run fixes against live resources using az CLI commands. Quick wins first, then ask before the slow storage migration.

The exact CLI commands per service live in the per-service references — pick the one(s) matching the resources discovered in Phase 2:

Fix Reference
Enable zone redundancy / configure health probes (Functions) references/services/functions/reliability.md
Enable zone redundancy / configure health probes (App Service) references/services/app-service/reliability.md
Upgrade storage replication (cross-service) references/configure-storage.md
Set up multi-region (cross-service) references/configure-multi-region.md
Platform overview / verification references/configure-zone-redundancy.md, references/configure-health-probes.md

Execution order — always quick wins first:

  1. Zone redundancy on compute (fast, in-place property update on the App's plan).

  2. Health probes (Premium / Dedicated only — in-place; for FC1 / Consumption, follow the consent gate in configure-health-probes.md).

  3. Verify the compute changes succeeded before doing anything else.

  4. ⛔ STOP — Ask about storage upgrade. Compute is now zone-redundant, but storage may still be LRS or GRS. Ask the user explicitly:

    ✅ Compute is now zone-redundant.
    
    To be **fully zone-redundant**, your storage account also needs to be upgraded:
      • stii5trxva2ark4: currently `Standard_LRS` → needs `Standard_ZRS`
    
    ⚠️  This is a live storage redundancy conversion:
       • Takes hours to days depending on data volume
       • Small ongoing cost increase (~$0.01/GB/month more)
       • Only supported for Standard general-purpose v2 accounts
    
    Do you want me to start the storage migration now? (yes / no / later)
    
    • yes → run az storage account update --sku Standard_ZRS (or migration start if needed); poll az storage account show --query sku.name until it reports Standard_ZRS.
    • no / later → leave storage as-is; note in the re-assessment that ZR storage remains a gap.
  5. Multi-region — do NOT auto-run. Handled in Step 3 below as an explicit follow-up after re-assessment.

⚠️ Warning: If the user uses azd up or terraform apply later, CLI-only changes may be overwritten by the IaC definitions. Recommend also patching IaC after CLI fixes.

Path B: Patch IaC

Update the user's Bicep or Terraform files so reliability settings are persistent.

Step 1: Detect IaC type

  1. Look for infra/ folder in project root
  2. If not found, check project root for *.bicep or *.tf files
  3. If still not found, ask user: "Where are your IaC files located?"
  4. Check for *.bicep files → use Bicep patching
  5. Check for *.tf files → use Terraform patching
  6. If both exist, ask user which to patch
  7. If no IaC exists, fall back to Path A (CLI) and inform user

Step 2: Classify each fix by risk level

Fix Risk Level What Happens
Zone redundancy (App plan) 🟢 Safe patch In-place property update on next deploy
Storage LRS → ZRS 🟡 Pre-migration required Live storage migration must complete before the IaC SKU change can deploy. Never bundle with safe patches — use the two-deploy flow in Steps 3–5.
Health check path (Basic/Standard/Premium / Dedicated) 🟢 Safe patch In-place update, but causes app restart
Health check path (FC1 / Consumption) ⚪ Code-only — ask first healthCheckPath is unsupported. Adding a health endpoint requires adding an HTTP-triggered /api/health function to app code. Always ask the user for explicit consent before touching source code. Do not patch IaC.

Step 3: Apply patches in two deploys (quick wins first)

The IaC patching framework (detection, AVM-module guidance, deploy-order rule, storage SKU patch) lives in:

IaC Type Framework reference
Bicep references/iac-patching-bicep.md
Terraform references/iac-patching-terraform.md

The actual per-service compute patches (Function App plan ZR, App Service Plan ZR, etc.) live in the per-service references — load the matching service file from Phase 2 for the exact Bicep / Terraform / AVM snippets. Only Azure Functions and App Service have per-service references in this skill version; Container Apps is out of scope.

Deploy 1 — Quick wins only. Patch the 🟢 Safe items (zone redundancy on the App Service/Function App plan, health probes on Basic/Standard/Premium / Dedicated). Do NOT include the storage SKU patch in this deploy.

After patching, the skill runs the deploy itself (do not stop and tell the user to run it). Detect the deployment tool and confirm once before executing:

📦 Patches applied to your IaC. Ready to deploy:
   Tool detected: azd (found azure.yaml)
   Command:       azd up

Proceed with deployment? (yes / no)

On yes, run the appropriate command, stream output back to the user, and continue to the next step on success:

  • AZD project (has azure.yaml): azd up
  • Bicep-only: az deployment group create --resource-group <rg> --template-file infra/main.bicep --parameters @infra/main.parameters.json
  • Terraform: terraform plan -out tfplan → (show plan summary) → terraform apply tfplan

On no, stop and report the patched files; do not proceed to Step 4 / Re-Assess.

If deployment fails, surface the error and stop — do not continue to the storage step.

⛔ STOP — Ask about storage upgrade before Deploy 2. After Deploy 1 succeeds, ask the user explicitly:

✅ Quick-win patches deployed. Compute is now zone-redundant.

To be **fully zone-redundant**, your storage account also needs to be upgraded:
  • stii5trxva2ark4: currently `Standard_LRS` → needs `Standard_ZRS`

⚠️  This is a two-part change:
   1. Live storage migration (`az storage account migration start`) — takes hours to days
   2. A second deploy to update your IaC's storage SKU to match

Do you want me to start the storage migration now? (yes / no / later)
  • yes → the skill runs the migration command itself, polls until complete, then patches the storage SKU in IaC and runs Deploy 2 (now a no-op confirmation). The user does not need to run anything manually.
  • no / later → leave the storage SKU patch unapplied. Note in the re-assessment that ZR storage remains a gap; suggest revisiting later.

Step 4: Storage migration (only if user said yes in Step 3)

The skill runs these commands itself — do not ask the user to run them. Show progress as you go:

🔄 Starting storage migration (this can take up to 72 hours)...

   az storage account migration start --name stii5trxva2ark4 \
     --resource-group rg-example --sku Standard_ZRS --no-wait

   Polling: az storage account show --name stii5trxva2ark4 --query sku.name
   ...
   ✅ Migration complete: sku.name = Standard_ZRS

For very long migrations, you may surface a checkpoint to the user ("this is still running, check back later") rather than blocking the entire conversation.

Step 5: Deploy 2 — storage SKU patch

After the migration completes, the skill patches the storage SKU in IaC and runs the same deploy command as Step 3 (e.g. azd up). This deploy is a no-op confirmation that the IaC matches the live state. Confirm once with the user before executing, then run it directly.

Step 2 (both paths): Re-Assess

After changes are applied (CLI) or deployed (IaC), automatically re-run the assessment and show the same feature-pivoted table as Phase 3, with each feature row's status updated to reflect the new state. Briefly call out what changed since the previous run.

🔄 Reliability Re-Assessment — rg-eventhubs-python-jan13 (eastus)
───────────────────────────────────────────────────────────────────────────────────────
Reliability Feature              Status      Resources
───────────────────────────────────────────────────────────────────────────────────────
Zone redundancy — compute        🟢 ON       • plan-ii5trxva2ark4 (FC1)              — now ON
                                             • plan-web-ii5trxva2ark4 (P1v3)         — now ON

Zone-redundant storage           🟢 ZRS      • stii5trxva2ark4                       — GRS → ZRS

Health probes                    🟡 PARTIAL  • func-api-ii5trxva2ark4                — still off (FC1, code change declined)
                                             • app-web-ii5trxva2ark4                 — now ON

Multi-region failover            🔴 OFF      • Single region (eastus) only
───────────────────────────────────────────────────────────────────────────────────────

What changed: Function App and App Service plan zone redundancy, storage replication and health probes on App Service.
(Multi-region offered next — see Step 3.)
Step 3 (both paths): Multi-region follow-up — ASK and WAIT

Multi-region is a significant cost/complexity step. Do NOT start it automatically. After re-assessment, only if all core single-region reliability features are 🟢 ON (zone-redundant compute, ZRS/GZRS storage, health probes), explicitly ask the user and wait for their response before doing anything:

🟢 Your app is now fully zone-redundant in {region}.

The next step (optional) is multi-region failover with Azure Front Door:
   • Deploys compute + storage in a second region (paired region recommended)
   • Adds Azure Front Door for global load balancing with health-probe-driven failover
   • Protects against full region outages
   • Estimated additional cost: ~2x compute (active-passive); Front Door ~$35/month base

Do you want me to set up multi-region failover now? (yes / no / later)
  • yes → proceed with references/configure-multi-region.md. Confirm secondary region choice with the user, then:
    1. Generate the multi-region IaC (Bicep / Terraform additions for the secondary region + Front Door).
    2. Confirm once with the user: 📦 Multi-region IaC generated. Ready to deploy with \azd up`. Proceed? (yes / no)`
    3. On yes, the skill runs the deploy itself (azd up / az deployment group create / terraform apply) and streams output. Do not stop and tell the user to run it.
    4. After successful deploy, run a final re-assessment so the user sees Multi-region failover flip to 🟢 ON.
  • no / later → leave the deployment as-is. Note that single-region zone-redundant is a reliable end state; multi-region can be revisited anytime.

⛔ Do not skip the wait. Do not generate multi-region IaC, deploy a Front Door, or modify any files until the user has explicitly said yes. If core reliability is not yet all 🟢, do not ask about multi-region — finish the core gaps first.

Priority Classification

Priority Criteria Action
Critical No zone redundancy AND production workload Fix immediately
High LRS storage on zone-redundant compute Fix within days
Medium No multi-region (single region but zone-redundant) Plan for next sprint
Low Missing health probes or monitoring gaps Track and fix

Error Handling

Error Message Remediation
Authentication required "Please login" Run az login and retry
Access denied "Forbidden" Confirm Reader/Contributor role assignment
Plan doesn't support ZR "Upgrade required" Inform user of plan upgrade path + cost delta
Region doesn't support AZ "Region limitation" Suggest supported regions

Best Practices

  • Run reliability assessments after every significant infrastructure change
  • Test failover scenarios periodically (at least quarterly)

Skill Boundaries

Action This skill does Hand off to
Assess reliability posture ✅ Yes —
Recommend improvements ✅ Yes —
Enable zone redundancy (CLI commands) ✅ Yes —
Patch Bicep/Terraform for reliability ✅ Yes —
Generate multi-region IaC ✅ Yes (additions for the secondary region + Front Door) azure-prepare for full new-app IaC scaffolding
Deploy IaC for reliability changes ✅ Yes (runs azd up / terraform apply / az deployment itself, after user confirmation) azure-deploy for general/non-reliability deploys
Validate pre-deployment Reliability checks only azure-validate for full validation
1---
2name: azure-reliability
3description: "Assess and improve the reliability posture of PaaS Applications (Azure Functions and Azure App Service). Scans deployed resources for zone redundancy, ZRS storage, health probes, and multi-region failover. Presents a feature-pivoted checklist, then drives staged remediation (CLI or IaC patches) end-to-end with user confirmation. WHEN: \"assess reliability\", \"check reliability\", \"zone redundant\", \"multi-region failover\", \"high availability\", \"disaster recovery\", \"single points of failure\", \"reliability posture\", \"resiliency\"."
4license: MIT
5metadata:
6 author: Microsoft
7 version: "1.1.2"
8---
9 
10# Azure Reliability Assessment & Configuration
11 
12## Quick Reference
13 
14| Property | Details |
15|---|---|
16| Best for | Reliability posture assessment, zone redundancy enablement, multi-region failover setup |
17| Primary capabilities | Reliability assessment table, Zone Redundancy Configuration, Multi-Region IaC Generation |
18| Supported services | Azure Functions, App Service (Container Apps planned for a future version) |
19| MCP tools | Azure Resource Graph queries, Azure CLI commands |
20 
21## When to Use This Skill
22 
23Activate this skill when user wants to:
24- "Assess my Function app's reliability"
25- "Assess my Web app's reliability"
26- "Check the reliability of my resource group" (App Service and Functions resources only)
27- "Is my app zone redundant?" (App Service and Functions resources only)
28- "Is my app service plan zone redundant?"
29- "Make my app zone redundant" (App Service and Functions resources only)
30- "Make my app service plan zone redundant"
31- "Set up multi-region failover for my app" (App Service and Functions resources only)
32- "Check my reliability posture"
33- "Find single points of failure" (App Service and Functions resources only)
34- "Enable high availability for my app" (App Service and Functions resources only)
35- "Check disaster recovery readiness"
36- "Improve my app's resilience" (App Service and Functions resources only)
37 
38> **Scope note:** This skill currently covers **Azure Functions and Azure App Service** only. If the user asks about Azure Container Apps reliability, acknowledge that support is planned but not yet available, and only proceed with the parts that apply to App Service and Functions resources in scope.
39 
40## Prerequisites
41 
42- Authentication: user is logged in to Azure via `az login`
43- Permissions: Reader access on target subscription/resource group (for assessment)
44- Permissions: Contributor access (for configuration changes)
45- Azure Resource Graph extension: `az extension add --name resource-graph`
46 
47## MCP Tools
48 
49| Tool | Purpose |
50|------|---------|
51| `mcp_azure_mcp_extension_cli_generate` | Generate `az` CLI commands for resource queries and configuration |
52| `mcp_azure_mcp_subscription_list` | List available subscriptions |
53| `mcp_azure_mcp_group_list` | List resource groups |
54 
55Primary query method: Azure Resource Graph via `az graph query` (requires `az extension add --name resource-graph`).
56 
57## Assessment Workflow
58 
59### Phase 1: Discover Resources
60 
611. **Identify scope** — Ask user for resource group, subscription, or app name
622. **Query Azure Resource Graph** to discover all resources in scope
633. **Classify resources** by service type (Functions, Storage, etc.). If Container Apps is found, **note it but do not deep-dive**.
64 
65**Important:** Always scope queries to the user's specified resource group or subscription. Add these filters to every Resource Graph query:
66- Resource group: `| where resourceGroup =~ '<rg-name>'`
67- Subscription: Use `--subscriptions <sub-id>` flag on `az graph query`
68- App name: `| where name =~ '<app-name>'`
69 
70### Phase 2: Assess Reliability
71 
72Two-step assessment: **platform-level discovery first, then per-service deep dive.**
73 
74**Step 1 — Platform discovery (find what's there).** Use these to enumerate resources in scope and detect cross-cutting reliability gaps:
75 
76| Platform check | Reference |
77|---|---|
78| Zone redundancy — discovery | [references/zone-redundancy-checks.md](references/zone-redundancy-checks.md) |
79| Storage redundancy (cross-service) | [references/storage-redundancy-checks.md](references/storage-redundancy-checks.md) |
80| Multi-region & global load balancers | [references/multi-region-checks.md](references/multi-region-checks.md) |
81| Front Door / Traffic Manager / App Insights probes | [references/health-probe-checks.md](references/health-probe-checks.md) |
82 
83**Step 2 — Per-service deep dive.** For each compute resource discovered in Step 1, load the matching service reference. The service reference is the single source of truth for that service's plan/SKU rules, assessment queries, CLI commands, IaC patches (Bicep + Terraform + AVM), and reporting hints.
84 
85This skill version ships **only the Azure Functions and App Service** per-service references. Other compute services are listed below explicitly so the dispatch logic is unambiguous: if a resource matches an unsupported row, do **not** attempt to load a reference, fabricate CLI commands, or generate IaC patches for it.
86 
87| Service detected | Reference |
88|---|---|
89| Azure Functions (`microsoft.web/serverfarms` with `kind contains 'functionapp'`) | [references/services/functions/reliability.md](references/services/functions/reliability.md) |
90| Azure App Service (non-Functions sites: `microsoft.web/sites` without `kind contains 'functionapp'`, `microsoft.web/serverfarms` without `kind contains 'functionapp'`) | [references/services/app-service/reliability.md](references/services/app-service/reliability.md) |
91| Azure Container Apps (`microsoft.app/containerapps`, `microsoft.app/managedenvironments`) | ⚪ Not yet shipped — planned for a future version |
92 
93> **Handling unsupported services:** If a resource matches an unsupported row above, surface it in the discovery summary, mark it as `⚪ not assessed (planned)` in the Phase 3 table, and skip the per-service remediation steps for it. Do **not** attempt to fabricate CLI commands or IaC patches for those services.
94 
95### Phase 3: Generate Reliability Checklist
96 
97Present findings as a **feature-pivoted** table: one row per reliability feature (Zone redundancy on compute, Zone-redundant storage, Health probes, Multi-region failover), with a single status indicator and the **specific resources** that are relevant to that feature. This avoids the noise of one-row-per-resource with mostly `n/a` cells. Do **not** assign numeric scores or grades.
98 
99```
100🔍 Reliability Assessment — {scope}
101─────────────────────────────────────────────────────────────────────────────────────────────
102Reliability Feature Status Resources
103─────────────────────────────────────────────────────────────────────────────────────────────
104Zone redundancy — compute 🔴 OFF • plan-web-ii5trxva2ark4 (P1v3)
105 • plan-ii5trxva2ark4 (FC1)
106 
107Zone-redundant storage 🔴 GRS • stii5trxva2ark4 (defaulted; no SKU set in IaC)
108 
109Health probes 🔴 OFF • func-api-ii5trxva2ark4 — needs code change (FC1)
110 • app-web-ii5trxva2ark4 — no health check path
111 
112Multi-region failover 🔴 OFF • Single region (eastus) only — Front Door not configured
113─────────────────────────────────────────────────────────────────────────────────────────────
114 
115Want me to fix the 🔴 items? I'll do the quick wins first (App
116plan zone redundancy + health checks on supported plans), then ask before
117storage migration and multi-region setup. (yes/no)
118```
119 
120**Rules for the table:**
121 
122- **Four feature rows, in this order:** Zone redundancy — compute · Zone-redundant storage · Health probes · Multi-region failover. Omit a row entirely only if no resource in scope could ever apply to it.
123- **Status column** is one symbol + one short word, no other characters:
124 - `🟢 ON` — feature is fully enabled across all relevant resources in scope
125 - `🟡 PARTIAL` — some resources have it, some don't (or partial config like liveness-only)
126 - `🔴 OFF` — feature is missing on all relevant resources
127 - For storage, replace `OFF` with the current SKU when relevant (`🔴 LRS`, `🔴 GRS`, `🟢 ZRS`, `🟢 GZRS`). When no SKU is set in IaC, label as `🔴 GRS` (ARM/AVM default) and note that in the resource line.
128- **Resources column** lists only what's relevant to that feature, one bullet per resource:
129 - For "needs fixing" resources, include a short inline reason (`(FC1)`, `(defaulted; no SKU set)`, `liveness only`, `needs code change (FC1)`).
130 - For resources that are **already ON** for that feature, mention them on the same row with `— already ON` so the user sees credit for what's right.
131- **Do not** include `n/a`, `—`, or empty cells. If a feature doesn't apply to any resource in scope, drop the row.
132- **Do not** include numeric scores, grades, or point totals.
133- End the assessment with a **single yes/no question** that kicks off the staged remediation flow. Do not enumerate the per-resource fix list here — the user will see it after they say yes (Configuration Workflow Step 1).
134 
135> **UX Note:** If the assessment finds the app **already has** all core reliability features (zone redundancy, ZRS/GZRS storage, health probes), skip the fix-it question and jump straight to Configuration Workflow [Step 3](#step-3-both-paths-multi-region-followup--ask-and-wait) (Multi-region follow-up). Do **NOT** start any multi-region work without explicit consent.
136 
137## Configuration Workflow
138 
139When user wants to **fix** findings from the assessment:
140 
141> **⛔ ALWAYS confirm with user before executing changes.** Show what will change, any cost implications, and any destructive actions (e.g., environment recreation).
142 
143### Step 1: Present Fix Plan + Choose Path
144 
145After assessment, if user says "fix it" / "improve my reliability" / "enable zone redundancy":
146 
1471. List each fixable finding with the specific action
1482. Flag any cost implications or breaking changes
1493. **Ask user which path they want:**
150 
151```
152I'll start with the quick wins (no downtime, fast):
153 
1541. ✏️ Enable zone redundancy on plan-ii5trxva2ark4 (Flex Consumption — no cost change)
1552. ✏️ Set health check path to /api/health on func-api-ii5trxva2ark4
156 
157Then, separately, I'll ask if you want to upgrade storage:
158 
1593. 🕒 Upgrade stii5trxva2ark4 from LRS → ZRS (small cost increase, migration takes hours)
160 — Required for full zone redundancy, but I'll confirm with you before starting.
161 
162How would you like to apply these changes?
163 
164 A) Fix now — Run az CLI commands against your live resources (immediate, one-time)
165 B) Patch my IaC — Update your Bicep/Terraform files so changes persist across deploys
166 
167(If you use azd or Terraform, option B is recommended so `azd up` won't overwrite changes.)
168```
169 
170### Path A: Fix Now (CLI)
171 
172Run fixes against live resources using `az` CLI commands. **Quick wins first, then ask before the slow storage migration.**
173 
174The exact CLI commands per service live in the per-service references — pick the one(s) matching the resources discovered in Phase 2:
175 
176| Fix | Reference |
177|---|---|
178| Enable zone redundancy / configure health probes (Functions) | [references/services/functions/reliability.md](references/services/functions/reliability.md) |
179| Enable zone redundancy / configure health probes (App Service) | [references/services/app-service/reliability.md](references/services/app-service/reliability.md) |
180| Upgrade storage replication (cross-service) | [references/configure-storage.md](references/configure-storage.md) |
181| Set up multi-region (cross-service) | [references/configure-multi-region.md](references/configure-multi-region.md) |
182| Platform overview / verification | [references/configure-zone-redundancy.md](references/configure-zone-redundancy.md), [references/configure-health-probes.md](references/configure-health-probes.md) |
183 
184**Execution order — always quick wins first:**
185 
1861. **Zone redundancy on compute** (fast, in-place property update on the App's plan).
1872. **Health probes** (Premium / Dedicated only — in-place; for FC1 / Consumption, follow the consent gate in [configure-health-probes.md](references/configure-health-probes.md)).
1883. **Verify** the compute changes succeeded before doing anything else.
1894. **⛔ STOP — Ask about storage upgrade.** Compute is now zone-redundant, but storage may still be LRS or GRS. Ask the user explicitly:
190 
191 ```
192 ✅ Compute is now zone-redundant.
193 
194 To be **fully zone-redundant**, your storage account also needs to be upgraded:
195 • stii5trxva2ark4: currently `Standard_LRS` → needs `Standard_ZRS`
196 
197 ⚠️ This is a live storage redundancy conversion:
198 • Takes hours to days depending on data volume
199 • Small ongoing cost increase (~$0.01/GB/month more)
200 • Only supported for Standard general-purpose v2 accounts
201 
202 Do you want me to start the storage migration now? (yes / no / later)
203 ```
204 
205 - **yes** → run `az storage account update --sku Standard_ZRS` (or `migration start` if needed); poll `az storage account show --query sku.name` until it reports `Standard_ZRS`.
206 - **no / later** → leave storage as-is; note in the re-assessment that ZR storage remains a gap.
207 
2085. **Multi-region** — do NOT auto-run. Handled in **Step 3** below as an explicit follow-up after re-assessment.
209 
210> **⚠️ Warning:** If the user uses `azd up` or `terraform apply` later, CLI-only changes may be overwritten by the IaC definitions. Recommend also patching IaC after CLI fixes.
211 
212### Path B: Patch IaC
213 
214Update the user's Bicep or Terraform files so reliability settings are persistent.
215 
216**Step 1: Detect IaC type**
2171. Look for `infra/` folder in project root
2182. If not found, check project root for `*.bicep` or `*.tf` files
2193. If still not found, ask user: "Where are your IaC files located?"
2204. Check for `*.bicep` files → use Bicep patching
2215. Check for `*.tf` files → use Terraform patching
2226. If both exist, ask user which to patch
2237. If no IaC exists, fall back to Path A (CLI) and inform user
224 
225**Step 2: Classify each fix by risk level**
226 
227| Fix | Risk Level | What Happens |
228|-----|-----------|--------------|
229| Zone redundancy (App plan) | 🟢 Safe patch | In-place property update on next deploy |
230| Storage LRS → ZRS | 🟡 Pre-migration required | Live storage migration must complete before the IaC SKU change can deploy. **Never bundle with safe patches** — use the two-deploy flow in Steps 3–5. |
231| Health check path (Basic/Standard/Premium / Dedicated) | 🟢 Safe patch | In-place update, but causes app restart |
232| Health check path (FC1 / Consumption) | ⚪ Code-only — ask first | `healthCheckPath` is unsupported. Adding a health endpoint requires adding an HTTP-triggered `/api/health` function to **app code**. **Always ask the user for explicit consent before touching source code.** Do **not** patch IaC. |
233 
234**Step 3: Apply patches in two deploys (quick wins first)**
235 
236The IaC patching framework (detection, AVM-module guidance, deploy-order rule, storage SKU patch) lives in:
237 
238| IaC Type | Framework reference |
239|---|---|
240| Bicep | [references/iac-patching-bicep.md](references/iac-patching-bicep.md) |
241| Terraform | [references/iac-patching-terraform.md](references/iac-patching-terraform.md) |
242 
243The actual **per-service compute patches** (Function App plan ZR, App Service Plan ZR, etc.) live in the per-service references — load the matching service file from Phase 2 for the exact Bicep / Terraform / AVM snippets. Only Azure Functions and App Service have per-service references in this skill version; Container Apps is out of scope.
244 
245**Deploy 1 — Quick wins only.** Patch the 🟢 Safe items (zone redundancy on the App Service/Function App plan, health probes on Basic/Standard/Premium / Dedicated). Do **NOT** include the storage SKU patch in this deploy.
246 
247After patching, **the skill runs the deploy itself** (do not stop and tell the user to run it). Detect the deployment tool and confirm once before executing:
248 
249```
250📦 Patches applied to your IaC. Ready to deploy:
251 Tool detected: azd (found azure.yaml)
252 Command: azd up
253 
254Proceed with deployment? (yes / no)
255```
256 
257On **yes**, run the appropriate command, stream output back to the user, and continue to the next step on success:
258- AZD project (has `azure.yaml`): `azd up`
259- Bicep-only: `az deployment group create --resource-group <rg> --template-file infra/main.bicep --parameters @infra/main.parameters.json`
260- Terraform: `terraform plan -out tfplan` → (show plan summary) → `terraform apply tfplan`
261 
262On **no**, stop and report the patched files; do not proceed to Step 4 / Re-Assess.
263 
264If deployment fails, surface the error and stop — do not continue to the storage step.
265 
266**⛔ STOP — Ask about storage upgrade before Deploy 2.** After Deploy 1 succeeds, ask the user explicitly:
267 
268```
269✅ Quick-win patches deployed. Compute is now zone-redundant.
270 
271To be **fully zone-redundant**, your storage account also needs to be upgraded:
272 • stii5trxva2ark4: currently `Standard_LRS` → needs `Standard_ZRS`
273 
274⚠️ This is a two-part change:
275 1. Live storage migration (`az storage account migration start`) — takes hours to days
276 2. A second deploy to update your IaC's storage SKU to match
277 
278Do you want me to start the storage migration now? (yes / no / later)
279```
280 
281- **yes** → the skill runs the migration command itself, polls until complete, then patches the storage SKU in IaC and runs **Deploy 2** (now a no-op confirmation). The user does not need to run anything manually.
282- **no / later** → leave the storage SKU patch unapplied. Note in the re-assessment that ZR storage remains a gap; suggest revisiting later.
283 
284**Step 4: Storage migration (only if user said yes in Step 3)**
285 
286The skill runs these commands itself — do not ask the user to run them. Show progress as you go:
287 
288```
289🔄 Starting storage migration (this can take up to 72 hours)...
290 
291 az storage account migration start --name stii5trxva2ark4 \
292 --resource-group rg-example --sku Standard_ZRS --no-wait
293 
294 Polling: az storage account show --name stii5trxva2ark4 --query sku.name
295 ...
296 ✅ Migration complete: sku.name = Standard_ZRS
297```
298 
299For very long migrations, you may surface a checkpoint to the user ("this is still running, check back later") rather than blocking the entire conversation.
300 
301**Step 5: Deploy 2 — storage SKU patch**
302 
303After the migration completes, the skill patches the storage SKU in IaC and runs the same deploy command as Step 3 (e.g. `azd up`). This deploy is a no-op confirmation that the IaC matches the live state. Confirm once with the user before executing, then run it directly.
304 
305### Step 2 (both paths): Re-Assess
306 
307After changes are applied (CLI) or deployed (IaC), automatically re-run the assessment and show the **same feature-pivoted table** as Phase 3, with each feature row's status updated to reflect the new state. Briefly call out what changed since the previous run.
308 
309```
310🔄 Reliability Re-Assessment — rg-eventhubs-python-jan13 (eastus)
311───────────────────────────────────────────────────────────────────────────────────────
312Reliability Feature Status Resources
313───────────────────────────────────────────────────────────────────────────────────────
314Zone redundancy — compute 🟢 ON • plan-ii5trxva2ark4 (FC1) — now ON
315 • plan-web-ii5trxva2ark4 (P1v3) — now ON
316 
317Zone-redundant storage 🟢 ZRS • stii5trxva2ark4 — GRS → ZRS
318 
319Health probes 🟡 PARTIAL • func-api-ii5trxva2ark4 — still off (FC1, code change declined)
320 • app-web-ii5trxva2ark4 — now ON
321 
322Multi-region failover 🔴 OFF • Single region (eastus) only
323───────────────────────────────────────────────────────────────────────────────────────
324 
325What changed: Function App and App Service plan zone redundancy, storage replication and health probes on App Service.
326(Multi-region offered next — see Step 3.)
327```
328 
329### Step 3 (both paths): Multi-region follow-up — ASK and WAIT
330 
331Multi-region is a significant cost/complexity step. Do **NOT** start it automatically. After re-assessment, only if **all core single-region reliability features are 🟢 ON** (zone-redundant compute, ZRS/GZRS storage, health probes), explicitly ask the user and **wait for their response** before doing anything:
332 
333```
334🟢 Your app is now fully zone-redundant in {region}.
335 
336The next step (optional) is multi-region failover with Azure Front Door:
337 • Deploys compute + storage in a second region (paired region recommended)
338 • Adds Azure Front Door for global load balancing with health-probe-driven failover
339 • Protects against full region outages
340 • Estimated additional cost: ~2x compute (active-passive); Front Door ~$35/month base
341 
342Do you want me to set up multi-region failover now? (yes / no / later)
343```
344 
345- **yes** → proceed with [references/configure-multi-region.md](references/configure-multi-region.md). Confirm secondary region choice with the user, then:
346 1. Generate the multi-region IaC (Bicep / Terraform additions for the secondary region + Front Door).
347 2. Confirm once with the user: `📦 Multi-region IaC generated. Ready to deploy with \`azd up\`. Proceed? (yes / no)`
348 3. On **yes**, **the skill runs the deploy itself** (`azd up` / `az deployment group create` / `terraform apply`) and streams output. Do not stop and tell the user to run it.
349 4. After successful deploy, run a final re-assessment so the user sees Multi-region failover flip to 🟢 ON.
350- **no / later** → leave the deployment as-is. Note that single-region zone-redundant is a reliable end state; multi-region can be revisited anytime.
351 
352> **⛔ Do not skip the wait.** Do not generate multi-region IaC, deploy a Front Door, or modify any files until the user has explicitly said yes. If core reliability is not yet all 🟢, do **not** ask about multi-region — finish the core gaps first.
353 
354## Priority Classification
355 
356| Priority | Criteria | Action |
357|---|---|---|
358| Critical | No zone redundancy AND production workload | Fix immediately |
359| High | LRS storage on zone-redundant compute | Fix within days |
360| Medium | No multi-region (single region but zone-redundant) | Plan for next sprint |
361| Low | Missing health probes or monitoring gaps | Track and fix |
362 
363## Error Handling
364 
365| Error | Message | Remediation |
366|---|---|---|
367| Authentication required | "Please login" | Run `az login` and retry |
368| Access denied | "Forbidden" | Confirm Reader/Contributor role assignment |
369| Plan doesn't support ZR | "Upgrade required" | Inform user of plan upgrade path + cost delta |
370| Region doesn't support AZ | "Region limitation" | Suggest supported regions |
371 
372## Best Practices
373 
374- Run reliability assessments after every significant infrastructure change
375- Test failover scenarios periodically (at least quarterly)
376 
377## Skill Boundaries
378 
379| Action | This skill does | Hand off to |
380|---|---|---|
381| Assess reliability posture | ✅ Yes | — |
382| Recommend improvements | ✅ Yes | — |
383| Enable zone redundancy (CLI commands) | ✅ Yes | — |
384| Patch Bicep/Terraform for reliability | ✅ Yes | — |
385| Generate multi-region IaC | ✅ Yes (additions for the secondary region + Front Door) | `azure-prepare` for full new-app IaC scaffolding |
386| Deploy IaC for reliability changes | ✅ Yes (runs `azd up` / `terraform apply` / `az deployment` itself, after user confirmation) | `azure-deploy` for general/non-reliability deploys |
387| Validate pre-deployment | Reliability checks only | `azure-validate` for full validation |
388 

Discussion

Alternatives

Azure app onboardEnd-to-end orchestrator: from a business idea, app idea, or existing app to running Azure deployment with cost estimates and pre-deploy approval. Analyzes your app, auto-detects the right Azure services, scaffolds infrastructure code, and deploys — tailored to your app, not a template. Handles moving existing apps to Azure without rewriting or with minimal changes. WHEN: bring your app to Azure, plan my app, cost to run, is my code ready to deploy, deploy my app to the cloud, deploy all my services, what Azure services do I need, plan my Azure deployment, deploy my new app to Azure, one-click deploy, I have an app and want it on Azure, migrate my app to Azure, help me get started, build an app, no code yet, starter project. DO NOT USE FOR: use azd for deployment(use azure-deploy), optimizing existing costs (use cost-optimization), code readiness checks only (use azure-app-onboard-prereq).Infrastructure & ops · MITAzure App Onboard Prereq — Repository EvaluationAssess whether source code is ready to deploy to Azure — the check BEFORE infrastructure work. Evaluates build health, app completeness, dependencies and local services, stack compatibility, and deployment feasibility. Answers questions about what your app needs before it can be deployed — frameworks, dependencies, and configuration. Checks whether dependencies are compatible and identifies deployment blockers and unsupported frameworks. WHEN: "evaluate my repo", "is my app ready to deploy", "what does my app need to deploy", "what do I need before deploying", "does my app need", "can I ship this to Azure", "scan my repo for issues", "is this app deployable", "check if my app is ready for Azure", "do I need a Dockerfile", "what's blocking my deployment", "are there any blockers", "are my dependencies compatible", "does Azure support my framework", "what needs to change before deploying", "check my app configuration".Infrastructure & ops · MITDocker MCP gatewayDocker's own CLI plugin: run any server from the Docker MCP Catalog in its own container, behind one connection, with secrets kept out of env vars.Coding · MITAzure cloud migrateAssess and migrate cross-cloud workloads to Azure with reports and code conversion. Supports Lambda→Functions, Beanstalk/Heroku/App Engine→App Service, Fargate/Kubernetes/Cloud Run/Spring Boot→Container Apps. WHEN: migrate Lambda to Functions, AWS to Azure, migrate Beanstalk, migrate Heroku, migrate App Engine, Cloud Run migration, Fargate to ACA, ECS/Kubernetes/GKE/EKS to Container Apps, Spring Boot to Container Apps, cross-cloud migration.Infrastructure & ops · MIT