GKE Golden Path Configuration skill

Provides GKE golden path configuration defaults, production readiness checklists, and cluster default patterns.

by google·Apache-2.0 license·★ 20,901 Stars on the repo·GitHub ↗

Use now

Files of GKE Golden Path Configuration

google/main1 file shown
SKILL.md
Show the full text154 lines

GKE Golden Path Configuration

The golden path is the recommended Autopilot configuration for production clusters. It defines sensible defaults — when the user requests different settings, apply them and note relevant trade-offs. For setting up autoscaling specifically, use gke-cluster-autoscaler for node autoscaling or gke-workload-scaling for workload autoscaling (HPA/VPA).

MCP Tools: get_cluster, create_cluster, update_cluster

Rules

  1. Default to the golden path. Use golden path values unless the user requests otherwise. When deviating, note trade-offs but respect the user's choice.
  2. Day-0 vs Day-1. Flag Day-0 decisions (networking, private nodes, subnets, IP allocation) prominently — they are hard/impossible to change after creation.
  3. Tool preference: MCP > gcloud > kubectl. MCP is preferred as it directly interfaces with GKE APIs with structured data, reducing shell syntax errors and parsing ambiguities. See the gke-basics skill's CLI reference for full coverage matrix and override options. If the user says "use gcloud" or "use kubectl", respect that for the session.
  4. Document decisions and rationale, especially for Day-0 choices and golden path deviations.

Required Inputs

If the user is unsure, use golden path defaults.

  • Project ID (required)
  • Region (required, e.g., us-central1)
  • Cluster name (required)
  • Environment type: dev/test or production (defaults to production)
  • Networking: bring-your-own VPC/subnet or auto-create (default: auto-create)
  • Scale expectations: expected node/pod count, workload types
  • Cost constraints: Spot VM tolerance, budget considerations

Always-Apply Defaults

Recommended best practices applied by default. If the user requests a different setting, apply it and briefly note the security or operational trade-off.

Setting Golden Path Value
autopilot.enabled true
privateClusterConfig.enablePrivateNodes true
masterAuthorizedNetworksConfig.privateEndpointEnforcementEnabled true
secretManagerConfig.enabled + rotationInterval: 120s true
rbacBindingConfig.enableInsecureBinding* false (both)
workloadIdentityConfig.workloadPool enabled
networkConfig.datapathProvider ADVANCED_DATAPATH
networkConfig.dnsConfig.clusterDns CLOUD_DNS
autoscaling.autoscalingProfile OPTIMIZE_UTILIZATION
verticalPodAutoscaling.enabled true
monitoringConfig components SYSTEM_COMPONENTS, STORAGE, POD, DEPLOYMENT, STATEFULSET, DAEMONSET, HPA, JOBSET, CADVISOR, KUBELET, DCGM, APISERVER, SCHEDULER, CONTROLLER_MANAGER
loggingConfig components SYSTEM_COMPONENTS, WORKLOADS (enabled by default)
advancedDatapathObservabilityConfig.enableMetrics true
nodeConfig.shieldedInstanceConfig.enableSecureBoot true
nodeConfig.workloadMetadataConfig.mode GKE_METADATA
nodeConfig.gcfsConfig.enabled / gvnic.enabled true / true
addonsConfig.statefulHaConfig.enabled true
Storage CSI drivers (Filestore, GCS FUSE, Parallelstore) enabled
Pod Security Standards restricted on production namespaces

Customer-Configurable Settings

These have golden path defaults but customers may deviate with valid justification. Ask before changing.

Setting Default Why Deviate
dnsEndpointConfig.allowExternalTraffic true Restrict if cluster only accessed from within VPC
autoIpamConfig / createSubnetwork true / true Customer has pre-existing VPC/subnets
maxPodsPerNode 48 (this golden path's choice) Halves per-node IP consumption (/25 instead of /24). Not a GKE default (Standard defaults to 110, Autopilot to 32); raise for high pod-density at the cost of more CIDR space
subnetwork auto-created Customer brings existing subnets
Release channel + maintenance windows REGULAR channel with a recurring maintenance window Add targeted maintenance exclusions (keep under ~6 months) only for critical freezes — see the gke-upgrades skill
nodeConfig.bootDisk.diskType pd-balanced pd-ssd for I/O-intensive, pd-standard for cost

Note: Autopilot selects node machine types automatically (e.g., ek-standard-8 may appear in describe output); the machine type is not customer-configurable in Autopilot. Steer workload placement via ComputeClasses instead.

Guardrails

  • Do not request or output secrets (tokens, keys, service account JSON).
  • Resolve project/cluster context from the conversation, MCP tools, or gcloud config get-value project; ask the user only if it cannot be resolved.
  • For Day-0 decisions, always ask clarifying questions before proceeding.
  • For Day-1 features, propose golden path defaults with trade-offs and let the customer confirm.
  • Do not promise zero downtime — see Upgrade Disruption below for what to advise instead.
  • When auditing existing clusters, compare against golden path and report deviations with severity and remediation.

Upgrade Disruption

Never promise zero downtime for node upgrades, on any configuration. Node upgrades cordon and drain nodes, which evicts Pods. Draining honors PodDisruptionBudgets and terminationGracePeriodSeconds for up to one hour, after which GKE forcefully evicts the remaining Pods so the upgrade can proceed. A PDB narrows the window; it cannot veto the upgrade. Say so plainly rather than implying the disruption can be eliminated.

What to recommend, all four — not a subset:

  • PodDisruptionBudgets with minAvailable set so eviction cannot take the last healthy replica. A PDB that can never be satisfied stalls the drain for an hour and then loses anyway.

  • At least 2 replicas, spread across zones with topology spread constraints. A single-replica Deployment has downtime by definition.

  • Readiness probes that reflect real serving health, so traffic drains before the Pod dies.

  • Surge upgrade settings on the node pool. Surge is the default strategy; the default is maxSurge=1, maxUnavailable=0 — one extra node is created and made ready before an old one is drained.

    Setting Controls Default
    maxSurge Additional nodes added per zone during the upgrade 1
    maxUnavailable Nodes simultaneously unavailable per zone 0

    Nodes upgraded at once is the sum of the two, capped at 20 (Autopilot) and 100 (Standard). Multi-zone node pools upgrade one zone at a time. Raising maxUnavailable trades availability for speed; raising maxSurge trades cost for availability.

Caveat: externalTrafficPolicy: Local does not work with parallel node drains, so it constrains aggressive surge configurations.

For rollback procedures and maintenance windows, see the gke-upgrades skill.

Golden Path Config

See golden-path-autopilot.yaml for the full cluster-level policy settings.

1---
2name: gke-golden-path
3description: >-
4 Provides GKE golden path configuration defaults, production readiness
5 checklists, and cluster default patterns. Use when designing GKE clusters,
6 verifying GKE production readiness, or checking configurations against
7 GKE defaults. Don't use for setting up workload autoscaling specifically (use
8 gke-workload-scaling instead).
9metadata:
10 version: "1.1.0"
11 category: Containers
12---
13 
14# GKE Golden Path Configuration
15 
16The golden path is the recommended Autopilot configuration for production
17clusters. It defines sensible defaults — when the user requests different
18settings, apply them and note relevant trade-offs. For setting up autoscaling
19specifically, use `gke-cluster-autoscaler` for node autoscaling or
20`gke-workload-scaling` for workload autoscaling (HPA/VPA).
21 
22> **MCP Tools:** `get_cluster`, `create_cluster`, `update_cluster`
23 
24## Rules
25 
261. **Default to the golden path.** Use golden path values unless the user
27 requests otherwise. When deviating, note trade-offs but respect the user's
28 choice.
292. **Day-0 vs Day-1.** Flag Day-0 decisions (networking, private nodes,
30 subnets, IP allocation) prominently — they are hard/impossible to change
31 after creation.
323. **Tool preference: MCP > gcloud > kubectl.** MCP is preferred as it directly
33 interfaces with GKE APIs with structured data, reducing shell syntax errors
34 and parsing ambiguities. See the `gke-basics` skill's CLI reference for full
35 coverage matrix and override options. If the user
36 says "use gcloud" or "use kubectl", respect that for the session.
374. **Document decisions and rationale**, especially for Day-0 choices and
38 golden path deviations.
39 
40## Required Inputs
41 
42If the user is unsure, use golden path defaults.
43 
44- **Project ID** (required)
45- **Region** (required, e.g., `us-central1`)
46- **Cluster name** (required)
47- **Environment type**: dev/test or production (defaults to production)
48- **Networking**: bring-your-own VPC/subnet or auto-create (default:
49 auto-create)
50- **Scale expectations**: expected node/pod count, workload types
51- **Cost constraints**: Spot VM tolerance, budget considerations
52 
53## Always-Apply Defaults
54 
55Recommended best practices applied by default. If the user requests a different
56setting, apply it and briefly note the security or operational trade-off.
57 
58Setting | Golden Path Value
59------------------------------------------------------------------ | -----------------
60`autopilot.enabled` | `true`
61`privateClusterConfig.enablePrivateNodes` | `true`
62`masterAuthorizedNetworksConfig.privateEndpointEnforcementEnabled` | `true`
63`secretManagerConfig.enabled` + `rotationInterval: 120s` | `true`
64`rbacBindingConfig.enableInsecureBinding*` | `false` (both)
65`workloadIdentityConfig.workloadPool` | enabled
66`networkConfig.datapathProvider` | `ADVANCED_DATAPATH`
67`networkConfig.dnsConfig.clusterDns` | `CLOUD_DNS`
68`autoscaling.autoscalingProfile` | `OPTIMIZE_UTILIZATION`
69`verticalPodAutoscaling.enabled` | `true`
70`monitoringConfig` components | SYSTEM_COMPONENTS, STORAGE, POD, DEPLOYMENT, STATEFULSET, DAEMONSET, HPA, JOBSET, CADVISOR, KUBELET, DCGM, APISERVER, SCHEDULER, CONTROLLER_MANAGER
71`loggingConfig` components | SYSTEM_COMPONENTS, WORKLOADS (enabled by default)
72`advancedDatapathObservabilityConfig.enableMetrics` | `true`
73`nodeConfig.shieldedInstanceConfig.enableSecureBoot` | `true`
74`nodeConfig.workloadMetadataConfig.mode` | `GKE_METADATA`
75`nodeConfig.gcfsConfig.enabled` / `gvnic.enabled` | `true` / `true`
76`addonsConfig.statefulHaConfig.enabled` | `true`
77Storage CSI drivers (Filestore, GCS FUSE, Parallelstore) | enabled
78Pod Security Standards | `restricted` on production namespaces
79 
80## Customer-Configurable Settings
81 
82These have golden path defaults but customers may deviate with valid
83justification. **Ask before changing.**
84 
85Setting | Default | Why Deviate
86---------------------------------------- | ----------------------------------- | -----------
87`dnsEndpointConfig.allowExternalTraffic` | `true` | Restrict if cluster only accessed from within VPC
88`autoIpamConfig` / `createSubnetwork` | `true` / `true` | Customer has pre-existing VPC/subnets
89`maxPodsPerNode` | `48` (this golden path's choice) | Halves per-node IP consumption (/25 instead of /24). Not a GKE default (Standard defaults to `110`, Autopilot to `32`); raise for high pod-density at the cost of more CIDR space
90`subnetwork` | auto-created | Customer brings existing subnets
91Release channel + maintenance windows | `REGULAR` channel with a recurring maintenance window | Add targeted maintenance exclusions (keep under ~6 months) only for critical freezes — see the `gke-upgrades` skill
92`nodeConfig.bootDisk.diskType` | `pd-balanced` | `pd-ssd` for I/O-intensive, `pd-standard` for cost
93 
94> **Note**: Autopilot selects node machine types automatically (e.g.,
95> `ek-standard-8` may appear in describe output); the machine type is not
96> customer-configurable in Autopilot. Steer workload placement via
97> ComputeClasses instead.
98 
99## Guardrails
100 
101- Do not request or output secrets (tokens, keys, service account JSON).
102- Resolve project/cluster context from the conversation, MCP tools, or
103 `gcloud config get-value project`; ask the user only if it cannot be
104 resolved.
105- For Day-0 decisions, always ask clarifying questions before proceeding.
106- For Day-1 features, propose golden path defaults with trade-offs and let the
107 customer confirm.
108- Do not promise zero downtime — see Upgrade Disruption below for what to
109 advise instead.
110- When auditing existing clusters, compare against golden path and report
111 deviations with severity and remediation.
112 
113## Upgrade Disruption
114 
115**Never promise zero downtime for node upgrades, on any configuration.** Node
116upgrades cordon and drain nodes, which evicts Pods. Draining honors
117PodDisruptionBudgets and `terminationGracePeriodSeconds` for **up to one hour**,
118after which GKE forcefully evicts the remaining Pods so the upgrade can proceed.
119A PDB narrows the window; it cannot veto the upgrade. Say so plainly rather than
120implying the disruption can be eliminated.
121 
122What to recommend, all four — not a subset:
123 
124- **PodDisruptionBudgets** with `minAvailable` set so eviction cannot take the
125 last healthy replica. A PDB that can never be satisfied stalls the drain for
126 an hour and then loses anyway.
127- **At least 2 replicas**, spread across zones with topology spread
128 constraints. A single-replica Deployment has downtime by definition.
129- **Readiness probes** that reflect real serving health, so traffic drains
130 before the Pod dies.
131- **Surge upgrade settings** on the node pool. Surge is the default strategy;
132 the default is `maxSurge=1`, `maxUnavailable=0` — one extra node is created
133 and made ready before an old one is drained.
134 
135 Setting | Controls | Default
136 ---------------- | --------------------------------------------------- | -------
137 `maxSurge` | Additional nodes added per zone during the upgrade | `1`
138 `maxUnavailable` | Nodes simultaneously unavailable per zone | `0`
139 
140 Nodes upgraded at once is the **sum** of the two, capped at 20 (Autopilot)
141 and 100 (Standard). Multi-zone node pools upgrade one zone at a time. Raising
142 `maxUnavailable` trades availability for speed; raising `maxSurge` trades
143 cost for availability.
144 
145> **Caveat**: `externalTrafficPolicy: Local` does not work with parallel node
146> drains, so it constrains aggressive surge configurations.
147 
148For rollback procedures and maintenance windows, see the `gke-upgrades` skill.
149 
150## Golden Path Config
151 
152See [golden-path-autopilot.yaml](./assets/golden-path-autopilot.yaml) for the
153full cluster-level policy settings.
154 

Discussion