Google Cloud Well-Architected Framework skill for the Cost Optimization pillar

Evaluates Google Cloud workloads for cost efficiency and FinOps alignment using the Google Cloud Well-Architected Framework Cost Optimization pillar.

by google·Apache-2.0 license·★ 20,901 Stars on the repo·GitHub ↗

Use now

Files of Google Cloud Well-Architected Framework skill for the Cost Optimization pillar

google/main1 file shown
SKILL.md
Show the full text181 lines

Google Cloud Well-Architected Framework skill for the Cost Optimization pillar

Overview

The Cost Optimization pillar of the Google Cloud Well-Architected Framework provides a structured approach to optimize the costs of your cloud workloads while maximizing business value. Cloud costs differ significantly from on-premises capital expenditure (CapEx) models, requiring a shift to operational expenditure (OpEx) management and a culture of accountability (FinOps). The FinOps lifecycle consists of three iterative phases:

  • Inform: Visibility and allocation. Always start with Cloud Billing reports for built-in console visibility (filtered by department labels). Complement with Looker Studio for custom, shareable cross-departmental dashboards.

  • Optimize: Rates and usage. Eliminate waste, right-size resources, and leverage commitments (CUDs, SUDs).

  • Operate: Continuous improvement. Integrate cost management into delivery pipelines and establish governance.

Operational instructions

  • Comprehensive First: Always provide a complete list of standard recommendations/strategies relevant to the service or scenario mentioned. Do not limit the scope or omit standard elements just because you are proposing follow-up questions.
  • Full Conjunction: When an instruction bundles multiple actions (e.g., "A AND B"), you must mention and explain BOTH actions in your response explicitly.
  • Conditional Assessment: Only use the 'Workload assessment questions' to refine advice after providing standard framework recommendations, unless the user explicitly asks for an interview or assessment first.

Core principles

The recommendations in the cost optimization pillar of the Well-Architected Framework are aligned with the following core principles:

Relevant Google Cloud products

The following are examples of Google Cloud products and features that are relevant to cost optimization:

  • Visibility and monitoring:

    • Cloud Billing reports: Built-in dashboards for visualizing spending and trends. Essential for visibility within the console.
    • BigQuery billing export: Enables granular, custom analysis of billing data using SQL and BI tools.
    • Looker Studio: Used for creating detailed, shared cost dashboards and reports. Use alongside Cloud Billing reports for custom visual insights.
    • Billing alerts and budgets: Automated notifications when spending reaches predefined thresholds.
    • Storage Insights: Used to analyze Cloud Storage access patterns and identify cost-saving opportunities.
  • Automation and optimization tools:

    • Recommender / Active Assist: Automatically identifies idle resources, rightsizing opportunities, and unused commitments.
    • Cloud Hub Optimization: Integrates billing and resource utilization data to help developers and application owners quickly identify their most expensive, fluctuating, or underutilized cloud resources.
    • FinOps hub: Presents active savings and optimization opportunities in one dashboard.
    • Billing quotas: Limits on resource consumption to prevent unexpected cost spikes.
  • Efficient infrastructure:

    • Managed services and serverless services: Services like Cloud Run, Cloud Run functions, and GKE Autopilot reduce operational overhead and pay-per-use scaling.

    • Compute Engine:

      • Committed Use Discounts (CUDs): Best for predictable, steady-state workloads.
      • Spot VMs: Best for fault-tolerant, interruptible, or unpredictable batch tasks.
      • Sustained Use Discounts (SUDs): Passive, automatic discounts for instances running a significant portion of the month without commitments. Always mention this as a passive alternative or complement to CUDs.
    • Cloud Storage Lifecycle Policies: Automated moves data to lower-cost storage classes (Nearline, Coldline, Archive) based on age or access. Note: Always recommend Storage Insights first to understand current access patterns before defining lifecycle rules.

    • Networking and Content Delivery:

      • Location awareness: Keep traffic within a single region where possible to avoid inter-region data transfer costs.
      • Cloud CDN: Caches content to reduce data egress from the origin.
      • Network Service Tiers: Offers Standard Tier as a lower-cost option compared to Premium Tier for latency-tolerant traffic.
      • Cloud Interconnect / Direct Peering: Optimizes costs for high-volume data transfer to on-premises environments. Note: Always clarify hybrid connectivity options when discussing egress control, as on-premises sync is often a factor in multi-regional designs.
    • Managed Databases:

      • Instance sizing: Right-size CPU and memory based on Cloud Monitoring metrics.
      • High Availability (HA): Implement HA only for production environments to avoid doubling node costs for all environments.
      • Storage Optimization: Optimize storage costs by managing backup retention policies AND explicitly identifying and deleting unused or idle resources (like unused and orphaned disks, expired snapshots, or idle instances).
  • Organization and governance:

    • Resource Manager: Logical structure (Organizations, Folders, Projects) for cost attribution.
    • Labels: Metadata tags for categorizing and filtering costs by environment, team, or application.
    • Organization Policy Service: Enforces constraints (e.g., restricted regions or machine types) to control costs.

Workload assessment questions

Ask appropriate questions to understand the cost-related requirements and constraints of the workload and the user's organization. Choose questions from the following list:

  • How do you incorporate cost considerations into your cloud architecture design process?
  • How do you foster a culture of cost awareness among your development teams?
  • How do you monitor and manage cloud costs across different projects or departments?
  • What strategies do you use to optimize the cost of your compute resources?
  • How do you balance cost optimization with the need for agility and innovation?
  • How do you ensure that you are not over-provisioning cloud resources?
  • How do you use data and analytics to drive cost optimization decisions?
  • How do you optimize costs in different environments (e.g., development, testing, production)?
  • How do you ensure that your cost optimization efforts are sustainable and ongoing?
  • How do you measure the success of your cloud cost optimization initiatives?

Validation checklist

Use the following checklist to evaluate the architecture's alignment with cost-optimization recommendations:

  • Cost Attribution: 100% of resources are labeled with key metadata (e.g., env, team, app).
  • Granular Visibility: BigQuery billing export is enabled and used for regular cost reviews.
  • Budgets and Alerts: Every project or business unit has defined budgets and active alerts.
  • Rightsizing: Resources are regularly adjusted based on rightsizing suggestions provided by Active Assist Recommender.
  • Commitment Strategy: Spend is reviewed monthly to optimize Committed Use Discount (CUD) coverage. For non-committed workloads, verify if Sustained Use Discounts (SUDs) are being captured automatically.
  • Idle Resource Management: Unused disks, IP addresses, and idle VMs are identified and removed monthly.
  • Managed Services: Serverless options are preferred for new workloads unless specific technical constraints exist.
  • Storage Tiers: Lifecycle policies are active for all major storage buckets to minimize archival costs. Be aware of retrieval fees associated with Nearline, Coldline and Archive storage classes.
  • Network Egress: Data transfer is minimized by keeping traffic regional, using Cloud CDN, and leveraging Standard Network Tier where appropriate. High-volume on-premises traffic uses Direct Peering or Cloud Interconnect (always verify if hybrid connectivity is involved).
  • Native Reporting: Cloud Billing reports must be used for standard views of spending trends in the console.
  • Custom Dashboards: Looker Studio is used for advanced, shareable, and customized reporting.
1---
2name: google-cloud-waf-cost-optimization
3metadata:
4 version: "1.0.1"
5 category: WellArchitectedFramework
6description: Evaluates Google Cloud workloads for cost efficiency and FinOps alignment using the Google Cloud Well-Architected Framework Cost Optimization pillar. Use when the user asks to analyze Google Cloud spend, reduce cloud bills, rightsize resources (for example, Compute, GKE, Storage, Databases), evaluate discount options (for example, CUDs, SUDs, Spot VMs), or eliminate idle capacity. Do not use for standalone product pricing lookups or non-cost architecture design.
7---
8 
9# Google Cloud Well-Architected Framework skill for the Cost Optimization pillar
10 
11## Overview
12 
13The Cost Optimization pillar of the Google Cloud Well-Architected Framework
14provides a structured approach to optimize the costs of your cloud workloads
15while maximizing business value. Cloud costs differ significantly from
16on-premises capital expenditure (CapEx) models, requiring a shift to operational
17expenditure (OpEx) management and a culture of accountability (FinOps).
18The FinOps lifecycle consists of three iterative phases:
19- **Inform**: Visibility and allocation. Always start with **Cloud Billing
20 reports** for built-in console visibility (filtered by department labels).
21 Complement with **Looker Studio** for custom, shareable cross-departmental
22 dashboards.
23 
24- **Optimize**: Rates and usage. Eliminate waste, right-size resources, and leverage commitments (**CUDs**, **SUDs**).
25- **Operate**: Continuous improvement. Integrate cost management into delivery pipelines and establish governance.
26 
27## Operational instructions
28 
29- **Comprehensive First**: Always provide a complete list of standard recommendations/strategies relevant to the service or scenario mentioned. Do not limit the scope or omit standard elements just because you are proposing follow-up questions.
30- **Full Conjunction**: When an instruction bundles multiple actions (e.g., "A **AND** B"), you must mention and explain **BOTH** actions in your response explicitly.
31- **Conditional Assessment**: Only use the 'Workload assessment questions' to refine advice *after* providing standard framework recommendations, unless the user explicitly asks for an interview or assessment first.
32 
33## Core principles
34 
35The recommendations in the cost optimization pillar of the Well-Architected
36Framework are aligned with the following core principles:
37 
38- **Align cloud spending with business value**: Ensure that your cloud
39 resources deliver measurable business value by aligning IT spending with
40 business objectives. Prioritize investments that directly contribute to
41 revenue, customer satisfaction, or operational efficiency. Grounding
42 document:
43 https://docs.cloud.google.com/architecture/framework/cost-optimization/align-cloud-spending-business-value.md.txt
44 
45- **Foster a culture of cost awareness**: Ensure that people across your
46 organization consider the cost impact of their decisions and activities.
47 Provide teams with the visibility and information they need to make informed,
48 cost-conscious choices. Grounding document:
49 https://docs.cloud.google.com/architecture/framework/cost-optimization/foster-culture-cost-awareness.md.txt
50 
51- **Optimize resource usage**: Provision only the resources that you need and
52 pay only for what you consume. Select the most cost-effective resource types,
53 sizes, and locations that meet your technical and business requirements.
54 Grounding document:
55 https://docs.cloud.google.com/architecture/framework/cost-optimization/optimize-resource-usage.md.txt
56 
57- **Optimize continuously**: Continuously monitor your cloud resource usage and
58 costs, and proactively make adjustments as needed to optimize your spending.
59 This iterative approach helps identify and address inefficiencies before they
60 become significant. Grounding document:
61 https://docs.cloud.google.com/architecture/framework/cost-optimization/optimize-continuously.md.txt
62 
63## Relevant Google Cloud products
64 
65The following are _examples_ of Google Cloud products and features that are
66relevant to cost optimization:
67 
68- **Visibility and monitoring**:
69 
70 - **Cloud Billing reports**: Built-in dashboards for visualizing spending and
71 trends. **Essential for visibility within the console.**
72 - **BigQuery billing export**: Enables granular, custom analysis of billing
73 data using SQL and BI tools.
74 - **Looker Studio**: Used for creating detailed, shared cost dashboards and
75 reports. **Use alongside Cloud Billing reports for custom visual insights.**
76 - **Billing alerts and budgets**: Automated notifications when spending
77 reaches predefined thresholds.
78 - **Storage Insights**: Used to analyze Cloud Storage access patterns and identify
79 cost-saving opportunities.
80 
81- **Automation and optimization tools**:
82 
83 - **Recommender / Active Assist**: Automatically identifies idle resources,
84 rightsizing opportunities, and unused commitments.
85 - **Cloud Hub Optimization**: Integrates billing and resource utilization data
86 to help developers and application owners quickly identify their most
87 expensive, fluctuating, or underutilized cloud resources.
88 - **FinOps hub**: Presents active savings and optimization opportunities in
89 one dashboard.
90 - **Billing quotas**: Limits on resource consumption to prevent unexpected
91 cost spikes.
92 
93- **Efficient infrastructure**:
94 
95 - **Managed services and serverless services**: Services like Cloud Run, Cloud
96 Run functions, and GKE Autopilot reduce operational overhead and pay-per-use
97 scaling.
98 - **Compute Engine**:
99 - **Committed Use Discounts (CUDs)**: Best for predictable, steady-state workloads.
100 - **Spot VMs**: Best for fault-tolerant, interruptible, or unpredictable batch tasks.
101 - **Sustained Use Discounts (SUDs)**: Passive, automatic discounts for instances running a significant portion of the month without commitments. **Always mention this as a passive alternative or complement to CUDs.**
102 
103 - **Cloud Storage Lifecycle Policies**: Automated moves data to lower-cost
104 storage classes (Nearline, Coldline, Archive) based on age or access.
105 **Note**: Always recommend **Storage Insights** first to understand current
106 access patterns before defining lifecycle rules.
107 
108 - **Networking and Content Delivery**:
109 - **Location awareness**: Keep traffic within a single region where possible to avoid inter-region data transfer costs.
110 - **Cloud CDN**: Caches content to reduce data egress from the origin.
111 - **Network Service Tiers**: Offers **Standard Tier** as a lower-cost option compared to **Premium Tier** for latency-tolerant traffic.
112 - **Cloud Interconnect / Direct Peering**: Optimizes costs for high-volume data transfer to on-premises environments. **Note**: Always clarify hybrid connectivity options when discussing egress control, as on-premises sync is often a factor in multi-regional designs.
113 - **Managed Databases**:
114 - **Instance sizing**: Right-size CPU and memory based on Cloud Monitoring metrics.
115 - **High Availability (HA)**: Implement HA only for production environments to avoid doubling node costs for all environments.
116 - **Storage Optimization**: Optimize storage costs by managing backup retention policies **AND** explicitly identifying and deleting unused or idle resources (like unused and orphaned disks, expired snapshots, or idle instances).
117 
118 
119- **Organization and governance**:
120 
121 - **Resource Manager**: Logical structure (Organizations, Folders, Projects)
122 for cost attribution.
123 - **Labels**: Metadata tags for categorizing and filtering costs by
124 environment, team, or application.
125 - **Organization Policy Service**: Enforces constraints (e.g., restricted
126 regions or machine types) to control costs.
127 
128## Workload assessment questions
129 
130Ask appropriate questions to understand the cost-related requirements and
131constraints of the workload and the user's organization. Choose questions from
132the following list:
133 
134- How do you incorporate cost considerations into your cloud architecture design
135 process?
136- How do you foster a culture of cost awareness among your development teams?
137- How do you monitor and manage cloud costs across different projects or
138 departments?
139- What strategies do you use to optimize the cost of your compute resources?
140- How do you balance cost optimization with the need for agility and innovation?
141- How do you ensure that you are not over-provisioning cloud resources?
142- How do you use data and analytics to drive cost optimization decisions?
143- How do you optimize costs in different environments (e.g., development,
144 testing, production)?
145- How do you ensure that your cost optimization efforts are sustainable and
146 ongoing?
147- How do you measure the success of your cloud cost optimization initiatives?
148 
149## Validation checklist
150 
151Use the following checklist to evaluate the architecture's alignment with
152cost-optimization recommendations:
153 
154- **Cost Attribution**: 100% of resources are labeled with key metadata
155 (e.g., `env`, `team`, `app`).
156- **Granular Visibility**: BigQuery billing export is enabled and used for
157 regular cost reviews.
158- **Budgets and Alerts**: Every project or business unit has defined budgets
159 and active alerts.
160- **Rightsizing**: Resources are regularly adjusted based on rightsizing
161 suggestions provided by Active Assist Recommender.
162- **Commitment Strategy**: Spend is reviewed monthly to optimize Committed Use Discount (CUD) coverage. For non-committed workloads, verify if **Sustained Use Discounts (SUDs)** are being captured automatically.
163- **Idle Resource Management**: Unused disks, IP addresses, and idle VMs are
164 identified and removed monthly.
165- **Managed Services**: Serverless options are preferred for new workloads
166 unless specific technical constraints exist.
167- **Storage Tiers**: Lifecycle policies are active for all major storage
168 buckets to minimize archival costs. Be aware of **retrieval fees**
169 associated with Nearline, Coldline and Archive storage classes.
170- **Network Egress**: Data transfer is minimized by keeping traffic regional,
171 using **Cloud CDN**, and leveraging **Standard Network Tier** where
172 appropriate. High-volume on-premises traffic uses **Direct Peering** or
173 **Cloud Interconnect** (always verify if hybrid connectivity is involved).
174- **Native Reporting**: **Cloud Billing reports** must be used for standard
175 views of spending trends in the console.
176- **Custom Dashboards**: **Looker Studio** is used for advanced, shareable, and
177 customized reporting.
178 
179 
180 
181 

Discussion

Alternatives