System design methodology skill

Drives an interactive system design session: classifies depth, elicits scale/SLO/consistency inputs, computes capacity, then reveals components one by one, each justified by a constraint.

by HoangNguyen0403·MIT license·★ 569 Stars on the repo·GitHub ↗

Use now

Files of System design methodology

HoangNguyen0403/develop1 file shown
SKILL.md
Show the full text100 lines

System Design Methodology

Priority: P0 (CRITICAL)

Requirements before solutions. Never draw a full architecture before numbers justify it.

Phase 0 - Classify Depth (always first)

  • Quick sketch: exploratory ask, no scale numbers available, answer needed now. Assume defaults, label each one ASSUMED, skip gates.
  • Full session: real build, migration, or budget commitment. Run every phase gate.
  • State depth and mode (new design | review existing | interview practice) in one line, then continue. Interview practice runs through system-design-interview-coaching: the round on a clock, the rubric after.
  • Escalate quick to full when a hard constraint or irreversible choice appears.

Phase 1 - Intake (gate)

  • Parse request: verbs to use cases, nouns to entities, adjectives to constraints.
  • Ask max 3 blocking questions per turn, each with a recommended default. See intake checklist.
  • Required before design: DAU/actors, top 3 use cases, read:write ratio, latency SLO, consistency need, retention, peak shape, budget, team size.
  • Freeze scope: list what is explicitly out of scope.

Phase 2 - Estimation (gate)

  • Compute QPS, storage, bandwidth, and working-set memory via system-design-estimation.
  • Present the numbers, name the one quantity that shapes the design, confirm before drawing anything.

Phase 3 - High-Level Design (incremental)

  • Price the null option first: do nothing, buy it, or let an existing service absorb it. Rejecting it needs a stated reason, not silence.
  • Start with the smallest system satisfying functional requirements: client, API, service, store.
  • Add one component at a time. For each, state constraint -> component -> cost in one line. No component without a named constraint.
  • Define API surface (one endpoint per functional requirement) and data ownership before optimizing.
  • Select views only when they answer a named question, per common-architecture-diagramming: a context/container, sequence, dataflow, deployment, or state view may be used when useful; prose or a table is sufficient otherwise. Carry metric and constraint only when the design states them; never invent a number to populate a node. See phase deliverables.

HLD, LLD, and Low-Level Design Routing

  • HLD answers audience-level boundaries, shaping constraints, ownership, failure domains, and the decision to make. Use context/container or prose only when that is enough; no diagram is mandatory.
  • LLD (the same lane as “low-level design”) answers one component or critical flow: data/state ownership, API or event contracts, ordering, idempotency, failure behavior, and verification. Use sequence, dataflow, or state only when that view resolves a named question.
  • Trace every handoff as requirement -> HLD decision -> component -> LLD contract -> verification. Give each link a stable ID and carry unresolved assumptions forward; an LLD must not silently change the HLD invariant.
  • Each selected view declares audience, question, decision, scenario, invariant, scope, status, evidence, and omissions. Lifecycle is proposed|implemented|retired; evidence is a citation, not confidence. Keep evidence_kind and evidence_confidence separate per the renderer-owned diagram spec and view manifest; never infer deployment from a code/document citation.
  • Views are evidence for a question, not a completeness checklist. Prefer a precise paragraph or table over a diagram that adds no decision value.

Specialist Deep-Dive Contract

  • Have the user pick the 2-3 riskiest components, then send one specialist brief per component with profile, audience/question, workload/SLO/team/budget, invariant, scope, evidence status, and current HLD decision. The specialist does not re-run intake, add neighboring components, or invent numbers.
  • Require options with rejection reasons, the recommended LLD contract, failure timeline/recovery, verification hooks, and any ADR reversal trigger. Merge the result back into the HLD-to-LLD trace before scoring.

Brownfield Path (review-existing mode)

  • Map current state before proposing anything: components, owners, traffic, incidents.
  • Measure, do not assume: pull real QPS, data volume, and p99 from the running system.
  • Find the binding constraint - the one that fails first at the next growth step.
  • Design the smallest change that moves it, then re-measure. A rewrite needs a structural constraint the current shape cannot satisfy.

Phase 4 - Deep Dives and Trade-offs

  • Stage what to build now, the enabling seam and metric threshold; record one ADR per irreversible decision with its reversal trigger, then score with system-design-review.

Design-to-Delivery Gate

  • Once HLD/LLD is fixed, list bounded docs/diagram slices: exact files, evidence, acceptance, verification, integrator. Route production to the cheapest qualified configured executor if available; lead owns decisions and final review.
  • If still defective after one focused correction, use the configured fallback or report BLOCKED. Log executor/model, corrections, exceptions and fallback reason; report actual usage/cost or unavailable, never assumed savings.

Anti-Patterns

  • No architecture before requirements: no diagram until Phase 1 answers exist or defaults are flagged.
  • No unjustified components: every box names the constraint it solves.
  • No design without the null option: state why doing nothing or buying loses before building.
  • No silent assumptions: an unknown input becomes a labeled ASSUMED default, never a hidden guess.
  • No full-stack reveal: never dump a finished diagram before incremental agreement.

Red Flags

  • Stop if "just give me the architecture": deliver a quick sketch with ASSUMED labels, not fake precision.
  • Stop if scale is unknown at Phase 3: return to Phase 2 and estimate from a stated assumption.

References

1---
2name: system-design-methodology
3description: "Drives an interactive system design session: classifies depth, elicits scale/SLO/consistency inputs, computes capacity, then reveals components one by one, each justified by a constraint. Use when designing a system or running a design session; diagrams go through `common-architecture-diagramming`."
4metadata:
5 triggers:
6 keywords:
7 - system design
8 - design a system
9 - design session
10 - high-level design
11 - low-level design
12 - HLD
13 - LLD
14 - requirements clarification
15 - capacity planning
16 - scale this
17---
18 
19# System Design Methodology
20 
21## **Priority: P0 (CRITICAL)**
22 
23Requirements before solutions. Never draw a full architecture before numbers justify it.
24 
25## Phase 0 - Classify Depth (always first)
26 
27- **Quick sketch**: exploratory ask, no scale numbers available, answer needed now. Assume defaults, label each one `ASSUMED`, skip gates.
28- **Full session**: real build, migration, or budget commitment. Run every phase gate.
29- State depth and mode (new design | review existing | interview practice) in one line, then continue. Interview practice runs through `system-design-interview-coaching`: the round on a clock, the rubric after.
30- Escalate quick to full when a hard constraint or irreversible choice appears.
31 
32## Phase 1 - Intake (gate)
33 
34- Parse request: verbs to use cases, nouns to entities, adjectives to constraints.
35- Ask max 3 blocking questions per turn, each with a recommended default. See [intake checklist](references/intake-checklist.md).
36- Required before design: DAU/actors, top 3 use cases, read:write ratio, latency SLO, consistency need, retention, peak shape, budget, team size.
37- Freeze scope: list what is explicitly out of scope.
38 
39## Phase 2 - Estimation (gate)
40 
41- Compute QPS, storage, bandwidth, and working-set memory via `system-design-estimation`.
42- Present the numbers, name the one quantity that shapes the design, confirm before drawing anything.
43 
44## Phase 3 - High-Level Design (incremental)
45 
46- Price the null option first: do nothing, buy it, or let an existing service absorb it. Rejecting it needs a stated reason, not silence.
47- Start with the smallest system satisfying functional requirements: client, API, service, store.
48- Add one component at a time. For each, state `constraint -> component -> cost` in one line. No component without a named constraint.
49- Define API surface (one endpoint per functional requirement) and data ownership before optimizing.
50- Select views only when they answer a named question, per `common-architecture-diagramming`: a context/container, sequence, dataflow, deployment, or state view may be used when useful; prose or a table is sufficient otherwise. Carry `metric` and `constraint` only when the design states them; never invent a number to populate a node. See [phase deliverables](references/phase-deliverables.md).
51 
52## HLD, LLD, and Low-Level Design Routing
53 
54- **HLD** answers audience-level boundaries, shaping constraints, ownership, failure domains, and the decision to make. Use context/container or prose only when that is enough; no diagram is mandatory.
55- **LLD** (the same lane as “low-level design”) answers one component or critical flow: data/state ownership, API or event contracts, ordering, idempotency, failure behavior, and verification. Use sequence, dataflow, or state only when that view resolves a named question.
56- Trace every handoff as `requirement -> HLD decision -> component -> LLD contract -> verification`. Give each link a stable ID and carry unresolved assumptions forward; an LLD must not silently change the HLD invariant.
57- Each selected view declares `audience`, `question`, `decision`, `scenario`, `invariant`, `scope`, `status`, `evidence`, and `omissions`. Lifecycle is `proposed|implemented|retired`; `evidence` is a citation, not confidence. Keep `evidence_kind` and `evidence_confidence` separate per the renderer-owned [diagram spec](../../common/common-architecture-diagramming/references/diagram-spec.md) and [view manifest](../../common/common-architecture-diagramming/references/view-manifest.md); never infer deployment from a code/document citation.
58- Views are evidence for a question, not a completeness checklist. Prefer a precise paragraph or table over a diagram that adds no decision value.
59 
60## Specialist Deep-Dive Contract
61 
62- Have the user pick the 2-3 riskiest components, then send one specialist brief per component with profile, audience/question, workload/SLO/team/budget, invariant, scope, evidence status, and current HLD decision. The specialist does not re-run intake, add neighboring components, or invent numbers.
63- Require options with rejection reasons, the recommended LLD contract, failure timeline/recovery, verification hooks, and any ADR reversal trigger. Merge the result back into the HLD-to-LLD trace before scoring.
64 
65## Brownfield Path (review-existing mode)
66 
67- Map current state before proposing anything: components, owners, traffic, incidents.
68- Measure, do not assume: pull real QPS, data volume, and p99 from the running system.
69- Find the binding constraint - the one that fails first at the next growth step.
70- Design the smallest change that moves it, then re-measure. A rewrite needs a structural constraint the current shape cannot satisfy.
71 
72## Phase 4 - Deep Dives and Trade-offs
73 
74- Stage what to build now, the enabling seam and metric threshold; record one ADR per irreversible decision with its reversal trigger, then score with `system-design-review`.
75 
76## Design-to-Delivery Gate
77 
78- Once HLD/LLD is fixed, list bounded docs/diagram slices: exact files, evidence, acceptance, verification, integrator. Route production to the cheapest qualified configured executor if available; lead owns decisions and final review.
79- If still defective after one focused correction, use the configured fallback or report BLOCKED. Log executor/model, corrections, exceptions and fallback reason; report actual usage/cost or `unavailable`, never assumed savings.
80 
81## Anti-Patterns
82 
83- **No architecture before requirements**: no diagram until Phase 1 answers exist or defaults are flagged.
84- **No unjustified components**: every box names the constraint it solves.
85- **No design without the null option**: state why doing nothing or buying loses before building.
86- **No silent assumptions**: an unknown input becomes a labeled `ASSUMED` default, never a hidden guess.
87- **No full-stack reveal**: never dump a finished diagram before incremental agreement.
88 
89## Red Flags
90 
91- **Stop if "just give me the architecture"**: deliver a quick sketch with `ASSUMED` labels, not fake precision.
92- **Stop if scale is unknown at Phase 3**: return to Phase 2 and estimate from a stated assumption.
93 
94## References
95 
96- [Four-Phase Process](references/four-phase-process.md) - per-phase gates, outputs, escalation rules
97- [Intake Checklist](references/intake-checklist.md) - question bank with defaults
98- [Phase Deliverables](references/phase-deliverables.md) - interview phase to artifact and diagram map
99- [Interview Coaching](../system-design-interview-coaching/SKILL.md) - timed mock rounds, rubric, mistakes
100 

Discussion

Alternatives

Research methodology design for health literacy and medication adherence in aotearoa new zealandExplore the methodological design for researching health literacy and its impact on medication adherence among adults with chronic diseases in Aotearoa New Zealand.Business & ops · CC0-1.0Scientific critical thinkingEvaluate scientific claims and evidence quality. Use for assessing experimental design validity, identifying biases and confounders, applying evidence grading frameworks (GRADE, Cochrane Risk of Bias), or teaching critical analysis. Best for understanding evidence quality, identifying flaws. For formal peer review writing use peer-review.Science · MITAcademic research synthesizerAcademic research synthesis specialist. Use PROACTIVELY for comprehensive research on academic topics, literature reviews, technical investigations, and well-cited analysis combining multiple sources. <example>Context: A podcast episode needs a segment grounded in peer-reviewed evidence with formal citations. user: "Research the current state of transformer efficiency techniques for the episode, with proper academic citations." assistant: "I'll use the academic-research-synthesizer agent to search arXiv and Semantic Scholar, extract full-text findings via WebFetch, and produce a cited literature synthesis with confidence levels." <commentary>Use academic-research-synthesizer (not comprehensive-researcher) when the episode segment needs peer-reviewed sourcing, formal citation format, and explicit confidence tagging rather than general-purpose multi-source coverage.</commentary></example> <example>Context: The episode-orchestrator has routed a "literature review" request for a technical deep-dive segment. user: "Summarize the research landscape on federated learning privacy guarantees." assistant: "I'll invoke academic-research-synthesizer to systematically search academic sources, note peer-review status per source, and synthesize consensus vs. open debates."</example>Business & ops · MITAcademic researcherAcademic research specialist for scholarly sources, peer-reviewed papers, and academic literature. Use PROACTIVELY for research paper analysis, literature reviews, citation tracking, and academic methodology evaluation. <example>Context: The research-orchestrator has kicked off Phase 4 parallel research on 'efficacy of intermittent fasting' and needs peer-reviewed evidence. user: "Find the academic evidence on intermittent fasting outcomes." assistant: "I'll use the academic-researcher agent to search Semantic Scholar, PubMed, and OpenAlex for peer-reviewed studies and write structured findings to academic-research.md." <commentary>The request is specifically for scholarly/peer-reviewed evidence rather than general web coverage or code, so academic-researcher (not web-researcher or technical-researcher) is the right specialist.</commentary></example> <example>Context: The user wants a literature review comparing methodologies across studies on a topic. user: "Can you review the literature on transformer model interpretability and identify research gaps?" assistant: "Let me invoke the academic-researcher agent to pull foundational and recent papers, extract methodologies, and surface open research gaps." <commentary>Literature review, methodology extraction, and research-gap identification are core academic-researcher capabilities, distinct from technical-researcher's focus on code repositories and implementations.</commentary></example>Business & ops · MIT