Science skill

The scientific method as a universal problem-solving algorithm — goal-first, plural falsifiable hypotheses, designed experiments, and honest measurement, scaling from TDD to feature validation to MVP launch.

by danielmiessler·MIT license·★ 19,269 Stars on the repo·GitHub ↗

Use now

Files of Science

danielmiessler/main1 file shown
SKILL.md
Show the full text191 lines

Customization

Before executing, check for user customizations at: ~/.claude/LIFEOS/USER/CUSTOMIZATIONS/SKILLS/Science/

If this directory exists, load and apply any PREFERENCES.md, configurations, or resources found there. These override default behavior. If the directory does not exist, proceed with skill defaults.

🚨 MANDATORY: Voice Notification (REQUIRED BEFORE ANY ACTION)

You MUST send this notification BEFORE doing anything else when this skill is invoked.

  1. Send voice notification:

    curl -s -X POST http://localhost:31337/notify \
      -H "Content-Type: application/json" \
      -d '{"message": "Running the WORKFLOWNAME workflow in the Science skill to ACTION"}' \
      > /dev/null 2>&1 &
    
  2. Output text notification:

    Running the **WorkflowName** workflow in the **Science** skill to ACTION...
    

This is not optional. Execute this curl command immediately upon skill invocation.

Science - The Universal Algorithm

What It Does

Applies the scientific method as a general problem-solving algorithm: define the goal first, generate multiple hypotheses, design experiments that can fail, measure honestly, analyze against the goal, iterate. Seven core workflows plus two diagnostic shortcuts (quick 15-minute debugging and structured multi-factor investigation). It scales from micro (TDD) to meso (feature validation) to macro (MVP launch).

The Problem

Most problem-solving is guessing dressed up as work. You pick the first idea that comes to mind, change something, and call it done when it "seems better" — which is confirmation bias, not progress. Without a clear definition of success you can't tell whether a change helped, so you keep tweaking forever or stop too early. Single-hypothesis thinking means you only ever test the idea you already believed. This skill forces the discipline that fixes all of that: a stated goal, at least three competing hypotheses, falsifiable tests, and measurement that compares to the goal rather than to your hopes.

How It Works

The whole thing is one repeating cycle, and the goal anchors it — without clear success criteria you cannot judge results:

GOAL -----> What does success look like?
   |
OBSERVE --> What is the current state?
   |
HYPOTHESIZE -> What might work? (Generate MULTIPLE)
   |
EXPERIMENT -> Design and run the test
   |
MEASURE --> What happened? (Data collection)
   |
ANALYZE --> How does it compare to the goal?
   |
ITERATE --> Adjust hypothesis and repeat
   |
   +------> Back to HYPOTHESIZE

The answer emerges from the cycle, not from guessing.


Workflow Routing

Output when executing: Running the **WorkflowName** workflow in the **Science** skill to ACTION...

Core Workflows
Workflow Trigger File
DefineGoal "define the goal", "what are we trying to achieve" Workflows/DefineGoal.md
GenerateHypotheses "what might work", "ideas", "hypotheses" Workflows/GenerateHypotheses.md
DesignExperiment "how do we test", "experiment design" Workflows/DesignExperiment.md
MeasureResults "what happened", "measure", "results" Workflows/MeasureResults.md
AnalyzeResults "analyze", "compare to goal" Workflows/AnalyzeResults.md
Iterate "iterate", "try again", "next cycle" Workflows/Iterate.md
FullCycle Full structured cycle Workflows/FullCycle.md
Diagnostic Workflows
Workflow Trigger File
QuickDiagnosis Quick debugging (15-min rule) Workflows/QuickDiagnosis.md
StructuredInvestigation Complex investigation Workflows/StructuredInvestigation.md

Resource Index

Resource Description
Methodology.md Deep dive into each phase
Protocol.md How skills implement Science
Templates.md Goal, Hypothesis, Experiment, Results templates
Examples.md Worked examples across scales

Domain Applications

Domain Manifestation Related Skill
Coding TDD (Red-Green-Refactor) Development
Products MVP -> Measure -> Iterate Development
Research Question -> Study -> Analyze Research
Prompts Prompt -> Eval -> Iterate Evals
Decisions Options -> Council -> Choose Council

Scale of Application

Level Cycle Time Example
Micro Minutes TDD: test, code, refactor
Meso Hours-Days Feature: spec, implement, validate
Macro Weeks-Months Product: MVP, launch, measure PMF

Integration Points

Phase Skills to Invoke
Goal Council for validation
Observe Research for context
Hypothesize Council for ideas, RedTeam for stress-test
Experiment Development (Worktrees) for parallel tests
Measure Evals for structured measurement
Analyze Council for multi-perspective analysis

Anti-Patterns

Bad Good
"Make it better" "Reduce load time from 3s to 1s"
"I think X will work" "Here are 3 approaches: X, Y, Z"
"Prove I'm right" "Design test that could disprove"
"Pretend failure didn't happen" "What did we learn?"
"Keep experimenting forever" "Ship and learn from production"

Gotchas

  • Minimum 3 hypotheses before testing. Single-hypothesis testing is confirmation bias — going straight to a single test is trial-and-error, not science.
  • Measurements must be specific and reproducible. "It seems better" is not a measurement.
  • Full cycle is for systematic investigation. For quick debugging, use quick diagnosis mode.

Examples

Example 1: Quick diagnosis

User: "figure out why Surface time filters show stale items"
→ Quick diagnosis mode
→ Hypothesis: timestamp format mismatch in D1
→ Test: query D1 for actual stored format
→ Analyze: compare stored vs expected format
→ Result: ISO string vs Unix timestamp mismatch

Example 2: Full systematic investigation

User: "experiment with different prompt structures for better output"
→ Full cycle mode
→ 3+ hypotheses generated
→ Controlled experiments with measurements
→ Analysis identifies winning approach
→ Iterates until convergence

Execution Log

After completing any workflow, append a single JSONL entry:

echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"Science","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl

Replace WORKFLOW_USED with the workflow executed, 8_WORD_SUMMARY with a brief input description, and SECONDS with approximate wall-clock time. Log status: "error" if the workflow failed.

1---
2name: Science
3version: 1.1.18
4description: "The scientific method as a universal problem-solving algorithm — goal-first, plural falsifiable hypotheses, designed experiments, and honest measurement, scaling from TDD to feature validation to MVP launch. USE WHEN think about, figure out, experiment, iterate, optimize, hypothesis, science, full cycle, quick diagnosis, structured investigation, how do we test, analyze results. NOT FOR multi-angle lens passes (use IterativeDepth)."
5---
6 
7## Customization
8 
9**Before executing, check for user customizations at:**
10`~/.claude/LIFEOS/USER/CUSTOMIZATIONS/SKILLS/Science/`
11 
12If this directory exists, load and apply any PREFERENCES.md, configurations, or resources found there. These override default behavior. If the directory does not exist, proceed with skill defaults.
13 
14 
15## 🚨 MANDATORY: Voice Notification (REQUIRED BEFORE ANY ACTION)
16 
17**You MUST send this notification BEFORE doing anything else when this skill is invoked.**
18 
191. **Send voice notification**:
20 ```bash
21 curl -s -X POST http://localhost:31337/notify \
22 -H "Content-Type: application/json" \
23 -d '{"message": "Running the WORKFLOWNAME workflow in the Science skill to ACTION"}' \
24 > /dev/null 2>&1 &
25 ```
26 
272. **Output text notification**:
28 ```
29 Running the **WorkflowName** workflow in the **Science** skill to ACTION...
30 ```
31 
32**This is not optional. Execute this curl command immediately upon skill invocation.**
33 
34# Science - The Universal Algorithm
35 
36## What It Does
37 
38Applies the scientific method as a general problem-solving algorithm: define the goal first, generate multiple hypotheses, design experiments that can fail, measure honestly, analyze against the goal, iterate. Seven core workflows plus two diagnostic shortcuts (quick 15-minute debugging and structured multi-factor investigation). It scales from micro (TDD) to meso (feature validation) to macro (MVP launch).
39 
40## The Problem
41 
42Most problem-solving is guessing dressed up as work. You pick the first idea that comes to mind, change something, and call it done when it "seems better" — which is confirmation bias, not progress. Without a clear definition of success you can't tell whether a change helped, so you keep tweaking forever or stop too early. Single-hypothesis thinking means you only ever test the idea you already believed. This skill forces the discipline that fixes all of that: a stated goal, at least three competing hypotheses, falsifiable tests, and measurement that compares to the goal rather than to your hopes.
43 
44## How It Works
45 
46The whole thing is one repeating cycle, and the goal anchors it — without clear success criteria you cannot judge results:
47 
48```
49GOAL -----> What does success look like?
50 |
51OBSERVE --> What is the current state?
52 |
53HYPOTHESIZE -> What might work? (Generate MULTIPLE)
54 |
55EXPERIMENT -> Design and run the test
56 |
57MEASURE --> What happened? (Data collection)
58 |
59ANALYZE --> How does it compare to the goal?
60 |
61ITERATE --> Adjust hypothesis and repeat
62 |
63 +------> Back to HYPOTHESIZE
64```
65 
66The answer emerges from the cycle, not from guessing.
67 
68---
69 
70 
71## Workflow Routing
72 
73**Output when executing:** `Running the **WorkflowName** workflow in the **Science** skill to ACTION...`
74 
75### Core Workflows
76 
77| Workflow | Trigger | File |
78|----------|---------|------|
79| **DefineGoal** | "define the goal", "what are we trying to achieve" | `Workflows/DefineGoal.md` |
80| **GenerateHypotheses** | "what might work", "ideas", "hypotheses" | `Workflows/GenerateHypotheses.md` |
81| **DesignExperiment** | "how do we test", "experiment design" | `Workflows/DesignExperiment.md` |
82| **MeasureResults** | "what happened", "measure", "results" | `Workflows/MeasureResults.md` |
83| **AnalyzeResults** | "analyze", "compare to goal" | `Workflows/AnalyzeResults.md` |
84| **Iterate** | "iterate", "try again", "next cycle" | `Workflows/Iterate.md` |
85| **FullCycle** | Full structured cycle | `Workflows/FullCycle.md` |
86 
87### Diagnostic Workflows
88 
89| Workflow | Trigger | File |
90|----------|---------|------|
91| **QuickDiagnosis** | Quick debugging (15-min rule) | `Workflows/QuickDiagnosis.md` |
92| **StructuredInvestigation** | Complex investigation | `Workflows/StructuredInvestigation.md` |
93 
94---
95 
96## Resource Index
97 
98| Resource | Description |
99|----------|-------------|
100| `Methodology.md` | Deep dive into each phase |
101| `Protocol.md` | How skills implement Science |
102| `Templates.md` | Goal, Hypothesis, Experiment, Results templates |
103| `Examples.md` | Worked examples across scales |
104 
105---
106 
107## Domain Applications
108 
109| Domain | Manifestation | Related Skill |
110|--------|---------------|---------------|
111| **Coding** | TDD (Red-Green-Refactor) | Development |
112| **Products** | MVP -> Measure -> Iterate | Development |
113| **Research** | Question -> Study -> Analyze | Research |
114| **Prompts** | Prompt -> Eval -> Iterate | Evals |
115| **Decisions** | Options -> Council -> Choose | Council |
116 
117---
118 
119## Scale of Application
120 
121| Level | Cycle Time | Example |
122|-------|-----------|---------|
123| **Micro** | Minutes | TDD: test, code, refactor |
124| **Meso** | Hours-Days | Feature: spec, implement, validate |
125| **Macro** | Weeks-Months | Product: MVP, launch, measure PMF |
126 
127---
128 
129## Integration Points
130 
131| Phase | Skills to Invoke |
132|-------|-----------------|
133| **Goal** | Council for validation |
134| **Observe** | Research for context |
135| **Hypothesize** | Council for ideas, RedTeam for stress-test |
136| **Experiment** | Development (Worktrees) for parallel tests |
137| **Measure** | Evals for structured measurement |
138| **Analyze** | Council for multi-perspective analysis |
139 
140---
141 
142## Anti-Patterns
143 
144| Bad | Good |
145|-----|------|
146| "Make it better" | "Reduce load time from 3s to 1s" |
147| "I think X will work" | "Here are 3 approaches: X, Y, Z" |
148| "Prove I'm right" | "Design test that could disprove" |
149| "Pretend failure didn't happen" | "What did we learn?" |
150| "Keep experimenting forever" | "Ship and learn from production" |
151 
152---
153 
154## Gotchas
155 
156- **Minimum 3 hypotheses before testing.** Single-hypothesis testing is confirmation bias — going straight to a single test is trial-and-error, not science.
157- **Measurements must be specific and reproducible.** "It seems better" is not a measurement.
158- **Full cycle is for systematic investigation.** For quick debugging, use quick diagnosis mode.
159 
160## Examples
161 
162**Example 1: Quick diagnosis**
163```
164User: "figure out why Surface time filters show stale items"
165→ Quick diagnosis mode
166→ Hypothesis: timestamp format mismatch in D1
167→ Test: query D1 for actual stored format
168→ Analyze: compare stored vs expected format
169→ Result: ISO string vs Unix timestamp mismatch
170```
171 
172**Example 2: Full systematic investigation**
173```
174User: "experiment with different prompt structures for better output"
175→ Full cycle mode
176→ 3+ hypotheses generated
177→ Controlled experiments with measurements
178→ Analysis identifies winning approach
179→ Iterates until convergence
180```
181 
182## Execution Log
183 
184After completing any workflow, append a single JSONL entry:
185 
186```bash
187echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"Science","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl
188```
189 
190Replace `WORKFLOW_USED` with the workflow executed, `8_WORD_SUMMARY` with a brief input description, and `SECONDS` with approximate wall-clock time. Log `status: "error"` if the workflow failed.
191 

Discussion