Autoresearch Skill

Run Karpathy-style autoresearch optimization on any content.

by ericosiu·MIT license·★ 3,615 Stars on the repo·GitHub ↗

Use now

Files of Autoresearch Skill

ericosiu/main1 file shown
SKILL.md
Show the full text260 lines

Autoresearch Skill

Karpathy-style optimization loops for any conversion-focused content. No traffic needed. Simulated expert panel. Minutes, not weeks.

When to use this: Pre-launch content optimization. Generate 50+ variants, score with 5 simulated experts, evolve winners, output the best version + full experiment log.

When NOT to use this: Post-launch real-traffic A/B testing — that requires real analytics, not simulated scoring.

The sequence: Run autoresearch FIRST to hit 85+ simulated score. Then deploy. Then validate with real traffic.


What You'll Produce

Every run outputs 3 files:

File Purpose
{name}-optimized.{ext} The winning optimized content
data/{name}-experiments.json Full experiment log — all variants + all scores
data/{name}-optimization-report.md Human-readable summary with winner rationale

Expert Panel (5 Personas)

Score every variant against all 5. Batch all variants into a single API call per round.

# Persona Scoring Lens
1 CMO at a mid-market B2B company (50M+ revenue) "Would this make me stop and engage?"
2 Skeptical founder "Do I believe this? Would I trust this company?"
3 Conversion rate optimizer "Is this clear, specific, and action-driving?"
4 Senior copywriter "Is this compelling, differentiated, and well-crafted?"
5 Your CEO/founder "Direct, ROI-obsessed, no BS. Would I put this on my site?"

Customization: Replace persona #5 with your own CEO/founder voice. Define their priorities and communication style in a references/founder-voice.md file.

Each judge scores 0–100. Final score = average across all 5 judges.


Round Structure (Per Content Element)

Round 1:
  → Generate 10 variants of the element
  → Batch-score all 10 with the 5-expert panel (1 API call)
  → Rank by average score
  → Keep top 3

Round 2 (Evolution):
  → Analyze what the top 3 did right
  → Generate 10 new variants that push those winning patterns further
  → Batch-score all 10 (1 API call)
  → Keep top 3

Round 3 (If score < threshold):
  → Identify weakest scoring dimension
  → Generate 10 variants optimized for that dimension
  → Batch-score → keep top 1

Multi-element cross-breeding:
  → Take top 1 winner from each element
  → Generate 5 combinations that mix winning elements
  → Score holistically as complete units
  → Output the single best combination

Stop condition: Top variant hits minimum score threshold (default: 80) OR 3 rounds complete.


Content Types & Score Dimensions

Landing Pages

Elements to optimize: Hero headline, subheadline, CTA text, problem section, social proof

Score dimensions:

  • first_impression — Does it grab immediately?
  • clarity — Is the offer instantly understood?
  • trust — Does it feel credible?
  • urgency — Is there a reason to act now?
  • would_convert — Would the judge actually click?
Email Sequences

Elements to optimize: Subject line, opening line, body copy, CTA, PS line

Score dimensions:

  • would_open — Subject line pass rate
  • would_read — Does the opening hook?
  • would_click — Is the CTA compelling?
  • would_reply — Does it feel personal enough to respond to?
  • spam_risk — Does it feel spammy? (lower = better; invert for final score)
Ad Copy

Elements to optimize: Headline, description, CTA

Score dimensions:

  • scroll_stopping — Does it interrupt the scroll?
  • clarity — Is the value prop clear in 3 seconds?
  • click_worthiness — Does the judge want to click?
  • relevance — Does it match likely audience intent?
  • differentiation — Does it stand out from competitors?
Form Pages

Elements to optimize: Headline, subtext, value prop bullets, button text, field order, thank-you copy

Score dimensions:

  • first_impression — Does it feel worth filling out?
  • trust — Do they believe their info is safe and the offer is real?
  • completion_likelihood — Would the judge start filling it out?
  • lead_quality — Would this attract serious prospects (not tire-kickers)?
  • would_fill_out — Final gut check: would they submit?

Step-by-Step Execution Protocol

Step 1: Intake & Parse

Read the source content. Identify content type automatically or confirm with user:

  • HTML file → landing page or form page
  • Markdown / plain text → email or ad copy
  • If ambiguous, ask: "Is this a landing page, email sequence, ad copy, or form page?"

Extract all optimizable elements. List them back to user:

Found 5 elements to optimize:
1. Hero headline: "We help B2B companies grow"
2. Subheadline: "Full-service digital marketing..."
3. CTA: "Get Started"
4. Problem statement: [excerpt]
5. Social proof: [excerpt]

Optimizing: all | Variants per round: 10 | Min score: 80
Step 2: Get API Key

Check for Anthropic API key: $ANTHROPIC_API_KEY environment variable.

export ANTHROPIC_API_KEY="your-api-key-here"
Step 3: Run Optimization Rounds

For each element, run the round structure above.

Critical API efficiency rule: ALWAYS batch all variants into a single prompt. Never call the API once per variant. A round with 10 variants = 1 API call.

Model preference (in order):

  1. claude-sonnet-4-5 (preferred — fast + smart)
  2. claude-opus-4 (if highest quality needed)
  3. Any claude-3.5+ model if the above aren't available
Step 4: Cross-Breed (Multi-Element)

After all elements have winners:

  1. Assemble the top winner from each element into a complete unit
  2. Generate 5 holistic variants that naturally combine the winning elements
  3. Score the complete units (not just individual parts)
  4. Pick the winner with the highest holistic score
Step 5: Write Output Files
# Create output directory
mkdir -p data

# Write optimized content
# Write experiments JSON
# Write optimization report

Experiments JSON structure:

{
  "run_id": "autoresearch-{name}-{timestamp}",
  "content_type": "landing_page",
  "source_file": "path/to/original",
  "min_score_threshold": 80,
  "rounds": [
    {
      "round": 1,
      "element": "hero_headline",
      "variants": [
        {
          "id": 1,
          "text": "...",
          "scores": {
            "cmo": 72,
            "skeptical_founder": 68,
            "cro": 75,
            "copywriter": 70,
            "founder": 65
          },
          "avg_score": 70
        }
      ],
      "top_3": [1, 4, 7],
      "winner_score": 82
    }
  ],
  "final_winner": {
    "hero_headline": "...",
    "subheadline": "...",
    "cta": "...",
    "holistic_score": 87
  }
}
Step 6: Report Back

Summarize results to user:

  • Final winning score
  • Biggest score jump (which element improved most)
  • Top 2 runner-up alternatives (in case winner doesn't feel right)
  • Path to all 3 output files
  • Clear next step

User Options

Option Default Description
elements all Which elements to optimize
variants_per_round 10 How many variants to generate per round
min_score 80 Stop when this score is hit
rounds 3 Max rounds before stopping
auto_apply false Whether to overwrite the source file with winners
content_type auto-detect Force a content type if auto-detect is wrong

Quality Gates

  • < 70: Don't ship. Something fundamental is broken.
  • 70-79: Marginal. One more round targeting the lowest-scoring dimension.
  • 80-84: Good. Shippable. Validate with real traffic.
  • 85-89: Strong. Ship with confidence.
  • 90+: Rare. Ship immediately.

Anti-Patterns to Avoid

  • Never call the API once per variant. Always batch. A 10-variant round = 1 call.
  • Don't over-optimize for one dimension. If you're hitting 95 on clarity but 45 on trust, the overall score is misleading.
  • Don't run more than 5 rounds. If you're not hitting 80 after 3 rounds, the problem is strategic (wrong positioning), not tactical (wrong words).
  • Don't cross-breed until each element has its own winner. Premature cross-breeding creates incoherent combinations.
1---
2name: autoresearch
3description: Run Karpathy-style autoresearch optimization on any content. Generates 50+ variants, scores with a 5-expert simulated panel, evolves winners through multiple rounds, outputs optimized version + full experiment log. Use when optimizing landing pages, email sequences, ad copy, headlines, form pages, CTA text, or any conversion-focused content. Triggers on "optimize this page", "run autoresearch", "score these variants", "A/B test this copy".
4---
5 
6# Autoresearch Skill
7 
8Karpathy-style optimization loops for any conversion-focused content. No traffic needed. Simulated expert panel. Minutes, not weeks.
9 
10**When to use this:** Pre-launch content optimization. Generate 50+ variants, score with 5 simulated experts, evolve winners, output the best version + full experiment log.
11 
12**When NOT to use this:** Post-launch real-traffic A/B testing — that requires real analytics, not simulated scoring.
13 
14> **The sequence:** Run autoresearch FIRST to hit 85+ simulated score. Then deploy. Then validate with real traffic.
15 
16---
17 
18## What You'll Produce
19 
20Every run outputs 3 files:
21 
22| File | Purpose |
23|------|---------|
24| `{name}-optimized.{ext}` | The winning optimized content |
25| `data/{name}-experiments.json` | Full experiment log — all variants + all scores |
26| `data/{name}-optimization-report.md` | Human-readable summary with winner rationale |
27 
28---
29 
30## Expert Panel (5 Personas)
31 
32Score every variant against all 5. Batch all variants into a **single API call** per round.
33 
34| # | Persona | Scoring Lens |
35|---|---------|-------------|
36| 1 | **CMO at a mid-market B2B company (50M+ revenue)** | "Would this make me stop and engage?" |
37| 2 | **Skeptical founder** | "Do I believe this? Would I trust this company?" |
38| 3 | **Conversion rate optimizer** | "Is this clear, specific, and action-driving?" |
39| 4 | **Senior copywriter** | "Is this compelling, differentiated, and well-crafted?" |
40| 5 | **Your CEO/founder** | "Direct, ROI-obsessed, no BS. Would I put this on my site?" |
41 
42> **Customization:** Replace persona #5 with your own CEO/founder voice. Define their priorities and communication style in a `references/founder-voice.md` file.
43 
44Each judge scores 0–100. **Final score = average across all 5 judges.**
45 
46---
47 
48## Round Structure (Per Content Element)
49 
50```
51Round 1:
52 → Generate 10 variants of the element
53 → Batch-score all 10 with the 5-expert panel (1 API call)
54 → Rank by average score
55 → Keep top 3
56 
57Round 2 (Evolution):
58 → Analyze what the top 3 did right
59 → Generate 10 new variants that push those winning patterns further
60 → Batch-score all 10 (1 API call)
61 → Keep top 3
62 
63Round 3 (If score < threshold):
64 → Identify weakest scoring dimension
65 → Generate 10 variants optimized for that dimension
66 → Batch-score → keep top 1
67 
68Multi-element cross-breeding:
69 → Take top 1 winner from each element
70 → Generate 5 combinations that mix winning elements
71 → Score holistically as complete units
72 → Output the single best combination
73```
74 
75**Stop condition:** Top variant hits minimum score threshold (default: 80) OR 3 rounds complete.
76 
77---
78 
79## Content Types & Score Dimensions
80 
81### Landing Pages
82**Elements to optimize:** Hero headline, subheadline, CTA text, problem section, social proof
83 
84**Score dimensions:**
85- `first_impression` — Does it grab immediately?
86- `clarity` — Is the offer instantly understood?
87- `trust` — Does it feel credible?
88- `urgency` — Is there a reason to act now?
89- `would_convert` — Would the judge actually click?
90 
91### Email Sequences
92**Elements to optimize:** Subject line, opening line, body copy, CTA, PS line
93 
94**Score dimensions:**
95- `would_open` — Subject line pass rate
96- `would_read` — Does the opening hook?
97- `would_click` — Is the CTA compelling?
98- `would_reply` — Does it feel personal enough to respond to?
99- `spam_risk` — Does it feel spammy? (lower = better; invert for final score)
100 
101### Ad Copy
102**Elements to optimize:** Headline, description, CTA
103 
104**Score dimensions:**
105- `scroll_stopping` — Does it interrupt the scroll?
106- `clarity` — Is the value prop clear in 3 seconds?
107- `click_worthiness` — Does the judge want to click?
108- `relevance` — Does it match likely audience intent?
109- `differentiation` — Does it stand out from competitors?
110 
111### Form Pages
112**Elements to optimize:** Headline, subtext, value prop bullets, button text, field order, thank-you copy
113 
114**Score dimensions:**
115- `first_impression` — Does it feel worth filling out?
116- `trust` — Do they believe their info is safe and the offer is real?
117- `completion_likelihood` — Would the judge start filling it out?
118- `lead_quality` — Would this attract serious prospects (not tire-kickers)?
119- `would_fill_out` — Final gut check: would they submit?
120 
121---
122 
123## Step-by-Step Execution Protocol
124 
125### Step 1: Intake & Parse
126 
127Read the source content. Identify content type automatically or confirm with user:
128- HTML file → landing page or form page
129- Markdown / plain text → email or ad copy
130- If ambiguous, ask: "Is this a landing page, email sequence, ad copy, or form page?"
131 
132Extract all optimizable elements. List them back to user:
133```
134Found 5 elements to optimize:
1351. Hero headline: "We help B2B companies grow"
1362. Subheadline: "Full-service digital marketing..."
1373. CTA: "Get Started"
1384. Problem statement: [excerpt]
1395. Social proof: [excerpt]
140 
141Optimizing: all | Variants per round: 10 | Min score: 80
142```
143 
144### Step 2: Get API Key
145 
146Check for Anthropic API key: `$ANTHROPIC_API_KEY` environment variable.
147 
148```bash
149export ANTHROPIC_API_KEY="your-api-key-here"
150```
151 
152### Step 3: Run Optimization Rounds
153 
154For each element, run the round structure above.
155 
156**Critical API efficiency rule:** ALWAYS batch all variants into a single prompt. Never call the API once per variant. A round with 10 variants = 1 API call.
157 
158Model preference (in order):
1591. `claude-sonnet-4-5` (preferred — fast + smart)
1602. `claude-opus-4` (if highest quality needed)
1613. Any claude-3.5+ model if the above aren't available
162 
163### Step 4: Cross-Breed (Multi-Element)
164 
165After all elements have winners:
1661. Assemble the top winner from each element into a complete unit
1672. Generate 5 holistic variants that naturally combine the winning elements
1683. Score the complete units (not just individual parts)
1694. Pick the winner with the highest holistic score
170 
171### Step 5: Write Output Files
172 
173```bash
174# Create output directory
175mkdir -p data
176 
177# Write optimized content
178# Write experiments JSON
179# Write optimization report
180```
181 
182**Experiments JSON structure:**
183```json
184{
185 "run_id": "autoresearch-{name}-{timestamp}",
186 "content_type": "landing_page",
187 "source_file": "path/to/original",
188 "min_score_threshold": 80,
189 "rounds": [
190 {
191 "round": 1,
192 "element": "hero_headline",
193 "variants": [
194 {
195 "id": 1,
196 "text": "...",
197 "scores": {
198 "cmo": 72,
199 "skeptical_founder": 68,
200 "cro": 75,
201 "copywriter": 70,
202 "founder": 65
203 },
204 "avg_score": 70
205 }
206 ],
207 "top_3": [1, 4, 7],
208 "winner_score": 82
209 }
210 ],
211 "final_winner": {
212 "hero_headline": "...",
213 "subheadline": "...",
214 "cta": "...",
215 "holistic_score": 87
216 }
217}
218```
219 
220### Step 6: Report Back
221 
222Summarize results to user:
223- Final winning score
224- Biggest score jump (which element improved most)
225- Top 2 runner-up alternatives (in case winner doesn't feel right)
226- Path to all 3 output files
227- Clear next step
228 
229---
230 
231## User Options
232 
233| Option | Default | Description |
234|--------|---------|-------------|
235| `elements` | all | Which elements to optimize |
236| `variants_per_round` | 10 | How many variants to generate per round |
237| `min_score` | 80 | Stop when this score is hit |
238| `rounds` | 3 | Max rounds before stopping |
239| `auto_apply` | false | Whether to overwrite the source file with winners |
240| `content_type` | auto-detect | Force a content type if auto-detect is wrong |
241 
242---
243 
244## Quality Gates
245 
246- **< 70:** Don't ship. Something fundamental is broken.
247- **70-79:** Marginal. One more round targeting the lowest-scoring dimension.
248- **80-84:** Good. Shippable. Validate with real traffic.
249- **85-89:** Strong. Ship with confidence.
250- **90+:** Rare. Ship immediately.
251 
252---
253 
254## Anti-Patterns to Avoid
255 
256- **Never call the API once per variant.** Always batch. A 10-variant round = 1 call.
257- **Don't over-optimize for one dimension.** If you're hitting 95 on clarity but 45 on trust, the overall score is misleading.
258- **Don't run more than 5 rounds.** If you're not hitting 80 after 3 rounds, the problem is strategic (wrong positioning), not tactical (wrong words).
259- **Don't cross-breed until each element has its own winner.** Premature cross-breeding creates incoherent combinations.
260 

Discussion