Deep Research — Disciplined Meta-Research skill

Run a disciplined, multi-source research investigation for a high-stakes question or decision — fan-out web search across many channels, parallel sub-agents, source triangulation (each claim backed by ≥3 independent sources), an adversarial review pass, and every source saved to its own file with verbatim quotes for reuse.

by alirezarezvani·MIT license·★ 26,349 Stars on the repo·GitHub ↗

Use now

Files of Deep Research — Disciplined Meta-Research

alirezarezvani/main1 file shown
SKILL.md
Show the full text87 lines

Deep Research — Disciplined Meta-Research

Turn "research this topic" into an auditable, reusable investigation instead of a one-shot wall of text. The output is a folder you can return to in a month: every claim traces to a specific source file, the plan documents why each choice was made, and a refresh protocol lets you update it later without re-running everything.

This is the heavy, methodical end of research. It is not a fast overview — it is the workflow you reach for when getting the answer wrong costs more than the tokens spent getting it right.

How it differs from a quick research router

A router-style research skill (keyword-classify → delegate → short sequential search → markdown brief) is optimal when you need an answer fast and the decision risk is low. deep-research is the opposite trade: it pays for rigor. Use it when the answer feeds a strategy, an irreversible decision, a published artifact, or a hypothesis you need to actually test — situations where a shallow fallback would be a liability.

Concretely, deep-research adds what a fast overview does not: falsifiable hypotheses up front, parallel sub-agent fan-out across many channels, triangulation with explicit source-type diversity, a mandatory adversarial pass, per-source files with verbatim quotes, and a refresh_targets.md for delta-updates later.

The pipeline (9 phases)

Depth scales with the task — shallow runs the core phases inline; medium/deep add capability discovery, verification, and refresh targets.

# Phase What it does
1 Reframe Rewrite the question, fix the underlying decision, state 2–4 falsifiable hypotheses
2 Genre & blocks Pick the report genre (qa / explainer / decision / landscape / validation / custom) and its building blocks
3 Plan Write plan.md: scope, structure, sourcing strategy, opposition queries, risk register, stop-criteria
3.5 Capability discovery Audit available API keys/channels in the environment; map subtopics to sources; fall back to HTML where needed
4 Search (loop) Dispatch sources → launch sub-agents in parallel → fetch & dedup → save each to sources/NN.md; re-evaluate between rounds
5 Score & triangulate Rate every source on Credibility / Recency / Bias; require ≥3 independent, differently-typed sources per thesis
6 Synthesize + adversarial Assemble the report from blocks, run 4 self-critique questions, add steel-manned counter-arguments
6.5 Verify Lightweight citation check before closing
7 Refresh targets Extract entities / numbers / hypotheses into refresh_targets.md — the entry point for future updates

Core mechanisms

These are what separate a documented investigation from a confident guess:

  • Triangulation. Every thesis must be backed by ≥3 independent sources of different types (primary / academic / industry / discussion). A claim with fewer is flagged "insufficient evidence," not stated as fact.
  • Source-grounding. Each source becomes its own sources/NN_slug.md with metadata, verbatim quotes, and scores. No dangling claim — every assertion links back to a specific file. An empty fetch produces an empty claim, never a fabricated citation.
  • Adversarial pass. Phase 6 always runs the strongest available reasoning: 4 self-critique questions plus an active search for counter-arguments and disconfirming evidence.
  • Falsifiable hypotheses. Phase 1 commits to 2–4 hypotheses; Phases 5–6 explicitly confirm or refute each against the evidence, or mark it under-determined.
  • Parallel sub-agents. Phase 4 launches search sub-agents concurrently (cheap models for broad web sweeps, stronger ones for reasoning-heavy subtopics) — never one-at-a-time.
  • Refresh protocol. Phase 7 emits refresh_targets.md; an update <slug> run produces a delta (new entrants, entity changes, refreshed numbers, adversarial triggers) instead of replaying the whole investigation.
  • Atomic findings. Reusable theses in findings/FN.md plus a sources.csv index — research compounds across questions instead of starting from zero each time.

Output structure

<root>/<slug>/
├── plan.md                  # scope, sourcing strategy, risk register, changelog
├── sources.csv              # index of every source with scores
├── sources/
│   ├── 01_<slug>.md         # one file = one source (metadata + verbatim quotes)
│   └── ...
├── findings/                # atomic, reusable theses (larger investigations)
│   └── F1_<short>.md
├── refresh_targets.md       # what to watch on update (medium/deep)
├── diffs/
│   └── YYYY-MM-DD_delta.md   # delta from an `update <slug>` run
└── YYYY-MM-DD_<genre>.md     # final report

When to use

  • A low-quality answer is expensive: strategy, business plan, report, or article groundwork.
  • Comparing N institutions, products, methodologies, or markets and you need defensible reasoning.
  • Validating a hypothesis or a decision against external data.
  • Meta-research: "understand how X works," "map the landscape of Y," answering a connected series of questions.

Anti-Patterns

  • Don't skip the existing-work check. Before searching, see whether the answer is already in the project or in a prior research folder — you risk re-researching something you already have.
  • Don't skip reframing, even when the request "seems clear." The decision behind the question usually changes the search.
  • Don't output to chat only. Always persist sources and the report to files — the reuse value is in the folder, not the transcript.
  • Don't fabricate citations. If a fetch returns nothing, the claim is empty — never invent a plausible URL. Bind every claim to a saved verbatim quote.
  • Don't build conclusions on a thin corpus. Too few sources, or sources that all share one type, means triangulation hasn't happened — say so rather than overstating confidence.
  • Don't skip the adversarial pass on medium/deep investigations. Confirmation-only research is the failure mode this skill exists to prevent.
  • Don't run sub-agents sequentially. Fan-out in parallel; serial search wastes the wall-clock advantage.
  • Don't collapse sources/ into one file. Per-source files are what make findings searchable and reusable across investigations.
  • Don't pick the heaviest model for everything. Match model to subtask — cheap for broad sweeps, strong for synthesis and the adversarial pass.

Cross-References

  • research router — for fast topic overviews where decision risk is low; deep-research is the heavyweight alternative when rigor matters more than speed.
  • competitive-teardown — for comparing N competitors on a structured 12-dimension matrix.
  • litreview / dossier / patent — domain specialists when the investigation is narrowly academic, person/company-focused, or patent-focused.
1---
2name: "deep-research"
3description: "Run a disciplined, multi-source research investigation for a high-stakes question or decision — fan-out web search across many channels, parallel sub-agents, source triangulation (each claim backed by ≥3 independent sources), an adversarial review pass, and every source saved to its own file with verbatim quotes for reuse. Use when a low-quality answer is expensive: strategy work, comparing N products/methods/markets, validating a hypothesis with external data, or mapping how a field works. NOT for quick fact-checks (answer directly), structured 12-dimension competitor scoring (use competitive-teardown), or fast topic overviews where the decision risk is low (use the research router instead)."
4---
5 
6# Deep Research — Disciplined Meta-Research
7 
8Turn "research this topic" into an auditable, reusable investigation instead of a one-shot wall of text. The output is a folder you can return to in a month: every claim traces to a specific source file, the plan documents *why* each choice was made, and a refresh protocol lets you update it later without re-running everything.
9 
10**This is the heavy, methodical end of research.** It is not a fast overview — it is the workflow you reach for when getting the answer *wrong* costs more than the tokens spent getting it right.
11 
12## How it differs from a quick research router
13 
14A router-style research skill (keyword-classify → delegate → short sequential search → markdown brief) is optimal when you need an answer fast and the decision risk is low. `deep-research` is the opposite trade: it pays for rigor. Use it when the answer feeds a strategy, an irreversible decision, a published artifact, or a hypothesis you need to actually test — situations where a shallow fallback would be a liability.
15 
16Concretely, `deep-research` adds what a fast overview does not: falsifiable hypotheses up front, parallel sub-agent fan-out across many channels, triangulation with explicit source-type diversity, a mandatory adversarial pass, per-source files with verbatim quotes, and a `refresh_targets.md` for delta-updates later.
17 
18## The pipeline (9 phases)
19 
20Depth scales with the task — `shallow` runs the core phases inline; `medium`/`deep` add capability discovery, verification, and refresh targets.
21 
22| # | Phase | What it does |
23|---|-------|--------------|
24| 1 | **Reframe** | Rewrite the question, fix the underlying decision, state 2–4 *falsifiable* hypotheses |
25| 2 | **Genre & blocks** | Pick the report genre (qa / explainer / decision / landscape / validation / custom) and its building blocks |
26| 3 | **Plan** | Write `plan.md`: scope, structure, sourcing strategy, opposition queries, risk register, stop-criteria |
27| 3.5 | **Capability discovery** | Audit available API keys/channels in the environment; map subtopics to sources; fall back to HTML where needed |
28| 4 | **Search** (loop) | Dispatch sources → launch sub-agents in parallel → fetch & dedup → save each to `sources/NN.md`; re-evaluate between rounds |
29| 5 | **Score & triangulate** | Rate every source on Credibility / Recency / Bias; require ≥3 independent, differently-typed sources per thesis |
30| 6 | **Synthesize + adversarial** | Assemble the report from blocks, run 4 self-critique questions, add steel-manned counter-arguments |
31| 6.5 | **Verify** | Lightweight citation check before closing |
32| 7 | **Refresh targets** | Extract entities / numbers / hypotheses into `refresh_targets.md` — the entry point for future updates |
33 
34## Core mechanisms
35 
36These are what separate a documented investigation from a confident guess:
37 
38- **Triangulation.** Every thesis must be backed by ≥3 independent sources of *different types* (primary / academic / industry / discussion). A claim with fewer is flagged "insufficient evidence," not stated as fact.
39- **Source-grounding.** Each source becomes its own `sources/NN_slug.md` with metadata, verbatim quotes, and scores. No dangling claim — every assertion links back to a specific file. An empty fetch produces an empty claim, never a fabricated citation.
40- **Adversarial pass.** Phase 6 always runs the strongest available reasoning: 4 self-critique questions plus an active search for counter-arguments and disconfirming evidence.
41- **Falsifiable hypotheses.** Phase 1 commits to 2–4 hypotheses; Phases 5–6 explicitly confirm or refute each against the evidence, or mark it under-determined.
42- **Parallel sub-agents.** Phase 4 launches search sub-agents concurrently (cheap models for broad web sweeps, stronger ones for reasoning-heavy subtopics) — never one-at-a-time.
43- **Refresh protocol.** Phase 7 emits `refresh_targets.md`; an `update <slug>` run produces a delta (new entrants, entity changes, refreshed numbers, adversarial triggers) instead of replaying the whole investigation.
44- **Atomic findings.** Reusable theses in `findings/FN.md` plus a `sources.csv` index — research compounds across questions instead of starting from zero each time.
45 
46## Output structure
47 
48```
49<root>/<slug>/
50├── plan.md # scope, sourcing strategy, risk register, changelog
51├── sources.csv # index of every source with scores
52├── sources/
53│ ├── 01_<slug>.md # one file = one source (metadata + verbatim quotes)
54│ └── ...
55├── findings/ # atomic, reusable theses (larger investigations)
56│ └── F1_<short>.md
57├── refresh_targets.md # what to watch on update (medium/deep)
58├── diffs/
59│ └── YYYY-MM-DD_delta.md # delta from an `update <slug>` run
60└── YYYY-MM-DD_<genre>.md # final report
61```
62 
63## When to use
64 
65- A low-quality answer is expensive: strategy, business plan, report, or article groundwork.
66- Comparing N institutions, products, methodologies, or markets and you need defensible reasoning.
67- Validating a hypothesis or a decision against external data.
68- Meta-research: "understand how X works," "map the landscape of Y," answering a connected series of questions.
69 
70## Anti-Patterns
71 
72- **Don't skip the existing-work check.** Before searching, see whether the answer is already in the project or in a prior research folder — you risk re-researching something you already have.
73- **Don't skip reframing**, even when the request "seems clear." The decision behind the question usually changes the search.
74- **Don't output to chat only.** Always persist sources and the report to files — the reuse value is in the folder, not the transcript.
75- **Don't fabricate citations.** If a fetch returns nothing, the claim is empty — never invent a plausible URL. Bind every claim to a saved verbatim quote.
76- **Don't build conclusions on a thin corpus.** Too few sources, or sources that all share one type, means triangulation hasn't happened — say so rather than overstating confidence.
77- **Don't skip the adversarial pass** on medium/deep investigations. Confirmation-only research is the failure mode this skill exists to prevent.
78- **Don't run sub-agents sequentially.** Fan-out in parallel; serial search wastes the wall-clock advantage.
79- **Don't collapse `sources/` into one file.** Per-source files are what make findings searchable and reusable across investigations.
80- **Don't pick the heaviest model for everything.** Match model to subtask — cheap for broad sweeps, strong for synthesis and the adversarial pass.
81 
82## Cross-References
83 
84- **research router** — for fast topic overviews where decision risk is low; `deep-research` is the heavyweight alternative when rigor matters more than speed.
85- **competitive-teardown** — for comparing N competitors on a structured 12-dimension matrix.
86- **litreview / dossier / patent** — domain specialists when the investigation is narrowly academic, person/company-focused, or patent-focused.
87 

Discussion

Alternatives

Research methodology design for health literacy and medication adherence in aotearoa new zealandExplore the methodological design for researching health literacy and its impact on medication adherence among adults with chronic diseases in Aotearoa New Zealand.Business & ops · CC0-1.0Scientific critical thinkingEvaluate scientific claims and evidence quality. Use for assessing experimental design validity, identifying biases and confounders, applying evidence grading frameworks (GRADE, Cochrane Risk of Bias), or teaching critical analysis. Best for understanding evidence quality, identifying flaws. For formal peer review writing use peer-review.Science · MITAcademic research synthesizerAcademic research synthesis specialist. Use PROACTIVELY for comprehensive research on academic topics, literature reviews, technical investigations, and well-cited analysis combining multiple sources. <example>Context: A podcast episode needs a segment grounded in peer-reviewed evidence with formal citations. user: "Research the current state of transformer efficiency techniques for the episode, with proper academic citations." assistant: "I'll use the academic-research-synthesizer agent to search arXiv and Semantic Scholar, extract full-text findings via WebFetch, and produce a cited literature synthesis with confidence levels." <commentary>Use academic-research-synthesizer (not comprehensive-researcher) when the episode segment needs peer-reviewed sourcing, formal citation format, and explicit confidence tagging rather than general-purpose multi-source coverage.</commentary></example> <example>Context: The episode-orchestrator has routed a "literature review" request for a technical deep-dive segment. user: "Summarize the research landscape on federated learning privacy guarantees." assistant: "I'll invoke academic-research-synthesizer to systematically search academic sources, note peer-review status per source, and synthesize consensus vs. open debates."</example>Business & ops · MITAcademic researcherAcademic research specialist for scholarly sources, peer-reviewed papers, and academic literature. Use PROACTIVELY for research paper analysis, literature reviews, citation tracking, and academic methodology evaluation. <example>Context: The research-orchestrator has kicked off Phase 4 parallel research on 'efficacy of intermittent fasting' and needs peer-reviewed evidence. user: "Find the academic evidence on intermittent fasting outcomes." assistant: "I'll use the academic-researcher agent to search Semantic Scholar, PubMed, and OpenAlex for peer-reviewed studies and write structured findings to academic-research.md." <commentary>The request is specifically for scholarly/peer-reviewed evidence rather than general web coverage or code, so academic-researcher (not web-researcher or technical-researcher) is the right specialist.</commentary></example> <example>Context: The user wants a literature review comparing methodologies across studies on a topic. user: "Can you review the literature on transformer model interpretability and identify research gaps?" assistant: "Let me invoke the academic-researcher agent to pull foundational and recent papers, extract methodologies, and surface open research gaps." <commentary>Literature review, methodology extraction, and research-gap identification are core academic-researcher capabilities, distinct from technical-researcher's focus on code repositories and implementations.</commentary></example>Business & ops · MIT