Universal scraping architect

Use for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts.

How to use it

Claude Code
  1. Run the line below. It pulls the whole folder into ~/.claude/skills/universal-scraping-architect, including the files SKILL.md points to.
  2. Describe your job in plain words. Claude Code follows the skill from there.
Claude Code — installs the whole folder, not just SKILL.md
npx degit alirezarezvani/claude-skills/engineering/universal-scraping-architect/skills/universal-scraping-architect#main ~/.claude/skills/universal-scraping-architect

For one project only, change the path to .claude/skills/universal-scraping-architect. This skill also uses output.json, project-context.md, extracted_output.json, robots.txt — copying SKILL.md alone won't be enough. See the folder on GitHub.

Claude (web or desktop app)
  1. On this page open ⋯ → Download .md.
  2. Save it as SKILL.md in a folder, zip the folder, then Customize → Skills → + → Create skill → Upload a skill.
  3. Pick the file and Save. Claude shows the name and description and runs a security scan.
  4. Check the skill is switched on.
  5. Start a new chat and describe your job in plain words. The AI follows the skill from there.
ChatGPT or another app
  1. ChatGPT: make a Project and paste it into Instructions.
  2. Neither? Paste it at the top of a new chat — it works for that chat.
Not working?
  • Check which app you pasted it into — the steps above name the right one.
  • Some skills need the paid tier of Claude or ChatGPT.
Step-by-step guide with screenshots · Ask in the forum

Paste into Claude, ChatGPT or Cursor.

Source of Universal scraping architect

Show the full text65 lines
namedescription
universal-scraping-architectUse for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts.

Universal Scraping Architect

Design complete, robust data-extraction pipelines with intelligent routing, validation, and token-budget tracking — not brittle one-off scripts.

Dependency Notice: BYOK (Bring Your Own Key) pattern for Firecrawl; API keys must only be loaded via environment variables. Per-script dependencies:

Script Dependencies Exact CLI
scripts/validate_extraction.py stdlib only python3 scripts/validate_extraction.py output.json --json
scripts/firecrawl_example.py firecrawl, requests (template; --sample runs offline) python3 scripts/firecrawl_example.py --sample
scripts/local_bs4_example.py beautifulsoup4, pandas (template; --sample runs offline) python3 scripts/local_bs4_example.py --sample

Before Starting

Check for context first: If project-context.md exists, read it before asking questions. Determine the target data format, scale of extraction, and deployment environment before writing any code.

How This Skill Works

This skill supports 3 extraction modes based on intelligent routing:

Mode 1: API-Driven (Firecrawl)

Use when the source is a public URL, heavily dynamic (JS/SPA), requires search-first discovery, or involves bulk crawling across a domain.

Mode 2: Local Python (Traditional)

Use when extracting from local files (PDF, Excel, CSV), the data is private/sensitive, or the target is a simple static HTML page where Firecrawl is overkill.

Mode 3: Hybrid Pipeline

Use when Firecrawl handles URL discovery/web extraction, but local Python (Pandas) is required to clean, normalize, and structure the output before saving.

The Extraction Pipeline

When executing a scraping task, always follow this sequence:

  1. Route the Approach: Explicitly state whether Firecrawl or Local Python is being used and why.
  2. Track Budgets: Estimate Firecrawl API quotas or LLM token context limits before executing large jobs.
  3. Extract Safely: Implement checkpointing for multi-page jobs. Handle pagination and dynamic layouts gracefully. Start from the editable runner templates — scripts/firecrawl_example.py (Mode 1) or scripts/local_bs4_example.py (Mode 2); run each with --sample first to see the expected summary shape without network access.
  4. Validate & Clean: Run python3 scripts/validate_extraction.py extracted_output.json --json on every extraction result before delivering it. It exits 0 only on {"status": "ok"}; warning (empty output) or error (malformed JSON) exit 1 — fix and re-extract, never ship unvalidated data. Beyond this structural gate, also check required fields and duplicates against the pipeline spec before delivering.
  5. Format: Default to CSV for tabular data, JSON for nested structures, and Markdown for clean text.

Proactive Triggers

Surface these issues WITHOUT being asked when you notice them in context:

  • Hardcoded API Keys → Flag immediately and rewrite to use os.getenv('FIRECRAWL_API_KEY').
  • Private Data Leakage → If the user asks to send local, sensitive files to an external API, flag the privacy risk and suggest Mode 2 (Local Python).
  • Missing Pagination → If the target implies hundreds of records but no pagination logic is requested, flag it and add checkpointing.

Output Artifacts

When you ask for... You get...
"Scrape this site" A fully validated Python extraction script with routing logic and error handling.
"Get data from this table" A clean CSV/JSON dataset with a summary log of row counts and empty values.
"Crawl these docs" A Markdown deliverable chunked for LLM token limits.

Anti-Patterns

  • Brittle Selectors: Never use highly nested CSS selectors (e.g., div > span > ul > li:nth-child(3)). Use data attributes or robust structural anchors.
  • Ignoring Etiquette: Never scrape without checking robots.txt or implementing sensible rate limits.
  • No Validation: Never blindly write scraped data to a file without checking if the array is empty or missing critical keys.
  • data-cleaning: Use when the scraped data requires complex statistical normalization or deduplication.
  • browser-automation: Use for highly interactive scraping requiring user emulation (clicks, logins) where Firecrawl is insufficient.
1---
2name: "universal-scraping-architect"
3description: "Use for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts."
4---
5 
6# Universal Scraping Architect
7 
8Design complete, robust data-extraction pipelines with intelligent routing, validation, and token-budget tracking — not brittle one-off scripts.
9 
10**Dependency Notice:** BYOK (Bring Your Own Key) pattern for Firecrawl; API keys must only be loaded via environment variables. Per-script dependencies:
11 
12| Script | Dependencies | Exact CLI |
13|---|---|---|
14| `scripts/validate_extraction.py` | stdlib only | `python3 scripts/validate_extraction.py output.json --json` |
15| `scripts/firecrawl_example.py` | `firecrawl`, `requests` (template; `--sample` runs offline) | `python3 scripts/firecrawl_example.py --sample` |
16| `scripts/local_bs4_example.py` | `beautifulsoup4`, `pandas` (template; `--sample` runs offline) | `python3 scripts/local_bs4_example.py --sample` |
17 
18## Before Starting
19**Check for context first:**
20If `project-context.md` exists, read it before asking questions. Determine the target data format, scale of extraction, and deployment environment before writing any code.
21 
22## How This Skill Works
23 
24This skill supports 3 extraction modes based on intelligent routing:
25 
26### Mode 1: API-Driven (Firecrawl)
27Use when the source is a public URL, heavily dynamic (JS/SPA), requires search-first discovery, or involves bulk crawling across a domain.
28### Mode 2: Local Python (Traditional)
29Use when extracting from local files (PDF, Excel, CSV), the data is private/sensitive, or the target is a simple static HTML page where Firecrawl is overkill.
30### Mode 3: Hybrid Pipeline
31Use when Firecrawl handles URL discovery/web extraction, but local Python (Pandas) is required to clean, normalize, and structure the output before saving.
32 
33## The Extraction Pipeline
34 
35When executing a scraping task, always follow this sequence:
361. **Route the Approach:** Explicitly state whether Firecrawl or Local Python is being used and why.
372. **Track Budgets:** Estimate Firecrawl API quotas or LLM token context limits before executing large jobs.
383. **Extract Safely:** Implement checkpointing for multi-page jobs. Handle pagination and dynamic layouts gracefully. Start from the editable runner templates — `scripts/firecrawl_example.py` (Mode 1) or `scripts/local_bs4_example.py` (Mode 2); run each with `--sample` first to see the expected summary shape without network access.
394. **Validate & Clean:** Run `python3 scripts/validate_extraction.py extracted_output.json --json` on every extraction result before delivering it. It exits 0 only on `{"status": "ok"}`; `warning` (empty output) or `error` (malformed JSON) exit 1 — fix and re-extract, never ship unvalidated data. Beyond this structural gate, also check required fields and duplicates against the pipeline spec before delivering.
405. **Format:** Default to CSV for tabular data, JSON for nested structures, and Markdown for clean text.
41 
42## Proactive Triggers
43 
44Surface these issues WITHOUT being asked when you notice them in context:
45- **Hardcoded API Keys** → Flag immediately and rewrite to use `os.getenv('FIRECRAWL_API_KEY')`.
46- **Private Data Leakage** → If the user asks to send local, sensitive files to an external API, flag the privacy risk and suggest Mode 2 (Local Python).
47- **Missing Pagination** → If the target implies hundreds of records but no pagination logic is requested, flag it and add checkpointing.
48 
49## Output Artifacts
50 
51| When you ask for... | You get... |
52|---------------------|------------|
53| "Scrape this site" | A fully validated Python extraction script with routing logic and error handling. |
54| "Get data from this table" | A clean CSV/JSON dataset with a summary log of row counts and empty values. |
55| "Crawl these docs" | A Markdown deliverable chunked for LLM token limits. |
56 
57## Anti-Patterns
58- **Brittle Selectors:** Never use highly nested CSS selectors (e.g., `div > span > ul > li:nth-child(3)`). Use data attributes or robust structural anchors.
59- **Ignoring Etiquette:** Never scrape without checking `robots.txt` or implementing sensible rate limits.
60- **No Validation:** Never blindly write scraped data to a file without checking if the array is empty or missing critical keys.
61 
62## Related Skills
63- **data-cleaning**: Use when the scraped data requires complex statistical normalization or deduplication.
64- **browser-automation**: Use for highly interactive scraping requiring user emulation (clicks, logins) where Firecrawl is insufficient.
65 

Discussion