Universal scraping architect
Use for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts.
How to use it
Claude Code
- Run the line below. It pulls the whole folder into
~/.claude/skills/universal-scraping-architect, including the files SKILL.md points to. - Describe your job in plain words. Claude Code follows the skill from there.
npx degit alirezarezvani/claude-skills/engineering/universal-scraping-architect/skills/universal-scraping-architect#main ~/.claude/skills/universal-scraping-architectFor one project only, change the path to .claude/skills/universal-scraping-architect. This skill also uses output.json, project-context.md, extracted_output.json, robots.txt — copying SKILL.md alone won't be enough. See the folder on GitHub.
Claude (web or desktop app)
- On this page open ⋯ → Download .md.
- Save it as SKILL.md in a folder, zip the folder, then Customize → Skills → + → Create skill → Upload a skill.
- Pick the file and Save. Claude shows the name and description and runs a security scan.
- Check the skill is switched on.
- Start a new chat and describe your job in plain words. The AI follows the skill from there.
ChatGPT or another app
- ChatGPT: make a Project and paste it into Instructions.
- Neither? Paste it at the top of a new chat — it works for that chat.
Not working?
- Check which app you pasted it into — the steps above name the right one.
- Some skills need the paid tier of Claude or ChatGPT.
Paste into Claude, ChatGPT or Cursor.
Source of Universal scraping architect
Show the full text65 lines
| name | description |
|---|---|
| universal-scraping-architect | Use for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts. |
Universal Scraping Architect
Design complete, robust data-extraction pipelines with intelligent routing, validation, and token-budget tracking — not brittle one-off scripts.
Dependency Notice: BYOK (Bring Your Own Key) pattern for Firecrawl; API keys must only be loaded via environment variables. Per-script dependencies:
| Script | Dependencies | Exact CLI |
|---|---|---|
scripts/validate_extraction.py |
stdlib only | python3 scripts/validate_extraction.py output.json --json |
scripts/firecrawl_example.py |
firecrawl, requests (template; --sample runs offline) |
python3 scripts/firecrawl_example.py --sample |
scripts/local_bs4_example.py |
beautifulsoup4, pandas (template; --sample runs offline) |
python3 scripts/local_bs4_example.py --sample |
Before Starting
Check for context first:
If project-context.md exists, read it before asking questions. Determine the target data format, scale of extraction, and deployment environment before writing any code.
How This Skill Works
This skill supports 3 extraction modes based on intelligent routing:
Mode 1: API-Driven (Firecrawl)
Use when the source is a public URL, heavily dynamic (JS/SPA), requires search-first discovery, or involves bulk crawling across a domain.
Mode 2: Local Python (Traditional)
Use when extracting from local files (PDF, Excel, CSV), the data is private/sensitive, or the target is a simple static HTML page where Firecrawl is overkill.
Mode 3: Hybrid Pipeline
Use when Firecrawl handles URL discovery/web extraction, but local Python (Pandas) is required to clean, normalize, and structure the output before saving.
The Extraction Pipeline
When executing a scraping task, always follow this sequence:
- Route the Approach: Explicitly state whether Firecrawl or Local Python is being used and why.
- Track Budgets: Estimate Firecrawl API quotas or LLM token context limits before executing large jobs.
- Extract Safely: Implement checkpointing for multi-page jobs. Handle pagination and dynamic layouts gracefully. Start from the editable runner templates —
scripts/firecrawl_example.py(Mode 1) orscripts/local_bs4_example.py(Mode 2); run each with--samplefirst to see the expected summary shape without network access. - Validate & Clean: Run
python3 scripts/validate_extraction.py extracted_output.json --jsonon every extraction result before delivering it. It exits 0 only on{"status": "ok"};warning(empty output) orerror(malformed JSON) exit 1 — fix and re-extract, never ship unvalidated data. Beyond this structural gate, also check required fields and duplicates against the pipeline spec before delivering. - Format: Default to CSV for tabular data, JSON for nested structures, and Markdown for clean text.
Proactive Triggers
Surface these issues WITHOUT being asked when you notice them in context:
- Hardcoded API Keys → Flag immediately and rewrite to use
os.getenv('FIRECRAWL_API_KEY'). - Private Data Leakage → If the user asks to send local, sensitive files to an external API, flag the privacy risk and suggest Mode 2 (Local Python).
- Missing Pagination → If the target implies hundreds of records but no pagination logic is requested, flag it and add checkpointing.
Output Artifacts
| When you ask for... | You get... |
|---|---|
| "Scrape this site" | A fully validated Python extraction script with routing logic and error handling. |
| "Get data from this table" | A clean CSV/JSON dataset with a summary log of row counts and empty values. |
| "Crawl these docs" | A Markdown deliverable chunked for LLM token limits. |
Anti-Patterns
- Brittle Selectors: Never use highly nested CSS selectors (e.g.,
div > span > ul > li:nth-child(3)). Use data attributes or robust structural anchors. - Ignoring Etiquette: Never scrape without checking
robots.txtor implementing sensible rate limits. - No Validation: Never blindly write scraped data to a file without checking if the array is empty or missing critical keys.
Related Skills
- data-cleaning: Use when the scraped data requires complex statistical normalization or deduplication.
- browser-automation: Use for highly interactive scraping requiring user emulation (clicks, logins) where Firecrawl is insufficient.
| 1 | |
| 2 | name "universal-scraping-architect" |
| 3 | description "Use for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts." |
| 4 | |
| 5 | |
| 6 | # Universal Scraping Architect |
| 7 | |
| 8 | Design complete, robust data-extraction pipelines with intelligent routing, validation, and token-budget tracking — not brittle one-off scripts. |
| 9 | |
| 10 | **Dependency Notice:** BYOK (Bring Your Own Key) pattern for Firecrawl; API keys must only be loaded via environment variables. Per-script dependencies: |
| 11 | |
| 12 | | Script | Dependencies | Exact CLI | |
| 13 | |---|---|---| |
| 14 | | `scripts/validate_extraction.py` | stdlib only | `python3 scripts/validate_extraction.py output.json --json` | |
| 15 | | `scripts/firecrawl_example.py` | `firecrawl`, `requests` (template; `--sample` runs offline) | `python3 scripts/firecrawl_example.py --sample` | |
| 16 | | `scripts/local_bs4_example.py` | `beautifulsoup4`, `pandas` (template; `--sample` runs offline) | `python3 scripts/local_bs4_example.py --sample` | |
| 17 | |
| 18 | ## Before Starting |
| 19 | **Check for context first:** |
| 20 | If `project-context.md` exists, read it before asking questions. Determine the target data format, scale of extraction, and deployment environment before writing any code. |
| 21 | |
| 22 | ## How This Skill Works |
| 23 | |
| 24 | This skill supports 3 extraction modes based on intelligent routing: |
| 25 | |
| 26 | ### Mode 1: API-Driven (Firecrawl) |
| 27 | Use when the source is a public URL, heavily dynamic (JS/SPA), requires search-first discovery, or involves bulk crawling across a domain. |
| 28 | ### Mode 2: Local Python (Traditional) |
| 29 | Use when extracting from local files (PDF, Excel, CSV), the data is private/sensitive, or the target is a simple static HTML page where Firecrawl is overkill. |
| 30 | ### Mode 3: Hybrid Pipeline |
| 31 | Use when Firecrawl handles URL discovery/web extraction, but local Python (Pandas) is required to clean, normalize, and structure the output before saving. |
| 32 | |
| 33 | ## The Extraction Pipeline |
| 34 | |
| 35 | When executing a scraping task, always follow this sequence: |
| 36 | **Route the Approach:** Explicitly state whether Firecrawl or Local Python is being used and why. |
| 37 | **Track Budgets:** Estimate Firecrawl API quotas or LLM token context limits before executing large jobs. |
| 38 | **Extract Safely:** Implement checkpointing for multi-page jobs. Handle pagination and dynamic layouts gracefully. Start from the editable runner templates — `scripts/firecrawl_example.py` (Mode 1) or `scripts/local_bs4_example.py` (Mode 2); run each with `--sample` first to see the expected summary shape without network access. |
| 39 | **Validate & Clean:** Run `python3 scripts/validate_extraction.py extracted_output.json --json` on every extraction result before delivering it. It exits 0 only on `{"status": "ok"}`; `warning` (empty output) or `error` (malformed JSON) exit 1 — fix and re-extract, never ship unvalidated data. Beyond this structural gate, also check required fields and duplicates against the pipeline spec before delivering. |
| 40 | **Format:** Default to CSV for tabular data, JSON for nested structures, and Markdown for clean text. |
| 41 | |
| 42 | ## Proactive Triggers |
| 43 | |
| 44 | Surface these issues WITHOUT being asked when you notice them in context: |
| 45 | **Hardcoded API Keys** → Flag immediately and rewrite to use `os.getenv('FIRECRAWL_API_KEY')`. |
| 46 | **Private Data Leakage** → If the user asks to send local, sensitive files to an external API, flag the privacy risk and suggest Mode 2 (Local Python). |
| 47 | **Missing Pagination** → If the target implies hundreds of records but no pagination logic is requested, flag it and add checkpointing. |
| 48 | |
| 49 | ## Output Artifacts |
| 50 | |
| 51 | | When you ask for... | You get... | |
| 52 | |---------------------|------------| |
| 53 | | "Scrape this site" | A fully validated Python extraction script with routing logic and error handling. | |
| 54 | | "Get data from this table" | A clean CSV/JSON dataset with a summary log of row counts and empty values. | |
| 55 | | "Crawl these docs" | A Markdown deliverable chunked for LLM token limits. | |
| 56 | |
| 57 | ## Anti-Patterns |
| 58 | **Brittle Selectors:** Never use highly nested CSS selectors (e.g., `div > span > ul > li:nth-child(3)`). Use data attributes or robust structural anchors. |
| 59 | **Ignoring Etiquette:** Never scrape without checking `robots.txt` or implementing sensible rate limits. |
| 60 | **No Validation:** Never blindly write scraped data to a file without checking if the array is empty or missing critical keys. |
| 61 | |
| 62 | ## Related Skills |
| 63 | **data-cleaning**: Use when the scraped data requires complex statistical normalization or deduplication. |
| 64 | **browser-automation**: Use for highly interactive scraping requiring user emulation (clicks, logins) where Firecrawl is insufficient. |
| 65 |
Discussion
Browse more free Claude skills.