Brightdata skill

4-tier progressive web scraping that auto-escalates WebFetch to curl to Interceptor to Bright Data proxy for bot detection and CAPTCHAs, with single-URL and multi-page crawl modes, output as markdown.

by danielmiessler·MIT license·★ 19,269 Stars on the repo·GitHub ↗

Use now

Files of Brightdata

danielmiessler/main1 file shown
SKILL.md
Show the full text76 lines

Customization

Before executing, check for user customizations at: ~/.claude/LIFEOS/USER/CUSTOMIZATIONS/SKILLS/BrightData/

If this directory exists, load and apply any PREFERENCES.md, configurations, or resources found there. These override default behavior. If the directory does not exist, proceed with skill defaults.

🚨 MANDATORY: Voice Notification (REQUIRED BEFORE ANY ACTION)

You MUST send this notification BEFORE doing anything else when this skill is invoked.

  1. Send voice notification:

    curl -s -X POST http://localhost:31337/notify \
      -H "Content-Type: application/json" \
      -d '{"message": "Running the WORKFLOWNAME workflow in the BrightData skill to ACTION"}' \
      > /dev/null 2>&1 &
    
  2. Output text notification:

    Running the **WorkflowName** workflow in the **BrightData** skill to ACTION...
    

This is not optional. Execute this curl command immediately upon skill invocation.

BrightData

Scrapes a single URL (FourTierScrape) or crawls a whole site (Crawl), escalating through four tiers only as far as each page needs. Output is always markdown. Start at Tier 1 and step up only when blocked — reaching for the heavy proxy every time wastes Tier-4 credits. A Cloudflare Accept: text/markdown pre-check runs before Tier 1 (recipe in FourTierScrape.md).

The four tiers (tool contract)

Tier Tool Wins on Cost / latency
1 WebFetch public content, no bot detection free · ~2-5s
2 curl + Chrome headers user-agent / basic header checks free · ~3-7s
3 Interceptor (real Chrome) JavaScript-rendered / SPA pages free · ~10-20s
4 Bright Data MCP mcp__Brightdata__scrape_as_markdown CAPTCHA, advanced fingerprinting, residential-IP needs Bright Data credits · ~5-15s

Playwright is banned across LifeOS — Tier 3 is Interceptor. Skip-ahead: explicit "use Bright Data" → Tier 4; "use browser" → Tier 3; a domain that already failed Tier 1 → start at Tier 2. The exact curl header block, Cloudflare pre-check, and Interceptor commands live in Workflows/FourTierScrape.md.

Workflows

When routing, output: Running the **WorkflowName** workflow in the **BrightData** skill to ACTION...

Workflow Trigger File
FourTierScrape "scrape/fetch/pull/get/retrieve [URL]", "can't access this site", "site is blocking me", "use Bright Data to fetch" Workflows/FourTierScrape.md
Crawl "crawl this site", "spider this domain", "map this website", "get all pages from", "scrape the whole site", "crawl all pages under /docs" Workflows/Crawl.md

Crawl picks Light Crawl (MCP scrape_batch + link loop, ≤50 pages, ~$0.006/page) for a section, or Full Crawl (Bright Data Crawl API api.brightdata.com/datasets/v3/trigger, $1.50/1K pages) for whole sites.

Gotchas

  • 4-tier escalation: WebFetch → curl → Interceptor → Bright Data proxy. Always start at Tier 1 and escalate only when blocked. Playwright is banned across LifeOS.
  • Bright Data proxy has usage costs. Don't use Tier 4 for sites accessible via Tier 1-3.
  • CAPTCHA-solving introduces latency. Allow extra time for Tier 4 responses.
  • Credentials in ~/.claude/.env — BRIGHTDATA_API_KEY.

Execution Log

After completing any workflow, append a single JSONL entry:

echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"BrightData","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl

Replace WORKFLOW_USED with the workflow executed, 8_WORD_SUMMARY with a brief input description, and SECONDS with approximate wall-clock time. Log status: "error" if the workflow failed.

1---
2name: BrightData
3version: 1.2.20
4description: "4-tier progressive web scraping that auto-escalates WebFetch to curl to Interceptor to Bright Data proxy for bot detection and CAPTCHAs, with single-URL and multi-page crawl modes, output as markdown. USE WHEN Bright Data, scrape URL, web scraping, bot detection, crawl site, CAPTCHA, can't access, site blocking, extract page content, scrape whole site, spider domain, convert URL to markdown, getting blocked. NOT FOR simple public content (use WebFetch directly), social platform scraping with named actors (use Apify), or real-Chrome bot bypass with logged-in sessions and zero CDP fingerprint (use Interceptor)."
5---
6 
7## Customization
8 
9**Before executing, check for user customizations at:**
10`~/.claude/LIFEOS/USER/CUSTOMIZATIONS/SKILLS/BrightData/`
11 
12If this directory exists, load and apply any PREFERENCES.md, configurations, or resources found there. These override default behavior. If the directory does not exist, proceed with skill defaults.
13 
14 
15## 🚨 MANDATORY: Voice Notification (REQUIRED BEFORE ANY ACTION)
16 
17**You MUST send this notification BEFORE doing anything else when this skill is invoked.**
18 
191. **Send voice notification**:
20 ```bash
21 curl -s -X POST http://localhost:31337/notify \
22 -H "Content-Type: application/json" \
23 -d '{"message": "Running the WORKFLOWNAME workflow in the BrightData skill to ACTION"}' \
24 > /dev/null 2>&1 &
25 ```
26 
272. **Output text notification**:
28 ```
29 Running the **WorkflowName** workflow in the **BrightData** skill to ACTION...
30 ```
31 
32**This is not optional. Execute this curl command immediately upon skill invocation.**
33 
34# BrightData
35 
36Scrapes a single URL (FourTierScrape) or crawls a whole site (Crawl), escalating through four tiers only as far as each page needs. Output is always markdown. Start at Tier 1 and step up only when blocked — reaching for the heavy proxy every time wastes Tier-4 credits. A Cloudflare `Accept: text/markdown` pre-check runs before Tier 1 (recipe in FourTierScrape.md).
37 
38## The four tiers (tool contract)
39 
40| Tier | Tool | Wins on | Cost / latency |
41|------|------|---------|----------------|
42| 1 | WebFetch | public content, no bot detection | free · ~2-5s |
43| 2 | curl + Chrome headers | user-agent / basic header checks | free · ~3-7s |
44| 3 | Interceptor (real Chrome) | JavaScript-rendered / SPA pages | free · ~10-20s |
45| 4 | Bright Data MCP `mcp__Brightdata__scrape_as_markdown` | CAPTCHA, advanced fingerprinting, residential-IP needs | Bright Data credits · ~5-15s |
46 
47Playwright is banned across LifeOS — Tier 3 is Interceptor. Skip-ahead: explicit "use Bright Data" → Tier 4; "use browser" → Tier 3; a domain that already failed Tier 1 → start at Tier 2. The exact curl header block, Cloudflare pre-check, and Interceptor commands live in `Workflows/FourTierScrape.md`.
48 
49## Workflows
50 
51When routing, output: `Running the **WorkflowName** workflow in the **BrightData** skill to ACTION...`
52 
53| Workflow | Trigger | File |
54|----------|---------|------|
55| FourTierScrape | "scrape/fetch/pull/get/retrieve [URL]", "can't access this site", "site is blocking me", "use Bright Data to fetch" | `Workflows/FourTierScrape.md` |
56| Crawl | "crawl this site", "spider this domain", "map this website", "get all pages from", "scrape the whole site", "crawl all pages under /docs" | `Workflows/Crawl.md` |
57 
58Crawl picks Light Crawl (MCP `scrape_batch` + link loop, ≤50 pages, ~$0.006/page) for a section, or Full Crawl (Bright Data Crawl API `api.brightdata.com/datasets/v3/trigger`, $1.50/1K pages) for whole sites.
59 
60## Gotchas
61 
62- **4-tier escalation: WebFetch → curl → Interceptor → Bright Data proxy.** Always start at Tier 1 and escalate only when blocked. Playwright is banned across LifeOS.
63- **Bright Data proxy has usage costs.** Don't use Tier 4 for sites accessible via Tier 1-3.
64- **CAPTCHA-solving introduces latency.** Allow extra time for Tier 4 responses.
65- **Credentials in `~/.claude/.env`** — BRIGHTDATA_API_KEY.
66 
67## Execution Log
68 
69After completing any workflow, append a single JSONL entry:
70 
71```bash
72echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"BrightData","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl
73```
74 
75Replace `WORKFLOW_USED` with the workflow executed, `8_WORD_SUMMARY` with a brief input description, and `SECONDS` with approximate wall-clock time. Log `status: "error"` if the workflow failed.
76 

Discussion

Alternatives

AI Crawler Access Analysis SkillAI crawler access analysis. Checks robots.txt, meta tags, and HTTP headers to determine which AI crawlers can access the site. Provides a complete access map and recommendations for maximizing AI visibility while maintaining appropriate control.Data & AI · MITBrowser Automation - POWERFULUse when the user asks to automate browser tasks, scrape websites, fill forms, capture screenshots, extract structured data from web pages, or build web automation workflows. NOT for testing — use playwright-pro for that.Data & AI · MITNotebookLM — Browser AutomationBrowser automation skill for controlling Google's NotebookLM. Use when the user wants anything done in NotebookLM (e.g., 'open NotebookLM', 'check my [name] notebook', 'ask my notebook about X', 'add [source] to NotebookLM', 'generate a Video Overview from my notebook', 'use NotebookLM Studio'). Handles reading and querying notebooks, adding sources (URLs, text, files, YouTube links, synthesized content), generating Studio outputs (Audio/Video Overviews, Mind Maps, Reports incl. Briefing Doc/Study Guide/FAQ, Flashcards, Quiz, slide decks, infographics — discover the exact set from the live Studio panel; the UI evolves fast), and creating new notebooks. Requires browser automation environment — fails gracefully when unavailable.Data & AI · MITBlog feed monitorScrape blog posts via RSS feeds (free, no API key) with Apify fallback for JS-heavy sites. Use when you need to monitor competitor blogs, track industry content, or aggregate blog posts by keyword.Data & AI · MIT