Riveter - Structured Web Scraping
Web scraping with structured data extraction - define your output schema
How to use it
- Hit Copy the whole skill.
- Claude: ⋯ → Download .md, then Customize → Skills → Add → Upload skill.
ChatGPT: make a Project and paste it into Instructions.
Neither? Paste it at the top of a new chat — it works for that chat. - Describe your job in plain words. The AI follows the skill from there.
npx degit gooseworks-ai/goose-skills/skills/research-tools/capabilities/structured-scraping-riveter#main ~/.claude/skills/structured-scraping-riveterFor one project only, change the path to .claude/skills/structured-scraping-riveter.
Not working?
- Check which app you pasted it into — the steps above name the right one.
- Some skills need the paid tier of Claude or ChatGPT.
Paste into Claude, ChatGPT or Cursor.
Show the full text133 lines
Riveter - Structured Web Scraping
Setup
Read your credentials from ~/.gooseworks/credentials.json:
export GOOSEWORKS_API_KEY=$(python3 -c "import json;print(json.load(open('$HOME/.gooseworks/credentials.json'))['api_key'])")
export GOOSEWORKS_API_BASE=$(python3 -c "import json;print(json.load(open('$HOME/.gooseworks/credentials.json')).get('api_base','https://api.gooseworks.ai'))")
If ~/.gooseworks/credentials.json does not exist, tell the user to run: npx gooseworks login
All endpoints use Bearer auth: -H "Authorization: Bearer $GOOSEWORKS_API_KEY"
Scrape web pages and extract data into your defined structure.
Capabilities
- Scrape: Scrape a webpage and return the text content
- Run: Copy link Define the structure of your output directly in the API request
- Run data: Retrieve the processed data from a completed project run (free)
- Run status: Check the current status of a project run (free)
- Stop run: Stop a currently running project (free)
Usage
Scrape
Scrape a webpage and return the text content. This endpoint allows you to extract text content from any public webpage.
Parameters:
- url* (string) - Example: "https://example.com"
- proxy_country_code (string) - Optional two-character country code for proxy (e.g., 'us', 'gb', 'de')
- skip_cache (boolean) - Default: false. Set to true to bypass cache and always fetch fresh content
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"riveter","path":"/v1/scrape","body":{"url":"https://example.com/article"}}'
Run
Copy link Define the structure of your output directly in the API request. This endpoint allows you to define both your input data and output configuration in a single request.
Parameters:
- input* (object) - The input object contains your source data: Keys are column/attribute names Values are arrays of strings (all arrays must be the same length) Maximum 1000 rows per request
- output* (object) - The output object defines what data you want to extract: Keys are the names of attributes you want to extract Each attribute requires: prompt: Instructions for finding/extracting this data contexts: Array of input or other output attribute names this depends on. Optional Output Configuration Each output attribute can optionally include: format: Data type ('number', 'json', 'url', 'text', 'email', 'tag', 'date', 'boolean') format_details: Format-specific configuration (varies by format type). For json format, you can provide either a description (string) or a schema (JSON Schema object) or both. tools: Array of tools to use (['web_search', 'web_scrape', 'query_pdf', 'query_image']) max_tool_calls: Number of tool calls allowed (0-10) run_when: When to run this extraction ('always', 'any_filled', 'all_filled')
- run_key (string) - Custom identifier for this run (optional, will be generated if not provided)
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"riveter","path":"/v1/run"}'
"input": {
"urls": ["https://example.com/products"]
},
"output": {
"name": {"prompt": "Product name", "contexts": ["urls"]},
"price": {"prompt": "Product price", "contexts": ["urls"], "format": "number"}
}
}'
Run data (free)
Retrieve the processed data from a completed project run
Parameters:
- run_key* (string) - The run key (UUID) of the project run to retrieve data for
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"riveter","path":"/v1/run_data","query":{"run_key":"abc123"}}'
Run status (free)
Check the current status of a project run
Parameters:
- run_key* (string) - The run key (UUID) of the project run to check
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"riveter","path":"/v1/run_status","query":{"run_key":"abc123"}}'
Stop run (free)
Stop a currently running project. This will halt all processing and mark the run as stopped. Behavior: If the run is already stopped or success, returns success with current status. If the run is in progress, stops all pending cells and marks the run as stopped. Stopped runs cannot be resumed
Parameters:
- run_key* (string) - The run key (UUID) of the project run to stop
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"riveter","path":"/v1/stop_run","query":{"run_key":"abc123"}}'
Use Cases
- E-commerce Scraping: Extract product data in consistent format
- Job Listings: Gather job postings with structured fields
- News Aggregation: Extract articles with title, date, content
- Price Monitoring: Track prices across competitor sites
Discover More
For full endpoint details and parameters:
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/search \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt":"riveter API endpoints"}' List all endpoints
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/details \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"riveter","path":"/v1/scrape"}' # Get endpoint details
| 1 | |
| 2 | name structured-scraping-riveter |
| 3 | description Web scraping with structured data extraction - define your output schema |
| 4 | source orthogonal |
| 5 | |
| 6 | |
| 7 | |
| 8 | # Riveter - Structured Web Scraping |
| 9 | |
| 10 | ## Setup |
| 11 | |
| 12 | Read your credentials from ~/.gooseworks/credentials.json: |
| 13 | |
| 14 | export GOOSEWORKS_API_KEY=$(python3 -c "import json;print(json.load(open('$HOME/.gooseworks/credentials.json'))['api_key'])") |
| 15 | export GOOSEWORKS_API_BASE=$(python3 -c "import json;print(json.load(open('$HOME/.gooseworks/credentials.json')).get('api_base','https://api.gooseworks.ai'))") |
| 16 | |
| 17 | |
| 18 | If ~/.gooseworks/credentials.json does not exist, tell the user to run: `npx gooseworks login` |
| 19 | |
| 20 | All endpoints use Bearer auth: `-H "Authorization: Bearer $GOOSEWORKS_API_KEY"` |
| 21 | |
| 22 | |
| 23 | Scrape web pages and extract data into your defined structure. |
| 24 | |
| 25 | ## Capabilities |
| 26 | |
| 27 | **Scrape**: Scrape a webpage and return the text content |
| 28 | **Run**: Copy link Define the structure of your output directly in the API request |
| 29 | **Run data**: Retrieve the processed data from a completed project run (free) |
| 30 | **Run status**: Check the current status of a project run (free) |
| 31 | **Stop run**: Stop a currently running project (free) |
| 32 | |
| 33 | ## Usage |
| 34 | |
| 35 | ### Scrape |
| 36 | Scrape a webpage and return the text content. This endpoint allows you to extract text content from any public webpage. |
| 37 | |
| 38 | Parameters: |
| 39 | url* (string) - Example: "https://example.com" |
| 40 | proxy_country_code (string) - Optional two-character country code for proxy (e.g., 'us', 'gb', 'de') |
| 41 | skip_cache (boolean) - Default: false. Set to true to bypass cache and always fetch fresh content |
| 42 | |
| 43 | |
| 44 | curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \ |
| 45 | -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \ |
| 46 | -H "Content-Type: application/json" \ |
| 47 | -d '{"api":"riveter","path":"/v1/scrape","body":{"url":"https://example.com/article"}}' |
| 48 | |
| 49 | |
| 50 | ### Run |
| 51 | Copy link Define the structure of your output directly in the API request. This endpoint allows you to define both your input data and output configuration in a single request. |
| 52 | |
| 53 | Parameters: |
| 54 | input* (object) - The input object contains your source data: Keys are column/attribute names Values are arrays of strings (all arrays must be the same length) Maximum 1000 rows per request |
| 55 | output* (object) - The output object defines what data you want to extract: Keys are the names of attributes you want to extract Each attribute requires: prompt: Instructions for finding/extracting this data contexts: Array of input or other output attribute names this depends on. Optional Output Configuration Each output attribute can optionally include: format: Data type ('number', 'json', 'url', 'text', 'email', 'tag', 'date', 'boolean') format_details: Format-specific configuration (varies by format type). For json format, you can provide either a description (string) or a schema (JSON Schema object) or both. tools: Array of tools to use (['web_search', 'web_scrape', 'query_pdf', 'query_image']) max_tool_calls: Number of tool calls allowed (0-10) run_when: When to run this extraction ('always', 'any_filled', 'all_filled') |
| 56 | run_key (string) - Custom identifier for this run (optional, will be generated if not provided) |
| 57 | |
| 58 | |
| 59 | curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \ |
| 60 | -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \ |
| 61 | -H "Content-Type: application/json" \ |
| 62 | -d '{"api":"riveter","path":"/v1/run"}' |
| 63 | "input": { |
| 64 | "urls": ["https://example.com/products"] |
| 65 | }, |
| 66 | "output": { |
| 67 | "name": {"prompt": "Product name", "contexts": ["urls"]}, |
| 68 | "price": {"prompt": "Product price", "contexts": ["urls"], "format": "number"} |
| 69 | } |
| 70 | }' |
| 71 | |
| 72 | |
| 73 | ### Run data (free) |
| 74 | Retrieve the processed data from a completed project run |
| 75 | |
| 76 | Parameters: |
| 77 | run_key* (string) - The run key (UUID) of the project run to retrieve data for |
| 78 | |
| 79 | |
| 80 | curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \ |
| 81 | -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \ |
| 82 | -H "Content-Type: application/json" \ |
| 83 | -d '{"api":"riveter","path":"/v1/run_data","query":{"run_key":"abc123"}}' |
| 84 | |
| 85 | |
| 86 | ### Run status (free) |
| 87 | Check the current status of a project run |
| 88 | |
| 89 | Parameters: |
| 90 | run_key* (string) - The run key (UUID) of the project run to check |
| 91 | |
| 92 | |
| 93 | curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \ |
| 94 | -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \ |
| 95 | -H "Content-Type: application/json" \ |
| 96 | -d '{"api":"riveter","path":"/v1/run_status","query":{"run_key":"abc123"}}' |
| 97 | |
| 98 | |
| 99 | ### Stop run (free) |
| 100 | Stop a currently running project. This will halt all processing and mark the run as stopped. Behavior: If the run is already stopped or success, returns success with current status. If the run is in progress, stops all pending cells and marks the run as stopped. Stopped runs cannot be resumed |
| 101 | |
| 102 | Parameters: |
| 103 | run_key* (string) - The run key (UUID) of the project run to stop |
| 104 | |
| 105 | |
| 106 | curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \ |
| 107 | -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \ |
| 108 | -H "Content-Type: application/json" \ |
| 109 | -d '{"api":"riveter","path":"/v1/stop_run","query":{"run_key":"abc123"}}' |
| 110 | |
| 111 | |
| 112 | ## Use Cases |
| 113 | |
| 114 | **E-commerce Scraping**: Extract product data in consistent format |
| 115 | **Job Listings**: Gather job postings with structured fields |
| 116 | **News Aggregation**: Extract articles with title, date, content |
| 117 | **Price Monitoring**: Track prices across competitor sites |
| 118 | |
| 119 | ## Discover More |
| 120 | |
| 121 | For full endpoint details and parameters: |
| 122 | |
| 123 | |
| 124 | curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/search \ |
| 125 | -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \ |
| 126 | -H "Content-Type: application/json" \ |
| 127 | -d '{"prompt":"riveter API endpoints"}' List all endpoints |
| 128 | curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/details \ |
| 129 | -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \ |
| 130 | -H "Content-Type: application/json" \ |
| 131 | -d '{"api":"riveter","path":"/v1/scrape"}' # Get endpoint details |
| 132 | |
| 133 |