Riveter - Structured Web Scraping

Web scraping with structured data extraction - define your output schema

How to use it

  1. Hit Copy the whole skill.
  2. Claude: ⋯ → Download .md, then Customize → Skills → Add → Upload skill.
    ChatGPT: make a Project and paste it into Instructions.
    Neither? Paste it at the top of a new chat — it works for that chat.
  3. Describe your job in plain words. The AI follows the skill from there.
Claude Code — installs the whole folder, not just SKILL.md
npx degit gooseworks-ai/goose-skills/skills/research-tools/capabilities/structured-scraping-riveter#main ~/.claude/skills/structured-scraping-riveter

For one project only, change the path to .claude/skills/structured-scraping-riveter.

Not working?
  • Check which app you pasted it into — the steps above name the right one.
  • Some skills need the paid tier of Claude or ChatGPT.
Step-by-step guide with screenshots · Ask in the forum

Paste into Claude, ChatGPT or Cursor.

Show the full text133 lines
structured-scraping-riveter/SKILL.md133 lines5.7 KBpushed 96d agoRawView on GitHub

Riveter - Structured Web Scraping

Setup

Read your credentials from ~/.gooseworks/credentials.json:

export GOOSEWORKS_API_KEY=$(python3 -c "import json;print(json.load(open('$HOME/.gooseworks/credentials.json'))['api_key'])")
export GOOSEWORKS_API_BASE=$(python3 -c "import json;print(json.load(open('$HOME/.gooseworks/credentials.json')).get('api_base','https://api.gooseworks.ai'))")

If ~/.gooseworks/credentials.json does not exist, tell the user to run: npx gooseworks login

All endpoints use Bearer auth: -H "Authorization: Bearer $GOOSEWORKS_API_KEY"

Scrape web pages and extract data into your defined structure.

Capabilities

  • Scrape: Scrape a webpage and return the text content
  • Run: Copy link Define the structure of your output directly in the API request
  • Run data: Retrieve the processed data from a completed project run (free)
  • Run status: Check the current status of a project run (free)
  • Stop run: Stop a currently running project (free)

Usage

Scrape

Scrape a webpage and return the text content. This endpoint allows you to extract text content from any public webpage.

Parameters:

  • url* (string) - Example: "https://example.com"
  • proxy_country_code (string) - Optional two-character country code for proxy (e.g., 'us', 'gb', 'de')
  • skip_cache (boolean) - Default: false. Set to true to bypass cache and always fetch fresh content
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"api":"riveter","path":"/v1/scrape","body":{"url":"https://example.com/article"}}'

Run

Copy link Define the structure of your output directly in the API request. This endpoint allows you to define both your input data and output configuration in a single request.

Parameters:

  • input* (object) - The input object contains your source data: Keys are column/attribute names Values are arrays of strings (all arrays must be the same length) Maximum 1000 rows per request
  • output* (object) - The output object defines what data you want to extract: Keys are the names of attributes you want to extract Each attribute requires: prompt: Instructions for finding/extracting this data contexts: Array of input or other output attribute names this depends on. Optional Output Configuration Each output attribute can optionally include: format: Data type ('number', 'json', 'url', 'text', 'email', 'tag', 'date', 'boolean') format_details: Format-specific configuration (varies by format type). For json format, you can provide either a description (string) or a schema (JSON Schema object) or both. tools: Array of tools to use (['web_search', 'web_scrape', 'query_pdf', 'query_image']) max_tool_calls: Number of tool calls allowed (0-10) run_when: When to run this extraction ('always', 'any_filled', 'all_filled')
  • run_key (string) - Custom identifier for this run (optional, will be generated if not provided)
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"api":"riveter","path":"/v1/run"}'
  "input": {
    "urls": ["https://example.com/products"]
  },
  "output": {
    "name": {"prompt": "Product name", "contexts": ["urls"]},
    "price": {"prompt": "Product price", "contexts": ["urls"], "format": "number"}
  }
}'

Run data (free)

Retrieve the processed data from a completed project run

Parameters:

  • run_key* (string) - The run key (UUID) of the project run to retrieve data for
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"api":"riveter","path":"/v1/run_data","query":{"run_key":"abc123"}}'

Run status (free)

Check the current status of a project run

Parameters:

  • run_key* (string) - The run key (UUID) of the project run to check
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"api":"riveter","path":"/v1/run_status","query":{"run_key":"abc123"}}'

Stop run (free)

Stop a currently running project. This will halt all processing and mark the run as stopped. Behavior: If the run is already stopped or success, returns success with current status. If the run is in progress, stops all pending cells and marks the run as stopped. Stopped runs cannot be resumed

Parameters:

  • run_key* (string) - The run key (UUID) of the project run to stop
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"api":"riveter","path":"/v1/stop_run","query":{"run_key":"abc123"}}'

Use Cases

  1. E-commerce Scraping: Extract product data in consistent format
  2. Job Listings: Gather job postings with structured fields
  3. News Aggregation: Extract articles with title, date, content
  4. Price Monitoring: Track prices across competitor sites

Discover More

For full endpoint details and parameters:

curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/search \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"riveter API endpoints"}' List all endpoints
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/details \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"api":"riveter","path":"/v1/scrape"}'   # Get endpoint details
1---
2name: structured-scraping-riveter
3description: Web scraping with structured data extraction - define your output schema
4source: orthogonal
5---
6 
7 
8# Riveter - Structured Web Scraping
9 
10## Setup
11 
12Read your credentials from ~/.gooseworks/credentials.json:
13```bash
14export GOOSEWORKS_API_KEY=$(python3 -c "import json;print(json.load(open('$HOME/.gooseworks/credentials.json'))['api_key'])")
15export GOOSEWORKS_API_BASE=$(python3 -c "import json;print(json.load(open('$HOME/.gooseworks/credentials.json')).get('api_base','https://api.gooseworks.ai'))")
16```
17 
18If ~/.gooseworks/credentials.json does not exist, tell the user to run: `npx gooseworks login`
19 
20All endpoints use Bearer auth: `-H "Authorization: Bearer $GOOSEWORKS_API_KEY"`
21 
22 
23Scrape web pages and extract data into your defined structure.
24 
25## Capabilities
26 
27- **Scrape**: Scrape a webpage and return the text content
28- **Run**: Copy link Define the structure of your output directly in the API request
29- **Run data**: Retrieve the processed data from a completed project run (free)
30- **Run status**: Check the current status of a project run (free)
31- **Stop run**: Stop a currently running project (free)
32 
33## Usage
34 
35### Scrape
36Scrape a webpage and return the text content. This endpoint allows you to extract text content from any public webpage.
37 
38Parameters:
39- url* (string) - Example: "https://example.com"
40- proxy_country_code (string) - Optional two-character country code for proxy (e.g., 'us', 'gb', 'de')
41- skip_cache (boolean) - Default: false. Set to true to bypass cache and always fetch fresh content
42 
43```bash
44curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
45 -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
46 -H "Content-Type: application/json" \
47 -d '{"api":"riveter","path":"/v1/scrape","body":{"url":"https://example.com/article"}}'
48```
49 
50### Run
51Copy link Define the structure of your output directly in the API request. This endpoint allows you to define both your input data and output configuration in a single request.
52 
53Parameters:
54- input* (object) - The input object contains your source data: Keys are column/attribute names Values are arrays of strings (all arrays must be the same length) Maximum 1000 rows per request
55- output* (object) - The output object defines what data you want to extract: Keys are the names of attributes you want to extract Each attribute requires: prompt: Instructions for finding/extracting this data contexts: Array of input or other output attribute names this depends on. Optional Output Configuration Each output attribute can optionally include: format: Data type ('number', 'json', 'url', 'text', 'email', 'tag', 'date', 'boolean') format_details: Format-specific configuration (varies by format type). For json format, you can provide either a description (string) or a schema (JSON Schema object) or both. tools: Array of tools to use (['web_search', 'web_scrape', 'query_pdf', 'query_image']) max_tool_calls: Number of tool calls allowed (0-10) run_when: When to run this extraction ('always', 'any_filled', 'all_filled')
56- run_key (string) - Custom identifier for this run (optional, will be generated if not provided)
57 
58```bash
59curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
60 -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
61 -H "Content-Type: application/json" \
62 -d '{"api":"riveter","path":"/v1/run"}'
63 "input": {
64 "urls": ["https://example.com/products"]
65 },
66 "output": {
67 "name": {"prompt": "Product name", "contexts": ["urls"]},
68 "price": {"prompt": "Product price", "contexts": ["urls"], "format": "number"}
69 }
70}'
71```
72 
73### Run data (free)
74Retrieve the processed data from a completed project run
75 
76Parameters:
77- run_key* (string) - The run key (UUID) of the project run to retrieve data for
78 
79```bash
80curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
81 -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
82 -H "Content-Type: application/json" \
83 -d '{"api":"riveter","path":"/v1/run_data","query":{"run_key":"abc123"}}'
84```
85 
86### Run status (free)
87Check the current status of a project run
88 
89Parameters:
90- run_key* (string) - The run key (UUID) of the project run to check
91 
92```bash
93curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
94 -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
95 -H "Content-Type: application/json" \
96 -d '{"api":"riveter","path":"/v1/run_status","query":{"run_key":"abc123"}}'
97```
98 
99### Stop run (free)
100Stop a currently running project. This will halt all processing and mark the run as stopped. Behavior: If the run is already stopped or success, returns success with current status. If the run is in progress, stops all pending cells and marks the run as stopped. Stopped runs cannot be resumed
101 
102Parameters:
103- run_key* (string) - The run key (UUID) of the project run to stop
104 
105```bash
106curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
107 -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
108 -H "Content-Type: application/json" \
109 -d '{"api":"riveter","path":"/v1/stop_run","query":{"run_key":"abc123"}}'
110```
111 
112## Use Cases
113 
1141. **E-commerce Scraping**: Extract product data in consistent format
1152. **Job Listings**: Gather job postings with structured fields
1163. **News Aggregation**: Extract articles with title, date, content
1174. **Price Monitoring**: Track prices across competitor sites
118 
119## Discover More
120 
121For full endpoint details and parameters:
122 
123```bash
124curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/search \
125 -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
126 -H "Content-Type: application/json" \
127 -d '{"prompt":"riveter API endpoints"}' List all endpoints
128curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/details \
129 -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
130 -H "Content-Type: application/json" \
131 -d '{"api":"riveter","path":"/v1/scrape"}' # Get endpoint details
132```
133 

Discussion

Alternatives

Also in Scraping & extraction