Olostep MCP server agent

MCP server for Olostep — the web scraping, crawling, and search infrastructure used by top AI companies.

by olostep·MIT license·★ 24 Stars on the repo·GitHub ↗

Files of Olostep MCP server

olostep/main1 file
README.md
Show the full text627 lines

Olostep MCP Server

Docker Hub npm version License: ISC

A Model Context Protocol (MCP) server implementation that integrates with Olostep for web scraping, content extraction, and search capabilities. To set up Olostep MCP Server, you need to have an API key. You can get the API key by signing up on the Olostep website.

Features

  • Scrape website content in HTML, Markdown, JSON or Plain Text (with optional parsers)
  • Parser-based web search with structured results
  • AI Answers with citations and optional JSON-shaped outputs
  • Batch scraping of up to 10k URLs
  • Autonomous site crawling from a start URL
  • Website URL discovery and mapping (with include/exclude filters)
  • Country-specific request routing for geo-targeted content
  • Configurable wait times for JavaScript-heavy websites
  • Comprehensive error handling and reporting
  • Simple API key configuration

Installation

There are multiple ways to connect to the Olostep MCP Server. Choose the one that best fits your workflow.

The simplest way — no local installation required. Connect directly to our hosted MCP server:

https://mcp.olostep.com/mcp

Authentication is done via a Bearer token in the Authorization header using your Olostep API key. See the Client Setup section below for configuration examples.

🐳 Docker Hub

Pull and run the official Docker image:

docker pull olostep/mcp-server

docker run -i --rm \
  -e OLOSTEP_API_KEY="your-api-key" \
  olostep/mcp-server
🔧 Local Docker Build

If you prefer to build the image yourself from source:

git clone https://github.com/olostep/olostep-mcp-server.git
cd olostep-mcp-server
npm install
npm run build
docker build -t olostep/mcp-server:local .

docker run -i --rm -e OLOSTEP_API_KEY="your-api-key" olostep/mcp-server:local
📦 npx

Run without any installation using npx:

env OLOSTEP_API_KEY=your-api-key npx -y olostep-mcp

On Windows (PowerShell):

$env:OLOSTEP_API_KEY = "your-api-key"; npx -y olostep-mcp

On Windows (CMD):

set OLOSTEP_API_KEY=your-api-key && npx -y olostep-mcp

Or install globally:

npm install -g olostep-mcp

Client Setup

Cursor

The easiest way is to use the remote endpoint. Create or edit .cursor/mcp.json in your project root:

{
  "mcpServers": {
    "olostep": {
      "url": "https://mcp.olostep.com/mcp",
      "headers": {
        "Authorization": "Bearer YOUR_API_KEY_HERE"
      }
    }
  }
}

Alternative (local): Go to Cursor Settings > Features > MCP Servers, click "+ Add New MCP Server":

  • Name: olostep
  • Type: command
  • Command: env OLOSTEP_API_KEY=your-api-key npx -y olostep-mcp
Claude Desktop

Add this to your claude_desktop_config.json:

{
  "mcpServers": {
    "mcp-server-olostep": {
      "command": "npx",
      "args": ["-y", "olostep-mcp"],
      "env": {
        "OLOSTEP_API_KEY": "YOUR_API_KEY_HERE"
      }
    }
  }
}

Alternative (Docker):

{
  "mcpServers": {
    "olostep": {
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "OLOSTEP_API_KEY=YOUR_API_KEY_HERE",
        "olostep/mcp-server"
      ]
    }
  }
}

Or install via the Smithery CLI in your device terminal:

npx -y @smithery/cli install @olostep/olostep-mcp-server --client claude
Claude Code

Add the remote endpoint to your Claude Code MCP configuration:

{
  "mcpServers": {
    "olostep": {
      "url": "https://mcp.olostep.com/mcp",
      "headers": {
        "Authorization": "Bearer YOUR_API_KEY_HERE"
      }
    }
  }
}

Alternative (local):

{
  "mcpServers": {
    "olostep": {
      "command": "npx",
      "args": ["-y", "olostep-mcp"],
      "env": {
        "OLOSTEP_API_KEY": "YOUR_API_KEY_HERE"
      }
    }
  }
}
Windsurf

Add this to your ./codeium/windsurf/model_config.json:

{
  "mcpServers": {
    "olostep": {
      "serverUrl": "https://mcp.olostep.com/mcp",
      "headers": {
        "Authorization": "Bearer YOUR_API_KEY_HERE"
      }
    }
  }
}

Alternative (local):

{
  "mcpServers": {
    "mcp-server-olostep": {
      "command": "npx",
      "args": ["-y", "olostep-mcp"],
      "env": {
        "OLOSTEP_API_KEY": "YOUR_API_KEY_HERE"
      }
    }
  }
}
VS Code

Add this to your .vscode/mcp.json:

{
  "servers": {
    "olostep": {
      "type": "http",
      "url": "https://mcp.olostep.com/mcp",
      "headers": {
        "Authorization": "Bearer YOUR_API_KEY_HERE"
      }
    }
  }
}

Alternative (local):

{
  "servers": {
    "olostep": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "olostep-mcp"],
      "env": {
        "OLOSTEP_API_KEY": "YOUR_API_KEY_HERE"
      }
    }
  }
}
Metorial

Option 1: One-Click Installation (Recommended)

  1. Open Metorial dashboard
  2. Navigate to MCP Servers directory
  3. Search for "Olostep"
  4. Click "Install" and enter your API key

Option 2: Manual Configuration

Add this to your Metorial MCP server configuration:

{
  "olostep": {
    "command": "npx",
    "args": ["-y", "olostep-mcp"],
    "env": {
      "OLOSTEP_API_KEY": "YOUR_API_KEY_HERE"
    }
  }
}

The Olostep tools will then be available in your Metorial AI chats.

Configuration

Environment Variables
  • OLOSTEP_API_KEY: Your Olostep API key (required)
  • ORBIT_KEY: An optional key for using Orbit to route requests.

Available Tools

1. Scrape Website (scrape_website)

Extract content from a single URL. Supports multiple formats and JavaScript rendering.

{
  "name": "scrape_website",
  "arguments": {
    "url_to_scrape": "https://example.com",
    "output_format": "markdown",
    "country": "US",
    "wait_before_scraping": 1000,
    "parser": "@olostep/amazon-product"
  }
}
Parameters:
  • url_to_scrape: The URL of the website you want to scrape (required)
  • output_format: Choose format (html, markdown, json, or text) - default: markdown
  • country: Optional country code (e.g., US, GB, CA) for location-specific scraping
  • wait_before_scraping: Wait time in milliseconds before scraping (0-10000)
  • parser: Optional parser ID for specialized extraction
Response (example):
{
  "content": [
    {
      "type": "text",
      "text": "{\n  \"id\": \"scrp_...\",\n  \"url\": \"https://example.com\",\n  \"markdown_content\": \"# ...\",\n  \"html_content\": null,\n  \"json_content\": null,\n  \"text_content\": null,\n  \"status\": \"succeeded\",\n  \"timestamp\": \"2025-11-14T12:34:56Z\",\n  \"screenshot_hosted_url\": null,\n  \"page_metadata\": { }\n}"
    }
  ]
}
2. Search the Web (search_web)

Search the Web for a given query and get structured results (non-AI, parser-based).

{
  "name": "search_web",
  "arguments": {
    "query": "your search query",
    "country": "US"
  }
}
Parameters:
  • query: Search query (required)
  • country: Optional country code for localized results (default: US)
Response:
  • Structured JSON (as text) representing parser-based results
3. Answers (AI) (answers)

Search the web and return AI-powered answers in the JSON structure you want, with sources and citations.

{
  "name": "answers",
  "arguments": {
    "task": "Who are the top 5 competitors to Acme Inc. in the EU?",
    "json": "Return a list of the top 5 competitors with name and homepage URL"
  }
}
Parameters:
  • task: Question or task to answer using web data (required)
  • json: Optional JSON schema/object or a short description of the desired output shape
Response includes:
  • answer_id, object, task, result (JSON if provided), sources, created
4. Batch Scrape URLs (batch_scrape_urls)

Scrape up to 10k URLs at the same time. Perfect for large-scale data extraction.

{
  "name": "batch_scrape_urls",
  "arguments": {
    "urls_to_scrape": [
      {"url": "https://example.com/a", "custom_id": "a"},
      {"url": "https://example.com/b", "custom_id": "b"}
    ],
    "output_format": "markdown",
    "country": "US",
    "wait_before_scraping": 500,
    "parser": "@olostep/amazon-product"
  }
}
Response includes:
  • batch_id, status, total_urls, created_at, formats, country, parser, urls
5. Create Crawl (create_crawl)

Start an async crawl that autonomously discovers and scrapes entire websites by following links. Returns a crawl_id — the crawl runs in the background and does not return content in this response. You must then call get_crawl_results with the crawl_id to poll status and retrieve the scraped pages (same two-step pattern as batch_scrape_urls + get_batch_results).

{
  "name": "create_crawl",
  "arguments": {
    "start_url": "https://example.com/docs",
    "max_pages": 25,
    "output_format": "markdown",
    "country": "US",
    "parser": "@olostep/doc-parser"
  }
}
Response includes:
  • crawl_id, object, status, start_url, max_pages, created, formats, country, parser

Pair this call with get_crawl_results — do not pass a crawl_id to get_batch_results (crawls and batches are separate resources).

6. Create Map (create_map)

Get all URLs on a website. Extract all URLs for discovery and analysis.

{
  "name": "create_map",
  "arguments": {
    "website_url": "https://example.com",
    "search_query": "blog",
    "top_n": 200,
    "include_url_patterns": ["/blog/**"],
    "exclude_url_patterns": ["/admin/**"]
  }
}
Response includes:
  • map_id, object, url, total_urls, urls, search_query, top_n
7. Get Webpage Content (get_webpage_content)

Retrieves webpage content in clean markdown format with support for JavaScript rendering.

{
  "name": "get_webpage_content",
  "arguments": {
    "url_to_scrape": "https://example.com",
    "wait_before_scraping": 1000,
    "country": "US"
  }
}
Parameters:
  • url_to_scrape: The URL of the webpage to scrape (required)
  • wait_before_scraping: Time to wait in milliseconds before starting the scrape (default: 0)
  • country: Residential country to load the request from (e.g., US, CA, GB) (optional)
Response:
{
  "content": [
    {
      "type": "text",
      "text": "# Example Website\n\nThis is the markdown content of the webpage..."
    }
  ]
}
8. Get Website URLs (get_website_urls)

Search and retrieve relevant URLs from a website, sorted by relevance to your query.

{
  "name": "get_website_urls",
  "arguments": {
    "url": "https://example.com",
    "search_query": "your search term"
  }
}
Parameters:
  • url: The URL of the website to map (required)
  • search_query: The search query to sort URLs by (required)
Response:
{
  "content": [
    {
      "type": "text",
      "text": "Found 42 URLs matching your query:\n\nhttps://example.com/page1\nhttps://example.com/page2\n..."
    }
  ]
}
9. Get Batch Results (get_batch_results)

Retrieve the results of a previously submitted batch scrape job using its batch_id.

{
  "name": "get_batch_results",
  "arguments": {
    "batch_id": "batch_abc123"
  }
}
Parameters:
  • batch_id: The batch ID returned from batch_scrape_urls (required)
Response includes:
  • batch_id, status (processing or completed), total_urls, completed_urls, items (array of scraped results per URL with url, custom_id, markdown_content, html_content, json_content, text_content, status, page_metadata)
10. Get Crawl Results (get_crawl_results)

Retrieve the status and scraped pages for an async crawl started with create_crawl. This is the required companion to create_crawl — create_crawl only kicks off the job and returns a crawl_id; this tool is how you actually fetch the discovered pages and their content.

{
  "name": "get_crawl_results",
  "arguments": {
    "crawl_id": "crawl_abc123",
    "formats": ["markdown"],
    "items_limit": 20,
    "cursor": 0
  }
}
Parameters:
  • crawl_id: The crawl ID returned from create_crawl (required)
  • formats: Array of formats to retrieve per page — markdown, html, json, text (default: ["markdown"])
  • items_limit: Max pages to retrieve content for, 1–100 (default: 20)
  • cursor: Pagination cursor into the list of discovered pages (default: 0)
  • search_query: Optional filter to rank/select pages by relevance to a query
Response includes:
  • While in progress: crawl_id, status (in_progress), pages_completed, pages_total, and a message prompting you to call again in ~10 seconds.
  • When completed: crawl_id, status (completed), pages_returned, next_cursor, has_more, and a pages array where each entry has url, custom_id, and the requested content fields (markdown_content, html_content, json_content, text_content).

Error Handling

The server provides robust error handling:

  • Detailed error messages for API issues
  • Network error reporting
  • Authentication failure handling
  • Rate limit information

Example error response:

{
  "isError": true,
  "content": [
    {
      "type": "text",
      "text": "Olostep API Error: 401 Unauthorized. Details: {\"error\":\"Invalid API key\"}"
    }
  ]
}

Distribution

Docker Images

The MCP server is available as a Docker image:

  • Docker Hub: [olostep/mcp-server](https://hub.docker.com/r/olostep/mcp-server)
  • Official Docker MCP Registry: mcp/olostep (coming soon - enhanced security with signatures & SBOMs)
  • GitHub Container Registry: ghcr.io/olostep/olostep-mcp-server
Docker Desktop MCP Toolkit

The Olostep MCP Server is being added to Docker Desktop's official MCP Toolkit, which means users will be able to:

  • Discover it in Docker Desktop's MCP Toolkit UI
  • Install it with one click
  • Configure it visually
  • Use it with any MCP-compatible client (Claude Desktop, Cursor, etc.)

Status: Submission in progress to Docker MCP Registry

Supported Platforms
  • linux/amd64
  • linux/arm64
Building Locally
# Clone the repository
git clone https://github.com/olostep/olostep-mcp-server.git
cd olostep-mcp-server

# Build the image
npm install
npm run build
docker build -t olostep/mcp-server .

# Run locally
docker run -i --rm -e OLOSTEP_API_KEY="your-key" olostep/mcp-server

License

ISC License

1# Olostep MCP Server
2 
3[Docker Hub](https://hub.docker.com/r/olostep/mcp-server)
4[npm version](https://www.npmjs.com/package/olostep-mcp)
5[License: ISC](https://opensource.org/licenses/ISC)
6 
7A Model Context Protocol (MCP) server implementation that integrates with [Olostep](https://olostep.com) for web scraping, content extraction, and search capabilities.
8To set up Olostep MCP Server, you need to have an API key. You can get the API key by signing up on the [Olostep website](https://olostep.com/auth).
9 
10## Features
11 
12- Scrape website content in HTML, Markdown, JSON or Plain Text (with optional parsers)
13- Parser-based web search with structured results
14- AI Answers with citations and optional JSON-shaped outputs
15- Batch scraping of up to 10k URLs
16- Autonomous site crawling from a start URL
17- Website URL discovery and mapping (with include/exclude filters)
18- Country-specific request routing for geo-targeted content
19- Configurable wait times for JavaScript-heavy websites
20- Comprehensive error handling and reporting
21- Simple API key configuration
22 
23## Installation
24 
25There are multiple ways to connect to the Olostep MCP Server. Choose the one that best fits your workflow.
26 
27### ☁️ Remote Endpoint (Recommended)
28 
29The simplest way — no local installation required. Connect directly to our hosted MCP server:
30 
31```
32https://mcp.olostep.com/mcp
33```
34 
35Authentication is done via a `Bearer` token in the `Authorization` header using your Olostep API key. See the [Client Setup](#client-setup) section below for configuration examples.
36 
37### 🐳 Docker Hub
38 
39Pull and run the official Docker image:
40 
41```bash
42docker pull olostep/mcp-server
43 
44docker run -i --rm \
45 -e OLOSTEP_API_KEY="your-api-key" \
46 olostep/mcp-server
47```
48 
49### 🔧 Local Docker Build
50 
51If you prefer to build the image yourself from source:
52 
53```bash
54git clone https://github.com/olostep/olostep-mcp-server.git
55cd olostep-mcp-server
56npm install
57npm run build
58docker build -t olostep/mcp-server:local .
59 
60docker run -i --rm -e OLOSTEP_API_KEY="your-api-key" olostep/mcp-server:local
61```
62 
63### 📦 npx
64 
65Run without any installation using npx:
66 
67```bash
68env OLOSTEP_API_KEY=your-api-key npx -y olostep-mcp
69```
70 
71On Windows (PowerShell):
72 
73```powershell
74$env:OLOSTEP_API_KEY = "your-api-key"; npx -y olostep-mcp
75```
76 
77On Windows (CMD):
78 
79```cmd
80set OLOSTEP_API_KEY=your-api-key && npx -y olostep-mcp
81```
82 
83Or install globally:
84 
85```bash
86npm install -g olostep-mcp
87```
88 
89## Client Setup
90 
91### Cursor
92 
93The easiest way is to use the remote endpoint. Create or edit `.cursor/mcp.json` in your project root:
94 
95```json
96{
97 "mcpServers": {
98 "olostep": {
99 "url": "https://mcp.olostep.com/mcp",
100 "headers": {
101 "Authorization": "Bearer YOUR_API_KEY_HERE"
102 }
103 }
104 }
105}
106```
107 
108**Alternative (local):** Go to Cursor Settings > Features > MCP Servers, click "+ Add New MCP Server":
109 
110- **Name:** `olostep`
111- **Type:** `command`
112- **Command:** `env OLOSTEP_API_KEY=your-api-key npx -y olostep-mcp`
113 
114### Claude Desktop
115 
116Add this to your `claude_desktop_config.json`:
117 
118```json
119{
120 "mcpServers": {
121 "mcp-server-olostep": {
122 "command": "npx",
123 "args": ["-y", "olostep-mcp"],
124 "env": {
125 "OLOSTEP_API_KEY": "YOUR_API_KEY_HERE"
126 }
127 }
128 }
129}
130```
131 
132**Alternative (Docker):**
133 
134```json
135{
136 "mcpServers": {
137 "olostep": {
138 "command": "docker",
139 "args": [
140 "run", "-i", "--rm",
141 "-e", "OLOSTEP_API_KEY=YOUR_API_KEY_HERE",
142 "olostep/mcp-server"
143 ]
144 }
145 }
146}
147```
148 
149Or install via the Smithery CLI in your device terminal:
150 
151```bash
152npx -y @smithery/cli install @olostep/olostep-mcp-server --client claude
153```
154 
155### Claude Code
156 
157Add the remote endpoint to your Claude Code MCP configuration:
158 
159```json
160{
161 "mcpServers": {
162 "olostep": {
163 "url": "https://mcp.olostep.com/mcp",
164 "headers": {
165 "Authorization": "Bearer YOUR_API_KEY_HERE"
166 }
167 }
168 }
169}
170```
171 
172**Alternative (local):**
173 
174```json
175{
176 "mcpServers": {
177 "olostep": {
178 "command": "npx",
179 "args": ["-y", "olostep-mcp"],
180 "env": {
181 "OLOSTEP_API_KEY": "YOUR_API_KEY_HERE"
182 }
183 }
184 }
185}
186```
187 
188### Windsurf
189 
190Add this to your `./codeium/windsurf/model_config.json`:
191 
192```json
193{
194 "mcpServers": {
195 "olostep": {
196 "serverUrl": "https://mcp.olostep.com/mcp",
197 "headers": {
198 "Authorization": "Bearer YOUR_API_KEY_HERE"
199 }
200 }
201 }
202}
203```
204 
205**Alternative (local):**
206 
207```json
208{
209 "mcpServers": {
210 "mcp-server-olostep": {
211 "command": "npx",
212 "args": ["-y", "olostep-mcp"],
213 "env": {
214 "OLOSTEP_API_KEY": "YOUR_API_KEY_HERE"
215 }
216 }
217 }
218}
219```
220 
221### VS Code
222 
223Add this to your `.vscode/mcp.json`:
224 
225```json
226{
227 "servers": {
228 "olostep": {
229 "type": "http",
230 "url": "https://mcp.olostep.com/mcp",
231 "headers": {
232 "Authorization": "Bearer YOUR_API_KEY_HERE"
233 }
234 }
235 }
236}
237```
238 
239**Alternative (local):**
240 
241```json
242{
243 "servers": {
244 "olostep": {
245 "type": "stdio",
246 "command": "npx",
247 "args": ["-y", "olostep-mcp"],
248 "env": {
249 "OLOSTEP_API_KEY": "YOUR_API_KEY_HERE"
250 }
251 }
252 }
253}
254```
255 
256### Metorial
257 
258**Option 1: One-Click Installation (Recommended)**
259 
2601. Open [Metorial](https://metorial.com) dashboard
2612. Navigate to MCP Servers directory
2623. Search for "Olostep"
2634. Click "Install" and enter your API key
264 
265**Option 2: Manual Configuration**
266 
267Add this to your Metorial MCP server configuration:
268 
269```json
270{
271 "olostep": {
272 "command": "npx",
273 "args": ["-y", "olostep-mcp"],
274 "env": {
275 "OLOSTEP_API_KEY": "YOUR_API_KEY_HERE"
276 }
277 }
278}
279```
280 
281The Olostep tools will then be available in your Metorial AI chats.
282 
283## Configuration
284 
285### Environment Variables
286 
287- `OLOSTEP_API_KEY`: Your Olostep API key (required)
288- `ORBIT_KEY`: An optional key for using Orbit to route requests.
289 
290## Available Tools
291 
292### 1. Scrape Website (`scrape_website`)
293 
294Extract content from a single URL. Supports multiple formats and JavaScript rendering.
295 
296```json
297{
298 "name": "scrape_website",
299 "arguments": {
300 "url_to_scrape": "https://example.com",
301 "output_format": "markdown",
302 "country": "US",
303 "wait_before_scraping": 1000,
304 "parser": "@olostep/amazon-product"
305 }
306}
307```
308 
309#### Parameters:
310 
311- `url_to_scrape`: The URL of the website you want to scrape (required)
312- `output_format`: Choose format (`html`, `markdown`, `json`, or `text`) - default: `markdown`
313- `country`: Optional country code (e.g., US, GB, CA) for location-specific scraping
314- `wait_before_scraping`: Wait time in milliseconds before scraping (0-10000)
315- `parser`: Optional parser ID for specialized extraction
316 
317#### Response (example):
318 
319```json
320{
321 "content": [
322 {
323 "type": "text",
324 "text": "{\n \"id\": \"scrp_...\",\n \"url\": \"https://example.com\",\n \"markdown_content\": \"# ...\",\n \"html_content\": null,\n \"json_content\": null,\n \"text_content\": null,\n \"status\": \"succeeded\",\n \"timestamp\": \"2025-11-14T12:34:56Z\",\n \"screenshot_hosted_url\": null,\n \"page_metadata\": { }\n}"
325 }
326 ]
327}
328```
329 
330### 2. Search the Web (`search_web`)
331 
332Search the Web for a given query and get structured results (non-AI, parser-based).
333 
334```json
335{
336 "name": "search_web",
337 "arguments": {
338 "query": "your search query",
339 "country": "US"
340 }
341}
342```
343 
344#### Parameters:
345 
346- `query`: Search query (required)
347- `country`: Optional country code for localized results (default: `US`)
348 
349#### Response:
350 
351- Structured JSON (as text) representing parser-based results
352 
353### 3. Answers (AI) (`answers`)
354 
355Search the web and return AI-powered answers in the JSON structure you want, with sources and citations.
356 
357```json
358{
359 "name": "answers",
360 "arguments": {
361 "task": "Who are the top 5 competitors to Acme Inc. in the EU?",
362 "json": "Return a list of the top 5 competitors with name and homepage URL"
363 }
364}
365```
366 
367#### Parameters:
368 
369- `task`: Question or task to answer using web data (required)
370- `json`: Optional JSON schema/object or a short description of the desired output shape
371 
372#### Response includes:
373 
374- `answer_id`, `object`, `task`, `result` (JSON if provided), `sources`, `created`
375 
376### 4. Batch Scrape URLs (`batch_scrape_urls`)
377 
378Scrape up to 10k URLs at the same time. Perfect for large-scale data extraction.
379 
380```json
381{
382 "name": "batch_scrape_urls",
383 "arguments": {
384 "urls_to_scrape": [
385 {"url": "https://example.com/a", "custom_id": "a"},
386 {"url": "https://example.com/b", "custom_id": "b"}
387 ],
388 "output_format": "markdown",
389 "country": "US",
390 "wait_before_scraping": 500,
391 "parser": "@olostep/amazon-product"
392 }
393}
394```
395 
396#### Response includes:
397 
398- `batch_id`, `status`, `total_urls`, `created_at`, `formats`, `country`, `parser`, `urls`
399 
400### 5. Create Crawl (`create_crawl`)
401 
402Start an **async** crawl that autonomously discovers and scrapes entire websites by following links. Returns a `crawl_id` — the crawl runs in the background and does **not** return content in this response. You must then call `get_crawl_results` with the `crawl_id` to poll status and retrieve the scraped pages (same two-step pattern as `batch_scrape_urls` + `get_batch_results`).
403 
404```json
405{
406 "name": "create_crawl",
407 "arguments": {
408 "start_url": "https://example.com/docs",
409 "max_pages": 25,
410 "output_format": "markdown",
411 "country": "US",
412 "parser": "@olostep/doc-parser"
413 }
414}
415```
416 
417#### Response includes:
418 
419- `crawl_id`, `object`, `status`, `start_url`, `max_pages`, `created`, `formats`, `country`, `parser`
420 
421> Pair this call with `get_crawl_results` — do **not** pass a `crawl_id` to `get_batch_results` (crawls and batches are separate resources).
422 
423### 6. Create Map (`create_map`)
424 
425Get all URLs on a website. Extract all URLs for discovery and analysis.
426 
427```json
428{
429 "name": "create_map",
430 "arguments": {
431 "website_url": "https://example.com",
432 "search_query": "blog",
433 "top_n": 200,
434 "include_url_patterns": ["/blog/**"],
435 "exclude_url_patterns": ["/admin/**"]
436 }
437}
438```
439 
440#### Response includes:
441 
442- `map_id`, `object`, `url`, `total_urls`, `urls`, `search_query`, `top_n`
443 
444### 7. Get Webpage Content (`get_webpage_content`)
445 
446Retrieves webpage content in clean markdown format with support for JavaScript rendering.
447 
448```json
449{
450 "name": "get_webpage_content",
451 "arguments": {
452 "url_to_scrape": "https://example.com",
453 "wait_before_scraping": 1000,
454 "country": "US"
455 }
456}
457```
458 
459#### Parameters:
460 
461- `url_to_scrape`: The URL of the webpage to scrape (required)
462- `wait_before_scraping`: Time to wait in milliseconds before starting the scrape (default: 0)
463- `country`: Residential country to load the request from (e.g., US, CA, GB) (optional)
464 
465#### Response:
466 
467```json
468{
469 "content": [
470 {
471 "type": "text",
472 "text": "# Example Website\n\nThis is the markdown content of the webpage..."
473 }
474 ]
475}
476```
477 
478### 8. Get Website URLs (`get_website_urls`)
479 
480Search and retrieve relevant URLs from a website, sorted by relevance to your query.
481 
482```json
483{
484 "name": "get_website_urls",
485 "arguments": {
486 "url": "https://example.com",
487 "search_query": "your search term"
488 }
489}
490```
491 
492#### Parameters:
493 
494- `url`: The URL of the website to map (required)
495- `search_query`: The search query to sort URLs by (required)
496 
497#### Response:
498 
499```json
500{
501 "content": [
502 {
503 "type": "text",
504 "text": "Found 42 URLs matching your query:\n\nhttps://example.com/page1\nhttps://example.com/page2\n..."
505 }
506 ]
507}
508```
509 
510### 9. Get Batch Results (`get_batch_results`)
511 
512Retrieve the results of a previously submitted batch scrape job using its `batch_id`.
513 
514```json
515{
516 "name": "get_batch_results",
517 "arguments": {
518 "batch_id": "batch_abc123"
519 }
520}
521```
522 
523#### Parameters:
524 
525- `batch_id`: The batch ID returned from `batch_scrape_urls` (required)
526 
527#### Response includes:
528 
529- `batch_id`, `status` (`processing` or `completed`), `total_urls`, `completed_urls`, `items` (array of scraped results per URL with `url`, `custom_id`, `markdown_content`, `html_content`, `json_content`, `text_content`, `status`, `page_metadata`)
530 
531### 10. Get Crawl Results (`get_crawl_results`)
532 
533Retrieve the status and scraped pages for an async crawl started with `create_crawl`. This is the required companion to `create_crawl` — `create_crawl` only kicks off the job and returns a `crawl_id`; this tool is how you actually fetch the discovered pages and their content.
534 
535```json
536{
537 "name": "get_crawl_results",
538 "arguments": {
539 "crawl_id": "crawl_abc123",
540 "formats": ["markdown"],
541 "items_limit": 20,
542 "cursor": 0
543 }
544}
545```
546 
547#### Parameters:
548 
549- `crawl_id`: The crawl ID returned from `create_crawl` (required)
550- `formats`: Array of formats to retrieve per page — `markdown`, `html`, `json`, `text` (default: `["markdown"]`)
551- `items_limit`: Max pages to retrieve content for, 1–100 (default: 20)
552- `cursor`: Pagination cursor into the list of discovered pages (default: 0)
553- `search_query`: Optional filter to rank/select pages by relevance to a query
554 
555#### Response includes:
556 
557- **While in progress:** `crawl_id`, `status` (`in_progress`), `pages_completed`, `pages_total`, and a `message` prompting you to call again in ~10 seconds.
558- **When completed:** `crawl_id`, `status` (`completed`), `pages_returned`, `next_cursor`, `has_more`, and a `pages` array where each entry has `url`, `custom_id`, and the requested content fields (`markdown_content`, `html_content`, `json_content`, `text_content`).
559 
560## Error Handling
561 
562The server provides robust error handling:
563 
564- Detailed error messages for API issues
565- Network error reporting
566- Authentication failure handling
567- Rate limit information
568 
569Example error response:
570 
571```json
572{
573 "isError": true,
574 "content": [
575 {
576 "type": "text",
577 "text": "Olostep API Error: 401 Unauthorized. Details: {\"error\":\"Invalid API key\"}"
578 }
579 ]
580}
581```
582 
583## Distribution
584 
585### Docker Images
586 
587The MCP server is available as a Docker image:
588 
589- **Docker Hub:** `[olostep/mcp-server](https://hub.docker.com/r/olostep/mcp-server)`
590- **Official Docker MCP Registry:** `mcp/olostep` (coming soon - enhanced security with signatures & SBOMs)
591- **GitHub Container Registry:** `ghcr.io/olostep/olostep-mcp-server`
592 
593### Docker Desktop MCP Toolkit
594 
595The Olostep MCP Server is being added to Docker Desktop's official MCP Toolkit, which means users will be able to:
596 
597- Discover it in Docker Desktop's MCP Toolkit UI
598- Install it with one click
599- Configure it visually
600- Use it with any MCP-compatible client (Claude Desktop, Cursor, etc.)
601 
602**Status**: Submission in progress to [Docker MCP Registry](https://github.com/docker/mcp-registry)
603 
604### Supported Platforms
605 
606- `linux/amd64`
607- `linux/arm64`
608 
609### Building Locally
610 
611```bash
612# Clone the repository
613git clone https://github.com/olostep/olostep-mcp-server.git
614cd olostep-mcp-server
615 
616# Build the image
617npm install
618npm run build
619docker build -t olostep/mcp-server .
620 
621# Run locally
622docker run -i --rm -e OLOSTEP_API_KEY="your-key" olostep/mcp-server
623```
624 
625## License
626 
627ISC License

Discussion

Alternatives