MCP Server

scrapewright

io.github.Ozymandias-Owens-2/scrapewright
Data & Analytics Public & reachable MCP 2025-11-25

What this MCP does

Detects website platforms and extracts structured data from individual pages or crawled sites using reusable parsing recipes.

account
Credits left and this month's usage for the key in use.
Input schema
{'type': 'object', 'title': 'accountArguments', 'properties': {}}
Output schema
{'type': 'object', 'title': 'accountDictOutput', 'additionalProperties': True}
crawl_site
Walk a site from one listing URL and extract every item. Waits up to four minutes; a longer crawl returns a job_id to pass to crawl_status.
Input schema
{'type': 'object', 'title': 'crawl_siteArguments', 'required': ['listing_url'], 'properties': {'js': {'type': 'boolean', 'title': 'Js', 'default': False}, 'fields': {'anyOf': [{'type': 'array', 'items': {'type': 'string'}}, {'type': 'null'}], 'title': 'Fields', 'default': None}, 'scroll': {'type': 'integer', 'title': 'Scroll', 'default': 0}, 'max_items': {'type': 'integer', 'title': 'Max Items', 'default': 25}, 'listing_url': {'type': 'string', 'title': 'Listing Url'}}}
Output schema
{'type': 'object', 'title': 'crawl_siteDictOutput', 'additionalProperties': True}
crawl_status
Fetch a crawl that outlived its call.
Input schema
{'type': 'object', 'title': 'crawl_statusArguments', 'required': ['job_id'], 'properties': {'job_id': {'type': 'string', 'title': 'Job Id'}}}
Output schema
{'type': 'object', 'title': 'crawl_statusDictOutput', 'additionalProperties': True}
detect_site
Report what platform a site runs on and which strategy to use. Cheap; call it before a large job.
Input schema
{'type': 'object', 'title': 'detect_siteArguments', 'required': ['url'], 'properties': {'url': {'type': 'string', 'title': 'Url'}}}
Output schema
{'type': 'object', 'title': 'detect_siteDictOutput', 'additionalProperties': True}
extract_page
Extract structured data from ONE page. ``fields`` declares your own schema, e.g. ["title", "salary:number", "tags:list"]; omit it for the product schema. First call on a new site compiles a recipe (300 credits); later calls replay it for 1 credit per row.
Input schema
{'type': 'object', 'title': 'extract_pageArguments', 'required': ['url'], 'properties': {'js': {'type': 'boolean', 'title': 'Js', 'default': False}, 'url': {'type': 'string', 'title': 'Url'}, 'fields': {'anyOf': [{'type': 'array', 'items': {'type': 'string'}}, {'type': 'null'}], 'title': 'Fields', 'default': None}}}
Output schema
{'type': 'object', 'title': 'extract_pageDictOutput', 'additionalProperties': True}
Added
account
Sept. 21, 2026, 2:40 a.m.
Added
crawl_status
Sept. 21, 2026, 2:40 a.m.
Added
crawl_site
Sept. 21, 2026, 2:40 a.m.
Added
extract_page
Sept. 21, 2026, 2:40 a.m.
Added
detect_site
Sept. 21, 2026, 2:40 a.m.