Servidor MCP

oxylabs-mcp

io.oxylabs/oxylabs-mcp

Qué hace este MCP

Searches, scrapes, crawls, maps, and browser-navigates websites and sources, including specialized Google and Amazon data extraction.

ai_browser_agent
Run the browser agent and return the data in the specified format. This tool is useful if you need navigate around the website and do some actions. It allows navigating to any url, clicking on links, filling forms, scrolling, etc. Finally it returns the data in the specified format. Schema is required only if output_format is json, csv or toon. 'task_prompt' describes what browser agent should achieve
Solo lectura
Esquema de entrada
{'type': 'object', 'required': ['url', 'task_prompt'], 'properties': {'url': {'type': 'string', 'description': 'The URL to start the browser agent navigation from.'}, 'schema': {'anyOf': [{'type': 'object', 'additionalProperties': True}, {'type': 'null'}], 'default': None, 'description': 'The schema to use for the scrape. Only required if output_format is json, csv or toon.'}, 'task_prompt': {'type': 'string', 'description': 'What browser agent should do.'}, 'geo_location': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'description': 'Two letter ISO country code to use for the browser proxy.'}, 'output_format': {'enum': ['json', 'markdown', 'html', 'csv', 'toon'], 'type': 'string', 'default': 'markdown', 'description': 'The output format. Markdown returns full text of the page including links. Toon(Token-Oriented Object Notation) returns data in Toon format, which is optimized for AI agents. If json, csv or toon, the schema is required.'}}, 'additionalProperties': False}
Esquema de salida
{'type': 'object', 'required': ['result'], 'properties': {'result': {'type': 'string'}}, 'x-fastmcp-wrap-result': True}
ai_crawler
Tool useful for crawling a website from starting url and returning data in a specified format. Schema is required only if output_format is json, csv or toon. 'render_javascript' is used to render javascript heavy websites. 'return_sources_limit' is used to limit the number of sources to return, for example if you expect results from single source, you can set it to 1.
Solo lectura
Esquema de entrada
{'type': 'object', 'required': ['url', 'user_prompt'], 'properties': {'url': {'type': 'string', 'description': 'The URL from which crawling will be started.'}, 'schema': {'anyOf': [{'type': 'object', 'additionalProperties': True}, {'type': 'null'}], 'default': None, 'description': 'The JSON schema to use for structured data extraction from the crawled pages. Only required if output_format is json, csv or toon.'}, 'user_prompt': {'type': 'string', 'description': 'What information user wants to extract from the domain.'}, 'geo_location': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'description': 'Two letter ISO country code to use for the crawl proxy.'}, 'output_format': {'enum': ['json', 'markdown', 'csv', 'toon'], 'type': 'string', 'default': 'markdown', 'description': 'The format of the output. If json, csv or toon, the schema is required. Markdown returns full text of the page. CSV returns data in CSV format. Toon(Token-Oriented Object Notation) returns data in Toon format, which is optimized for AI agents.'}, 'render_javascript': {'type': 'boolean', 'default': False, 'description': 'Whether to render the HTML of the page using javascript. Much slower, therefore use it only for websites that require javascript to render the page. Unless user asks to use it, first try to crawl the page without it. If results are unsatisfactory, try to use it.'}, 'return_sources_limit': {'type': 'integer', 'default': 25, 'maximum': 50, 'description': 'The maximum number of sources to return.'}}, 'additionalProperties': False}
Esquema de salida
{'type': 'object', 'required': ['result'], 'properties': {'result': {'type': 'string'}}, 'x-fastmcp-wrap-result': True}
ai_map
Tool useful for mapping website's URLs.
Solo lectura
Esquema de entrada
{'type': 'object', 'required': ['url'], 'properties': {'url': {'type': 'string', 'description': 'The URL from which URLs mapping will be started.'}, 'limit': {'type': 'integer', 'default': 25, 'maximum': 50, 'description': 'The maximum number of URLs to return.'}, 'user_prompt': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'description': "What kind of URLs user wants to find. Can be used together with 'search_keywords'."}, 'geo_location': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'description': 'Two letter ISO country code to use for the mapping proxy.'}, 'max_crawl_depth': {'type': 'integer', 'default': 1, 'maximum': 5, 'description': 'The maximum depth of the crawl.'}, 'search_keywords': {'anyOf': [{'type': 'array', 'items': {'type': 'string'}}, {'type': 'null'}], 'default': None, 'description': 'The keywords to use for URLs paths filtering. Keywords are matched as OR condition. Meaning, one keyword is enough to match the url path.'}, 'allow_subdomains': {'type': 'boolean', 'default': False, 'description': 'Whether to map subdomains URLs as well.'}, 'render_javascript': {'type': 'boolean', 'default': False, 'description': 'Whether to render the HTML of the page using javascript. Much slower, therefore use it only for websites that require javascript to render the page. Unless user asks to use it, first try to crawl the page without it. If results are unsatisfactory, try to use it.'}, 'allow_external_domains': {'type': 'boolean', 'default': False, 'description': 'Whether to include external domains URLs.'}}, 'additionalProperties': False}
Esquema de salida
{'type': 'object', 'required': ['result'], 'properties': {'result': {'type': 'string'}}, 'x-fastmcp-wrap-result': True}
ai_scraper
Scrape the contents of the web page and return the data in the specified format. Schema is required only if output_format is json or csv. 'render_javascript' is used to render javascript heavy websites.
Solo lectura
Esquema de entrada
{'type': 'object', 'required': ['url'], 'properties': {'url': {'type': 'string', 'description': 'The URL to scrape'}, 'schema': {'anyOf': [{'type': 'object', 'additionalProperties': True}, {'type': 'null'}], 'default': None, 'description': 'The JSON schema to use for structured data extraction from the scraped page. Only required if output_format is json, csv or toon.'}, 'geo_location': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'description': 'Two letter ISO country code to use for the scrape proxy.'}, 'output_format': {'enum': ['json', 'markdown', 'csv', 'toon'], 'type': 'string', 'default': 'markdown', 'description': 'The format of the output. If json, csv or toon, the schema is required. Markdown returns full text of the page. CSV returns data in CSV format, tabular like data. Toon(Token-Oriented Object Notation) returns data in Toon format, which is optimized for AI agents.'}, 'render_javascript': {'type': 'boolean', 'default': False, 'description': 'Whether to render the HTML of the page using javascript. Much slower, therefore use it only for websites that require javascript to render the page.Unless user asks to use it, first try to scrape the page without it. If results are unsatisfactory, try to use it.'}}, 'additionalProperties': False}
Esquema de salida
{'type': 'object', 'required': ['result'], 'properties': {'result': {'type': 'string'}}, 'x-fastmcp-wrap-result': True}
ai_search
Search the web based on a provided query. 'return_content' is used to return markdown content for each search result. If 'return_content' is set to True, you don't need to use ai_scraper to get the content of the search results urls, because it is already included in the search results. if 'return_content' is set to True, prefer lower 'limit' to reduce payload size.
Solo lectura
Esquema de entrada
{'type': 'object', 'required': ['query'], 'properties': {'limit': {'type': 'integer', 'default': 10, 'maximum': 50, 'description': 'Maximum number of results to return.'}, 'query': {'type': 'string', 'description': 'The query to search for.'}, 'geo_location': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'description': 'Two letter ISO country code to use for the search proxy.'}, 'return_content': {'type': 'boolean', 'default': False, 'description': 'Whether to return markdown content of the search results.'}, 'render_javascript': {'type': 'boolean', 'default': False, 'description': 'Whether to render the HTML of the page using javascript. Much slower, therefore use it only if user asks to use it.First try to search with setting it to False. '}}, 'additionalProperties': False}
Esquema de salida
{'type': 'object', 'required': ['result'], 'properties': {'result': {'type': 'string'}}, 'x-fastmcp-wrap-result': True}
amazon_product_scraper
Scrape Amazon products. Supports content parsing, different user agent types, domain, geolocation, locale parameters and different output formats. Supports Amazon specific parameters such as currency and getting more accurate pricing data with auto select variant.
Solo lectura
Esquema de entrada
{'type': 'object', 'required': ['query'], 'properties': {'parse': {'type': 'boolean', 'default': True, 'description': 'Should result be parsed. If the result is not parsed, the output_format parameter is applied.'}, 'query': {'type': 'string', 'description': 'Keyword to search for.'}, 'domain': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'examples': ['uk', 'us', 'fr'], 'description': "\n        Domain localization for Google.\n        Use country top level domains.\n        For example:\n            - 'co.uk' for United Kingdom\n            - 'us' for United States\n            - 'fr' for France\n        "}, 'locale': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'examples': ['en-US', 'de-AT', 'fr-FR'], 'description': "\n        Set 'Accept-Language' header value which changes your Google search page web interface language.\n        Examples:\n            - 'en-US' for English, United States\n            - 'de-AT' for German, Austria\n            - 'fr-FR' for French, France\n        "}, 'render': {'anyOf': [{'type': 'string', 'const': 'html'}, {'type': 'null'}], 'default': None, 'examples': ['html'], 'description': "\n        Whether a headless browser should be used to render the page.\n        For example:\n            - 'html' when browser is required to render the page.\n        "}, 'currency': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'examples': ['USD', 'EUR', 'AUD'], 'description': 'Currency that will be used to display the prices.'}, 'geo_location': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'examples': ['US', 'DE', 'FR'], 'description': "\n        The geographical location that the result should be adapted for.\n        Use ISO-3166 country codes.\n        Examples:\n            - 'California, United States'\n            - 'Mexico'\n            - 'US' for United States\n            - 'DE' for Germany\n            - 'FR' for France\n        "}, 'output_format': {'anyOf': [{'enum': ['links', 'md', 'html'], 'type': 'string'}, {'type': 'null'}], 'default': None, 'description': "\n        The format of the output. Works only when parse parameter is false.\n            - links - Most efficient when the goal is navigation or finding specific URLs. Use this first when you need to locate a specific page within a website.\n            - md - Best for extracting and reading visible content once you've found the right page. Use this to get structured content that's easy to read and process.\n            - html - Should be used sparingly only when you need the raw HTML structure, JavaScript code, or styling information.\n        "}, 'user_agent_type': {'anyOf': [{'enum': ['desktop', 'desktop_chrome', 'desktop_firefox', 'desktop_safari', 'desktop_edge', 'desktop_opera', 'mobile', 'mobile_ios', 'mobile_android', 'tablet'], 'type': 'string'}, {'type': 'null'}], 'default': None, 'description': 'Device type and browser that will be used to determine User-Agent header value.'}, 'autoselect_variant': {'type': 'boolean', 'default': False, 'description': 'To get accurate pricing/buybox data, set this parameter to true.'}}, 'additionalProperties': False}
Esquema de salida
{'type': 'object', 'required': ['result'], 'properties': {'result': {'type': 'string'}}, 'x-fastmcp-wrap-result': True}
amazon_search_scraper
Scrape Amazon search results. Supports content parsing, different user agent types, pagination, domain, geolocation, locale parameters and different output formats. Supports Amazon specific parameters such as category id, merchant id, currency.
Solo lectura
Esquema de entrada
{'type': 'object', 'required': ['query'], 'properties': {'pages': {'type': 'integer', 'default': 0, 'description': 'Number of pages to retrieve.'}, 'parse': {'type': 'boolean', 'default': True, 'description': 'Should result be parsed. If the result is not parsed, the output_format parameter is applied.'}, 'query': {'type': 'string', 'description': 'Keyword to search for.'}, 'domain': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'examples': ['uk', 'us', 'fr'], 'description': "\n        Domain localization for Google.\n        Use country top level domains.\n        For example:\n            - 'co.uk' for United Kingdom\n            - 'us' for United States\n            - 'fr' for France\n        "}, 'locale': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'examples': ['en-US', 'de-AT', 'fr-FR'], 'description': "\n        Set 'Accept-Language' header value which changes your Google search page web interface language.\n        Examples:\n            - 'en-US' for English, United States\n            - 'de-AT' for German, Austria\n            - 'fr-FR' for French, France\n        "}, 'render': {'anyOf': [{'type': 'string', 'const': 'html'}, {'type': 'null'}], 'default': None, 'examples': ['html'], 'description': "\n        Whether a headless browser should be used to render the page.\n        For example:\n            - 'html' when browser is required to render the page.\n        "}, 'currency': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'examples': ['USD', 'EUR', 'AUD'], 'description': 'Currency that will be used to display the prices.'}, 'start_page': {'type': 'integer', 'default': 0, 'description': 'Starting page number.'}, 'category_id': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'description': 'Search for items in a particular browse node (product category).'}, 'merchant_id': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'description': 'Search for items sold by a particular seller.'}, 'geo_location': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'examples': ['US', 'DE', 'FR'], 'description': "\n        The geographical location that the result should be adapted for.\n        Use ISO-3166 country codes.\n        Examples:\n            - 'California, United States'\n            - 'Mexico'\n            - 'US' for United States\n            - 'DE' for Germany\n            - 'FR' for France\n        "}, 'output_format': {'anyOf': [{'enum': ['links', 'md', 'html'], 'type': 'string'}, {'type': 'null'}], 'default': None, 'description': "\n        The format of the output. Works only when parse parameter is false.\n            - links - Most efficient when the goal is navigation or finding specific URLs. Use this first when you need to locate a specific page within a website.\n            - md - Best for extracting and reading visible content once you've found the right page. Use this to get structured content that's easy to read and process.\n            - html - Should be used sparingly only when you need the raw HTML structure, JavaScript code, or styling information.\n        "}, 'user_agent_type': {'anyOf': [{'enum': ['desktop', 'desktop_chrome', 'desktop_firefox', 'desktop_safari', 'desktop_edge', 'desktop_opera', 'mobile', 'mobile_ios', 'mobile_android', 'tablet'], 'type': 'string'}, {'type': 'null'}], 'default': None, 'description': 'Device type and browser that will be used to determine User-Agent header value.'}}, 'additionalProperties': False}
Esquema de salida
{'type': 'object', 'required': ['result'], 'properties': {'result': {'type': 'string'}}, 'x-fastmcp-wrap-result': True}
generate_schema
Generate a json schema in openapi format.
Solo lectura
Esquema de entrada
{'type': 'object', 'required': ['user_prompt', 'app_name'], 'properties': {'app_name': {'enum': ['ai_crawler', 'ai_scraper', 'browser_agent'], 'type': 'string'}, 'user_prompt': {'type': 'string'}}, 'additionalProperties': False}
Esquema de salida
{'type': 'object', 'required': ['result'], 'properties': {'result': {'type': 'string'}}, 'x-fastmcp-wrap-result': True}
google_search_scraper
Scrape Google Search results. Supports content parsing, different user agent types, pagination, domain, geolocation, locale parameters and different output formats.
Solo lectura
Esquema de entrada
{'type': 'object', 'required': ['query'], 'properties': {'limit': {'type': 'integer', 'default': 0, 'description': 'Number of results to retrieve in each page.'}, 'pages': {'type': 'integer', 'default': 0, 'description': 'Number of pages to retrieve.'}, 'parse': {'type': 'boolean', 'default': True, 'description': 'Should result be parsed. If the result is not parsed, the output_format parameter is applied.'}, 'query': {'type': 'string', 'description': 'URL-encoded keyword to search for.'}, 'domain': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'examples': ['uk', 'us', 'fr'], 'description': "\n        Domain localization for Google.\n        Use country top level domains.\n        For example:\n            - 'co.uk' for United Kingdom\n            - 'us' for United States\n            - 'fr' for France\n        "}, 'locale': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'examples': ['en-US', 'de-AT', 'fr-FR'], 'description': "\n        Set 'Accept-Language' header value which changes your Google search page web interface language.\n        Examples:\n            - 'en-US' for English, United States\n            - 'de-AT' for German, Austria\n            - 'fr-FR' for French, France\n        "}, 'render': {'anyOf': [{'type': 'string', 'const': 'html'}, {'type': 'null'}], 'default': None, 'examples': ['html'], 'description': "\n        Whether a headless browser should be used to render the page.\n        For example:\n            - 'html' when browser is required to render the page.\n        "}, 'ad_mode': {'type': 'boolean', 'default': False, 'description': 'If true will use the Google Ads source optimized for the paid ads.'}, 'start_page': {'type': 'integer', 'default': 0, 'description': 'Starting page number.'}, 'geo_location': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'examples': ['US', 'DE', 'FR'], 'description': "\n        The geographical location that the result should be adapted for.\n        Use ISO-3166 country codes.\n        Examples:\n            - 'California, United States'\n            - 'Mexico'\n            - 'US' for United States\n            - 'DE' for Germany\n            - 'FR' for France\n        "}, 'output_format': {'anyOf': [{'enum': ['links', 'md', 'html'], 'type': 'string'}, {'type': 'null'}], 'default': None, 'description': "\n        The format of the output. Works only when parse parameter is false.\n            - links - Most efficient when the goal is navigation or finding specific URLs. Use this first when you need to locate a specific page within a website.\n            - md - Best for extracting and reading visible content once you've found the right page. Use this to get structured content that's easy to read and process.\n            - html - Should be used sparingly only when you need the raw HTML structure, JavaScript code, or styling information.\n        "}, 'user_agent_type': {'anyOf': [{'enum': ['desktop', 'desktop_chrome', 'desktop_firefox', 'desktop_safari', 'desktop_edge', 'desktop_opera', 'mobile', 'mobile_ios', 'mobile_android', 'tablet'], 'type': 'string'}, {'type': 'null'}], 'default': None, 'description': 'Device type and browser that will be used to determine User-Agent header value.'}}, 'additionalProperties': False}
Esquema de salida
{'type': 'object', 'required': ['result'], 'properties': {'result': {'type': 'string'}}, 'x-fastmcp-wrap-result': True}
universal_scraper
Get a content of any webpage. Supports browser rendering, parsing of certain webpages and different output formats.
Solo lectura
Esquema de entrada
{'type': 'object', 'required': ['url'], 'properties': {'url': {'type': 'string', 'description': 'Website url to scrape.'}, 'render': {'anyOf': [{'type': 'string', 'const': 'html'}, {'type': 'null'}], 'default': None, 'examples': ['html'], 'description': "\n        Whether a headless browser should be used to render the page.\n        For example:\n            - 'html' when browser is required to render the page.\n        "}, 'geo_location': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'examples': ['US', 'DE', 'FR'], 'description': "\n        The geographical location that the result should be adapted for.\n        Use ISO-3166 country codes.\n        Examples:\n            - 'California, United States'\n            - 'Mexico'\n            - 'US' for United States\n            - 'DE' for Germany\n            - 'FR' for France\n        "}, 'output_format': {'anyOf': [{'enum': ['links', 'md', 'html'], 'type': 'string'}, {'type': 'null'}], 'default': None, 'description': "\n        The format of the output. Works only when parse parameter is false.\n            - links - Most efficient when the goal is navigation or finding specific URLs. Use this first when you need to locate a specific page within a website.\n            - md - Best for extracting and reading visible content once you've found the right page. Use this to get structured content that's easy to read and process.\n            - html - Should be used sparingly only when you need the raw HTML structure, JavaScript code, or styling information.\n        "}, 'user_agent_type': {'anyOf': [{'enum': ['desktop', 'desktop_chrome', 'desktop_firefox', 'desktop_safari', 'desktop_edge', 'desktop_opera', 'mobile', 'mobile_ios', 'mobile_android', 'tablet'], 'type': 'string'}, {'type': 'null'}], 'default': None, 'description': 'Device type and browser that will be used to determine User-Agent header value.'}}, 'additionalProperties': False}
Esquema de salida
{'type': 'object', 'required': ['result'], 'properties': {'result': {'type': 'string'}}, 'x-fastmcp-wrap-result': True}
Añadido
amazon_product_scraper
17 de September de 2026 a las 12:53
Añadido
amazon_search_scraper
17 de September de 2026 a las 12:53
Añadido
google_search_scraper
17 de September de 2026 a las 12:53
Añadido
universal_scraper
17 de September de 2026 a las 12:53
Añadido
ai_map
17 de September de 2026 a las 12:53
Añadido
generate_schema
17 de September de 2026 a las 12:53
Añadido
ai_search
17 de September de 2026 a las 12:53
Añadido
ai_browser_agent
17 de September de 2026 a las 12:53
Añadido
ai_scraper
17 de September de 2026 a las 12:53
Añadido
ai_crawler
17 de September de 2026 a las 12:53