Servidor MCP

X1-BaaS Scraping API

com.tazpal/x1-baas

Qué hace este MCP

Scrapes public URLs with browser rendering and returns cleaned Markdown, HTML, screenshots, PDFs, or CSV data, using crypto-based payment instructions.

get_pricing
Check current pricing and payment requirements for the BaaS scrape engine. Returns pricing info, supported networks, and payment instructions. No authentication required. Returns: Current pricing details and payment instructions.
Solo lectura Idempotente
Esquema de entrada
{'type': 'object', 'title': 'get_pricingArguments', 'properties': {}}
Esquema de salida
{'type': 'object', 'title': 'get_pricingOutput', 'required': ['result'], 'properties': {'result': {'type': 'string', 'title': 'Result'}}}
scrape
Scrape a URL and return content in your preferred format. Supported output formats: - markdown (default): Clean LLM-ready Markdown text - screenshot: PNG/JPEG image of the page - pdf: PDF document of the page - csv: Table data extracted as CSV - html: Sanitized HTML with scripts/ads removed This tool handles: - JavaScript rendering (SPA, dynamic content) - Anti-bot bypass (Cloudflare Turnstile, Datadome) - DOM cleaning (strips scripts, nav, footer, ads) - HTML-to-Markdown conversion (Mozilla Readability engine) - Automatic retry with escalating wait strategies - Domain cooldown to avoid rate-limiting - Response caching (5 min TTL) Args: url: The URL to scrape (must start with http:// or https://) output: Output format: "markdown" (default), "screenshot", "pdf", "csv", "html" wait_for_selector: Optional CSS selector to wait for before extraction (e.g., ".article-content") timeout_ms: Navigation timeout in milliseconds (default: 20000, max: 120000) block_media: Block images/fonts/video for faster loading (default: true) wait_strategy: Wait strategy: "default", "spa", "heavy", "cloudflare" (auto-detected if omitted) retry: Enable automatic retry on failure (default: true) bypass_cache: Skip cache, force fresh scrape (default: false) javascript: Custom JavaScript to execute after page load (e.g., "window.scrollTo(0, 1000)") Returns: Content in the requested format, or an error message.
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', 'title': 'scrapeArguments', 'required': ['url'], 'properties': {'url': {'type': 'string', 'title': 'Url'}, 'retry': {'type': 'boolean', 'title': 'Retry', 'default': True}, 'output': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'title': 'Output', 'default': None}, 'javascript': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'title': 'Javascript', 'default': None}, 'timeout_ms': {'type': 'integer', 'title': 'Timeout Ms', 'default': 20000}, 'block_media': {'type': 'boolean', 'title': 'Block Media', 'default': True}, 'bypass_cache': {'type': 'boolean', 'title': 'Bypass Cache', 'default': False}, 'wait_strategy': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'title': 'Wait Strategy', 'default': None}, 'wait_for_selector': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'title': 'Wait For Selector', 'default': None}}}
Esquema de salida
{'type': 'object', 'title': 'scrapeOutput', 'required': ['result'], 'properties': {'result': {'type': 'string', 'title': 'Result'}}}
server_status
Check the BaaS engine health and operational status. Returns browser status, uptime, x402 payment mode, and context count. Returns: Server health and diagnostic information.
Solo lectura Idempotente
Esquema de entrada
{'type': 'object', 'title': 'server_statusArguments', 'properties': {}}
Esquema de salida
{'type': 'object', 'title': 'server_statusOutput', 'required': ['result'], 'properties': {'result': {'type': 'string', 'title': 'Result'}}}
Añadido
server_status
17 de September de 2026 a las 12:37
Añadido
get_pricing
17 de September de 2026 a las 12:37
Añadido
scrape
17 de September de 2026 a las 12:37