Servidor MCP

doc.page PDF Extraction

page.doc/pdf-extract
Datos y analítica Conocimiento y documentación Público y accesible MCP 2026-07-28

Qué hace este MCP

Extracts PDFs into Markdown, structured tables, semantic RAG chunks, citations, and tracked shareable document links.

create_doc_link
Publish a PDF as a tracked doc.page Doc Link and get back a shareable URL. The link belongs to the API key's account and also appears in its doc.page library. Requires an API key. Free plan: up to 3 active links; custom vanity slugs are premium-only. Optional expiry and open-notification toggle.
Esquema de entrada
{'type': 'object', 'required': ['url'], 'properties': {'url': {'type': 'string', 'description': 'http(s) URL of the PDF to publish (max 25 MB).'}, 'name': {'type': 'string', 'description': 'Display name in the library. Defaults to the filename.'}, 'slug': {'type': 'string', 'description': 'Custom vanity slug (premium plans only). Lowercase letters, digits and hyphens.'}, 'expiresAt': {'type': 'string', 'description': 'ISO 8601 date-time after which the link stops working. Omit for no expiry.'}, 'notifyOnOpen': {'type': 'boolean', 'description': 'Email the account owner on the first open. Default true.'}}, 'additionalProperties': False}
extract_pdf
Extract a PDF into clean Markdown and structured elements (headings, paragraphs). Returns the canonical ExtractedDocument object. mode "hybrid" runs a heavier semantic engine that also reconstructs tables and bounding boxes; the default "fast" engine is prose-only (low confidence.tables).
Esquema de entrada
{'type': 'object', 'required': ['url'], 'properties': {'url': {'type': 'string', 'description': 'http(s) URL of the PDF to extract.'}, 'mode': {'enum': ['fast', 'hybrid'], 'type': 'string', 'description': 'fast = prose engine. hybrid = semantic engine with tables + bounding boxes when deployed; falls back to fast with a warning otherwise.'}, 'outputs': {'type': 'array', 'items': {'enum': ['markdown', 'elements', 'chunks', 'images'], 'type': 'string'}, 'description': 'Subset of outputs to include. Default: markdown and elements.'}, 'chunkTokens': {'type': 'integer', 'description': 'Target chunk size in tokens (when chunks are requested). Default 512.'}}, 'additionalProperties': False}
get_chunks
Split a PDF into semantic chunks ready for embeddings (RAG). Each chunk carries its text, estimated tokens, starting page, section heading and the source element ids for citation.
Esquema de entrada
{'type': 'object', 'required': ['url'], 'properties': {'url': {'type': 'string', 'description': 'http(s) URL of the PDF to chunk.'}, 'maxTokens': {'type': 'integer', 'description': 'Target chunk size in tokens. Default 512.'}}, 'additionalProperties': False}
get_doc_link_stats
Reading analytics for one Doc Link of the API key's account, by id or slug. Always returns the summary (total views, unique visitors, last visit). Premium plans additionally get countries, visitor companies (as_org) and per-page views + average dwell time; pass include:["visits"] for the recent visit rows. Requires an API key.
Esquema de entrada
{'type': 'object', 'properties': {'id': {'type': 'string', 'description': 'Doc Link item id (from create_doc_link or list_doc_links).'}, 'slug': {'type': 'string', 'description': 'Doc Link slug — alternative to id.'}, 'include': {'type': 'array', 'items': {'enum': ['visits'], 'type': 'string'}, 'description': 'Extra sections. "visits" adds the recent visit rows (enriched on premium plans).'}}, 'additionalProperties': False}
list_doc_links
List the Doc Links of the API key's account (id, slug, URL, name, disabled/expiry state, total views, last view). Use this to recover links created in earlier sessions before querying stats. Requires an API key.
Esquema de entrada
{'type': 'object', 'properties': {}, 'additionalProperties': False}
list_tables
Return every table in a PDF as structured JSON (reconstructed rows and columns) with page and bounding box for verifiable citations. Uses the semantic (hybrid) engine.
Esquema de entrada
{'type': 'object', 'required': ['url'], 'properties': {'url': {'type': 'string', 'description': 'http(s) URL of the PDF.'}}, 'additionalProperties': False}
revoke_doc_link
Disable a Doc Link of the API key's account (by id or slug) so the public URL stops serving. The item and its stats remain in the library; on the free plan this frees an active-link slot. Requires an API key.
Esquema de entrada
{'type': 'object', 'properties': {'id': {'type': 'string', 'description': 'Doc Link item id.'}, 'slug': {'type': 'string', 'description': 'Doc Link slug — alternative to id.'}}, 'additionalProperties': False}
Añadido
revoke_doc_link
17 de September de 2026 a las 12:54
Añadido
get_doc_link_stats
17 de September de 2026 a las 12:54
Añadido
list_doc_links
17 de September de 2026 a las 12:54
Añadido
create_doc_link
17 de September de 2026 a las 12:54
Añadido
list_tables
17 de September de 2026 a las 12:54
Añadido
get_chunks
17 de September de 2026 a las 12:54
Añadido
extract_pdf
17 de September de 2026 a las 12:54