doc.page PDF Extraction
What this MCP does
Extracts PDFs into Markdown, structured tables, semantic RAG chunks, citations, and tracked shareable document links.
Tools
Input schema
{'type': 'object', 'required': ['url'], 'properties': {'url': {'type': 'string', 'description': 'http(s) URL of the PDF to publish (max 25 MB).'}, 'name': {'type': 'string', 'description': 'Display name in the library. Defaults to the filename.'}, 'slug': {'type': 'string', 'description': 'Custom vanity slug (premium plans only). Lowercase letters, digits and hyphens.'}, 'expiresAt': {'type': 'string', 'description': 'ISO 8601 date-time after which the link stops working. Omit for no expiry.'}, 'notifyOnOpen': {'type': 'boolean', 'description': 'Email the account owner on the first open. Default true.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'required': ['url'], 'properties': {'url': {'type': 'string', 'description': 'http(s) URL of the PDF to extract.'}, 'mode': {'enum': ['fast', 'hybrid'], 'type': 'string', 'description': 'fast = prose engine. hybrid = semantic engine with tables + bounding boxes when deployed; falls back to fast with a warning otherwise.'}, 'outputs': {'type': 'array', 'items': {'enum': ['markdown', 'elements', 'chunks', 'images'], 'type': 'string'}, 'description': 'Subset of outputs to include. Default: markdown and elements.'}, 'chunkTokens': {'type': 'integer', 'description': 'Target chunk size in tokens (when chunks are requested). Default 512.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'required': ['url'], 'properties': {'url': {'type': 'string', 'description': 'http(s) URL of the PDF to chunk.'}, 'maxTokens': {'type': 'integer', 'description': 'Target chunk size in tokens. Default 512.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'properties': {'id': {'type': 'string', 'description': 'Doc Link item id (from create_doc_link or list_doc_links).'}, 'slug': {'type': 'string', 'description': 'Doc Link slug — alternative to id.'}, 'include': {'type': 'array', 'items': {'enum': ['visits'], 'type': 'string'}, 'description': 'Extra sections. "visits" adds the recent visit rows (enriched on premium plans).'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'properties': {}, 'additionalProperties': False}
Input schema
{'type': 'object', 'required': ['url'], 'properties': {'url': {'type': 'string', 'description': 'http(s) URL of the PDF.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'properties': {'id': {'type': 'string', 'description': 'Doc Link item id.'}, 'slug': {'type': 'string', 'description': 'Doc Link slug — alternative to id.'}}, 'additionalProperties': False}
Recent tool changes
Similar MCP servers
Data Govt Nz
Searches and retrieves metadata for datasets, organizations, groups, and resources in New Zealand’s data.govt.nz CKAN open-data c…
Figshare
Searches and retrieves research outputs from Figshare, including papers, datasets, posters, code, metadata, files, collections, a…
Ror
Resolves research organization names and affiliation strings to canonical Research Organization Registry identifiers and records.
branchly
Manages an AI application’s knowledge-base content, prompts, tools, data sources, retrieval context, sessions, and usage analytic…
emberverse
Provides access to a persistent knowledge graph for reading, reasoning, and depositing information.
Islam West Africa Collection (IWAC)
Provides read-only search, retrieval, archival metadata, OCR, transcription, geographic, temporal, sentiment, topic, and semantic…
NicheDB
Aggregates scheduled public-data sources into searchable collections, enriched items, profiles, feeds, and future-event listings.
Plumbline — a trust layer for LLM answers
Queries verified workforce and payroll facts, grounds answers in SQL and evidence, checks answer provenance, and reports trust ve…