MCP Server

docforge

io.github.steviepinero/docforge
Knowledge & Documentation Productivity Public & reachable MCP 2025-11-25

What this MCP does

Extracts text, tables, and invoice data from PDFs and images, performs OCR, and converts markdown or HTML to printable PDFs.

extract_tables
Extract Tables from PDF
Detect and reconstruct tables from a text-based PDF. Returns each table as structured rows plus ready-to-use markdown and CSV renderings. Works best on PDFs with clear columnar layout (invoices, reports, statements).
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'file_url': {'type': 'string', 'format': 'uri', 'description': 'Public http(s) URL of the file'}, 'file_base64': {'type': 'string', 'description': 'Base64-encoded file contents (data-URI prefix allowed)'}}, 'additionalProperties': False}
ocr_image
OCR Image to Text
Run optical character recognition on an image (png, jpg, webp, bmp) and return the recognized text with a confidence score. Supports 100+ languages via the language parameter (ISO 639-2 codes like 'eng', 'deu', 'fra', 'spa').
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'file_url': {'type': 'string', 'format': 'uri', 'description': 'Public http(s) URL of the file'}, 'language': {'type': 'string', 'description': "Tesseract language code, default 'eng'"}, 'file_base64': {'type': 'string', 'description': 'Base64-encoded file contents (data-URI prefix allowed)'}}, 'additionalProperties': False}
parse_invoice
Parse Invoice/Receipt
Extract structured data from an invoice or receipt: vendor, invoice number, dates, currency, subtotal, tax, total, and line items. Accepts a text-based PDF, or an image when is_image is true (OCR is applied first). Returns JSON.
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'file_url': {'type': 'string', 'format': 'uri', 'description': 'Public http(s) URL of the file'}, 'is_image': {'type': 'boolean', 'description': 'Set true when the file is a photo/scan image rather than a PDF'}, 'file_base64': {'type': 'string', 'description': 'Base64-encoded file contents (data-URI prefix allowed)'}}, 'additionalProperties': False}
pdf_to_markdown
PDF to Markdown
Extract the text of a PDF and convert it to clean markdown. Detects headings by font size and preserves lists and paragraphs. Input: a text-based PDF via file_url or file_base64. For scanned PDFs use ocr_image on page images instead.
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'file_url': {'type': 'string', 'format': 'uri', 'description': 'Public http(s) URL of the file'}, 'file_base64': {'type': 'string', 'description': 'Base64-encoded file contents (data-URI prefix allowed)'}}, 'additionalProperties': False}
render_pdf
Render Markdown/HTML to PDF
Render markdown (or simple HTML) into a clean, printable A4 PDF. Supports headings, paragraphs, bullet and numbered lists, blockquotes, code blocks, horizontal rules, and inline bold/italic/code. Returns the PDF as base64 plus page count.
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['content'], 'properties': {'title': {'type': 'string', 'description': 'PDF document title metadata'}, 'format': {'enum': ['markdown', 'html'], 'type': 'string', 'description': 'Input format, default markdown'}, 'content': {'type': 'string', 'minLength': 1, 'description': 'The markdown or HTML source to render'}}, 'additionalProperties': False}
Added
parse_invoice
Sept. 17, 2026, 12:52 p.m.
Added
render_pdf
Sept. 17, 2026, 12:52 p.m.
Added
extract_tables
Sept. 17, 2026, 12:52 p.m.
Added
ocr_image
Sept. 17, 2026, 12:52 p.m.
Added
pdf_to_markdown
Sept. 17, 2026, 12:52 p.m.