MCP 服务器

docforge

io.github.steviepinero/docforge
知识与文档 生产力 公开且可连接 MCP 2025-11-25

此 MCP 可以做什么

Extracts text, tables, and invoice data from PDFs and images, performs OCR, and converts markdown or HTML to printable PDFs.

extract_tables
Extract Tables from PDF
Detect and reconstruct tables from a text-based PDF. Returns each table as structured rows plus ready-to-use markdown and CSV renderings. Works best on PDFs with clear columnar layout (invoices, reports, statements).
输入模式
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'file_url': {'type': 'string', 'format': 'uri', 'description': 'Public http(s) URL of the file'}, 'file_base64': {'type': 'string', 'description': 'Base64-encoded file contents (data-URI prefix allowed)'}}, 'additionalProperties': False}
ocr_image
OCR Image to Text
Run optical character recognition on an image (png, jpg, webp, bmp) and return the recognized text with a confidence score. Supports 100+ languages via the language parameter (ISO 639-2 codes like 'eng', 'deu', 'fra', 'spa').
输入模式
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'file_url': {'type': 'string', 'format': 'uri', 'description': 'Public http(s) URL of the file'}, 'language': {'type': 'string', 'description': "Tesseract language code, default 'eng'"}, 'file_base64': {'type': 'string', 'description': 'Base64-encoded file contents (data-URI prefix allowed)'}}, 'additionalProperties': False}
parse_invoice
Parse Invoice/Receipt
Extract structured data from an invoice or receipt: vendor, invoice number, dates, currency, subtotal, tax, total, and line items. Accepts a text-based PDF, or an image when is_image is true (OCR is applied first). Returns JSON.
输入模式
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'file_url': {'type': 'string', 'format': 'uri', 'description': 'Public http(s) URL of the file'}, 'is_image': {'type': 'boolean', 'description': 'Set true when the file is a photo/scan image rather than a PDF'}, 'file_base64': {'type': 'string', 'description': 'Base64-encoded file contents (data-URI prefix allowed)'}}, 'additionalProperties': False}
pdf_to_markdown
PDF to Markdown
Extract the text of a PDF and convert it to clean markdown. Detects headings by font size and preserves lists and paragraphs. Input: a text-based PDF via file_url or file_base64. For scanned PDFs use ocr_image on page images instead.
输入模式
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'file_url': {'type': 'string', 'format': 'uri', 'description': 'Public http(s) URL of the file'}, 'file_base64': {'type': 'string', 'description': 'Base64-encoded file contents (data-URI prefix allowed)'}}, 'additionalProperties': False}
render_pdf
Render Markdown/HTML to PDF
Render markdown (or simple HTML) into a clean, printable A4 PDF. Supports headings, paragraphs, bullet and numbered lists, blockquotes, code blocks, horizontal rules, and inline bold/italic/code. Returns the PDF as base64 plus page count.
输入模式
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['content'], 'properties': {'title': {'type': 'string', 'description': 'PDF document title metadata'}, 'format': {'enum': ['markdown', 'html'], 'type': 'string', 'description': 'Input format, default markdown'}, 'content': {'type': 'string', 'minLength': 1, 'description': 'The markdown or HTML source to render'}}, 'additionalProperties': False}
已添加
parse_invoice
2026年9月17日 12:52
已添加
render_pdf
2026年9月17日 12:52
已添加
extract_tables
2026年9月17日 12:52
已添加
ocr_image
2026年9月17日 12:52
已添加
pdf_to_markdown
2026年9月17日 12:52