MCP-Server

paper-mcp

io.github.MCPServings/paper-mcp
Wissenschaft & Engineering Suche & Recherche Öffentlich und erreichbar MCP 2025-11-25

Was dieses MCP kann

Searches and retrieves academic papers and citation data across arXiv, Semantic Scholar, OpenAlex, PubMed, and Europe PMC, with PDF extraction, OCR, and LaTeX tools.

autocomplete_papers
Semantic Scholar: autocomplete paper titles for a partial query (fast type-ahead).
Eingabeschema
{'type': 'object', 'title': 'autocomplete_papersArguments', 'required': ['query'], 'properties': {'query': {'type': 'string', 'title': 'Query'}}}
extract_pdf
Extract a PDF to clean Markdown/LaTeX text via MinerU (great for papers behind no open-access full text — give the user's PDF and get readable text back). Provide pdf_url (downloaded server-side, SSRF-guarded) OR pdf_base64. formula/table toggle math/table reconstruction. Returns {task_id, status, cached, content, chars}: a recently-seen (cached) or small PDF comes back with `content` in one call; a fresh PDF (MinerU is GPU-heavy, minutes) returns status='running' + a task_id — then call extract_pdf_result(task_id) to fetch the text.
Eingabeschema
{'type': 'object', 'title': 'extract_pdfArguments', 'properties': {'table': {'type': 'boolean', 'title': 'Table', 'default': True}, 'formula': {'type': 'boolean', 'title': 'Formula', 'default': True}, 'pdf_url': {'type': 'string', 'title': 'Pdf Url', 'default': ''}, 'pdf_base64': {'type': 'string', 'title': 'Pdf Base64', 'default': ''}}}
extract_pdf_result
Fetch the result of an extract_pdf job by task_id. Returns {task_id, status, content, chars}: `content` is the extracted text once status='done'; while still 'running' content is null — call again shortly. Results expire server-side, so fetch reasonably soon.
Eingabeschema
{'type': 'object', 'title': 'extract_pdf_resultArguments', 'required': ['task_id'], 'properties': {'task_id': {'type': 'string', 'title': 'Task Id'}}}
get_author
Semantic Scholar: a single author's profile by id.
Eingabeschema
{'type': 'object', 'title': 'get_authorArguments', 'required': ['author_id'], 'properties': {'author_id': {'type': 'string', 'title': 'Author Id'}}}
get_author_papers
Semantic Scholar: all papers by a given author id, newest first.
Eingabeschema
{'type': 'object', 'title': 'get_author_papersArguments', 'required': ['author_id'], 'properties': {'start': {'type': 'integer', 'title': 'Start', 'default': 0}, 'author_id': {'type': 'string', 'title': 'Author Id'}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 20}}}
get_authors_batch
Semantic Scholar: fetch many authors at once by id.
Eingabeschema
{'type': 'object', 'title': 'get_authors_batchArguments', 'required': ['ids'], 'properties': {'ids': {'type': 'array', 'items': {'type': 'string'}, 'title': 'Ids'}}}
get_dataset_diffs
Semantic Scholar Datasets: incremental diff (added/updated/deleted) for a dataset between two releases. Needs the key.
Eingabeschema
{'type': 'object', 'title': 'get_dataset_diffsArguments', 'required': ['dataset_name', 'start_release'], 'properties': {'end_release': {'type': 'string', 'title': 'End Release', 'default': 'latest'}, 'dataset_name': {'type': 'string', 'title': 'Dataset Name'}, 'start_release': {'type': 'string', 'title': 'Start Release'}}}
get_dataset_download_links
Semantic Scholar Datasets: get download links (presigned URLs) for one dataset in a release. Needs the API key.
Eingabeschema
{'type': 'object', 'title': 'get_dataset_download_linksArguments', 'required': ['dataset_name'], 'properties': {'release_id': {'type': 'string', 'title': 'Release Id', 'default': 'latest'}, 'dataset_name': {'type': 'string', 'title': 'Dataset Name'}}}
get_dataset_release
Semantic Scholar Datasets: which datasets a release contains (papers, abstracts, citations, embeddings, s2orc, tldrs…). release_id defaults to 'latest'.
Eingabeschema
{'type': 'object', 'title': 'get_dataset_releaseArguments', 'properties': {'release_id': {'type': 'string', 'title': 'Release Id', 'default': 'latest'}}}
get_openalex_citations
OpenAlex: papers that CITE this work (forward citation graph), most-cited first.
Eingabeschema
{'type': 'object', 'title': 'get_openalex_citationsArguments', 'required': ['work_id'], 'properties': {'start': {'type': 'integer', 'title': 'Start', 'default': 0}, 'work_id': {'type': 'string', 'title': 'Work Id'}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 10}}}
get_openalex_references
OpenAlex: the works this one REFERENCES (its bibliography).
Eingabeschema
{'type': 'object', 'title': 'get_openalex_referencesArguments', 'required': ['work_id'], 'properties': {'work_id': {'type': 'string', 'title': 'Work Id'}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 25}}}
get_openalex_trends
OpenAlex: publication-trend analytics for a query — counts grouped by year (default), or by 'institutions.id', 'authorships.author.id', 'open_access.is_oa', 'type', 'language'. Returns aggregate counts only (cheap, no rows).
Eingabeschema
{'type': 'object', 'title': 'get_openalex_trendsArguments', 'required': ['query'], 'properties': {'query': {'type': 'string', 'title': 'Query'}, 'group_by': {'type': 'string', 'title': 'Group By', 'default': 'publication_year'}}}
get_openalex_work
OpenAlex: fetch one work's full record (316M-work, all-field corpus). id accepts OpenAlex Wxxxx, a DOI, or an arXiv id.
Eingabeschema
{'type': 'object', 'title': 'get_openalex_workArguments', 'required': ['work_id'], 'properties': {'work_id': {'type': 'string', 'title': 'Work Id'}}}
get_paper
Fetch one paper by id, with full abstract and PDF link.
Eingabeschema
{'type': 'object', 'title': 'get_paperArguments', 'required': ['paper_id'], 'properties': {'source': {'type': 'string', 'title': 'Source', 'default': 'arxiv'}, 'paper_id': {'type': 'string', 'title': 'Paper Id'}}}
get_paper_authors
Semantic Scholar: the authors of a paper (with h-index, paper/citation counts).
Eingabeschema
{'type': 'object', 'title': 'get_paper_authorsArguments', 'required': ['paper_id'], 'properties': {'start': {'type': 'integer', 'title': 'Start', 'default': 0}, 'paper_id': {'type': 'string', 'title': 'Paper Id'}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 100}}}
get_paper_citations
Semantic Scholar: papers that CITE this one (forward citation graph). id accepts S2 id / DOI: / ARXIV: / CorpusId:.
Eingabeschema
{'type': 'object', 'title': 'get_paper_citationsArguments', 'required': ['paper_id'], 'properties': {'start': {'type': 'integer', 'title': 'Start', 'default': 0}, 'paper_id': {'type': 'string', 'title': 'Paper Id'}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 10}}}
get_paper_references
Semantic Scholar: papers this one REFERENCES (its bibliography). id accepts S2 id / DOI: / ARXIV: / CorpusId:.
Eingabeschema
{'type': 'object', 'title': 'get_paper_referencesArguments', 'required': ['paper_id'], 'properties': {'start': {'type': 'integer', 'title': 'Start', 'default': 0}, 'paper_id': {'type': 'string', 'title': 'Paper Id'}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 10}}}
get_papers_batch
Semantic Scholar: fetch many papers at once by id (S2/DOI:/ARXIV:/CorpusId:), up to ~500 per call.
Eingabeschema
{'type': 'object', 'title': 'get_papers_batchArguments', 'required': ['ids'], 'properties': {'ids': {'type': 'array', 'items': {'type': 'string'}, 'title': 'Ids'}}}
lint_latex
Lint a LaTeX snippet: report errors and return an auto-fixed version. Input `code` (the LaTeX source). Returns {errors, fixed_code, summary_en, summary_zh, elapsed_ms}.
Eingabeschema
{'type': 'object', 'title': 'lint_latexArguments', 'required': ['code'], 'properties': {'code': {'type': 'string', 'title': 'Code'}}}
list_categories
List common subject category codes for filtering/recent.
Eingabeschema
{'type': 'object', 'title': 'list_categoriesArguments', 'properties': {'source': {'type': 'string', 'title': 'Source', 'default': 'arxiv'}}}
list_dataset_releases
Semantic Scholar Datasets: list all available release ids (dated snapshots of the full corpus).
Eingabeschema
{'type': 'object', 'title': 'list_dataset_releasesArguments', 'properties': {}}
list_ocr_models
List the OCR models available for recognize_formula / recognize_table.
Eingabeschema
{'type': 'object', 'title': 'list_ocr_modelsArguments', 'properties': {}}
list_openalex_topics
OpenAlex: search the topic taxonomy (~4500 topics) to find the right subject term for filtering or recent-work queries.
Eingabeschema
{'type': 'object', 'title': 'list_openalex_topicsArguments', 'required': ['query'], 'properties': {'query': {'type': 'string', 'title': 'Query'}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 15}}}
list_paper_sources
List available paper corpora.
Eingabeschema
{'type': 'object', 'title': 'list_paper_sourcesArguments', 'properties': {}}
list_recent
List the latest papers in a subject category, newest first.
Eingabeschema
{'type': 'object', 'title': 'list_recentArguments', 'required': ['category'], 'properties': {'start': {'type': 'integer', 'title': 'Start', 'default': 0}, 'source': {'type': 'string', 'title': 'Source', 'default': 'arxiv'}, 'category': {'type': 'string', 'title': 'Category'}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 10}}}
match_paper_title
Semantic Scholar: find the single paper whose title best matches the given text (exact-match lookup).
Eingabeschema
{'type': 'object', 'title': 'match_paper_titleArguments', 'required': ['title'], 'properties': {'title': {'type': 'string', 'title': 'Title'}}}
read_paper
Read a paper's full text. format='markdown' (default, body with formulas as $LaTeX$), 'html' (raw LaTeXML HTML), or 'latex' (the original LaTeX manuscript from the e-print source). arXiv only; id like 2401.01234.
Eingabeschema
{'type': 'object', 'title': 'read_paperArguments', 'required': ['paper_id'], 'properties': {'format': {'type': 'string', 'title': 'Format', 'default': 'markdown'}, 'source': {'type': 'string', 'title': 'Source', 'default': 'arxiv'}, 'paper_id': {'type': 'string', 'title': 'Paper Id'}}}
recognize_formula
Recognize a math formula from an image and return LaTeX. Provide image_url (downloaded server-side) OR image_base64. model: deepseek-ocr (default), paddleocr-vl, or texify. Returns {latex, model, elapsed_ms}.
Eingabeschema
{'type': 'object', 'title': 'recognize_formulaArguments', 'properties': {'model': {'type': 'string', 'title': 'Model', 'default': 'deepseek-ocr'}, 'image_url': {'type': 'string', 'title': 'Image Url', 'default': ''}, 'image_base64': {'type': 'string', 'title': 'Image Base64', 'default': ''}}}
recognize_table
Recognize a table from an image and return LaTeX tabular code. Provide image_url OR image_base64. model: deepseek-ocr (default), paddleocr-vl, or texify. Returns {latex, model, elapsed_ms}.
Eingabeschema
{'type': 'object', 'title': 'recognize_tableArguments', 'properties': {'model': {'type': 'string', 'title': 'Model', 'default': 'deepseek-ocr'}, 'image_url': {'type': 'string', 'title': 'Image Url', 'default': ''}, 'image_base64': {'type': 'string', 'title': 'Image Base64', 'default': ''}}}
recommend_papers_for_paper
Semantic Scholar: recommend papers similar to one paper. pool='recent' (last open corpus) or 'all-cs' (all of CS). If the 'recent' pool yields nothing (common for older papers), it automatically retries the 'all-cs' pool.
Eingabeschema
{'type': 'object', 'title': 'recommend_papers_for_paperArguments', 'required': ['paper_id'], 'properties': {'pool': {'type': 'string', 'title': 'Pool', 'default': 'recent'}, 'paper_id': {'type': 'string', 'title': 'Paper Id'}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 10}}}
recommend_papers_from_examples
Semantic Scholar: recommend papers from positive (and optional negative) example paper ids.
Eingabeschema
{'type': 'object', 'title': 'recommend_papers_from_examplesArguments', 'required': ['positive_ids'], 'properties': {'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 10}, 'negative_ids': {'anyOf': [{'type': 'array', 'items': {'type': 'string'}}, {'type': 'null'}], 'title': 'Negative Ids', 'default': None}, 'positive_ids': {'type': 'array', 'items': {'type': 'string'}, 'title': 'Positive Ids'}}}
search_all
Aggregated search across arXiv, Semantic Scholar and OpenAlex at once. Fans out concurrently, de-duplicates the same work across corpora (by DOI or title) and re-ranks with Reciprocal Rank Fusion, so papers found by several sources rank highest. Each hit lists which `sources` found it and an `ids` map ({source: id}) you can pass to get_paper / read_paper / the citation tools. Prefer this over search_papers for a broad lookup.
Eingabeschema
{'type': 'object', 'title': 'search_allArguments', 'required': ['query'], 'properties': {'query': {'type': 'string', 'title': 'Query'}, 'sources': {'type': 'string', 'title': 'Sources', 'default': 'arxiv,semanticscholar,openalex'}, 'per_source': {'type': 'integer', 'title': 'Per Source', 'default': 0}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 10}}}
search_authors
Semantic Scholar: search for authors by name; returns profiles with h-index and paper/citation counts.
Eingabeschema
{'type': 'object', 'title': 'search_authorsArguments', 'required': ['query'], 'properties': {'query': {'type': 'string', 'title': 'Query'}, 'start': {'type': 'integer', 'title': 'Start', 'default': 0}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 10}}}
search_by_author
Find papers by a specific author, newest first.
Eingabeschema
{'type': 'object', 'title': 'search_by_authorArguments', 'required': ['author'], 'properties': {'start': {'type': 'integer', 'title': 'Start', 'default': 0}, 'author': {'type': 'string', 'title': 'Author'}, 'source': {'type': 'string', 'title': 'Source', 'default': 'arxiv'}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 10}}}
search_medical
Evidence-graded MEDICAL literature search (PubMed + Europe PMC). Unlike search_all (generic, ranks high-cited reviews/guidelines above trials), this filters by research type via PubMed Publication-Type tags and re-ranks by the evidence pyramid (meta-analysis / systematic review > RCT > cohort > ...), so the actual clinical trials surface first. Open-access full text is pulled from Europe PMC by PMID. `query` should be English keyword/boolean text (PubMed maps it); do natural-language/multilingual understanding upstream. Returns hits with pmid/doi/study_type/evidence_level/citations/abstract and, when open-access, fulltext.
Eingabeschema
{'type': 'object', 'title': 'search_medicalArguments', 'required': ['query'], 'properties': {'query': {'type': 'string', 'title': 'Query'}, 'year_from': {'type': 'integer', 'title': 'Year From', 'default': 0}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 10}, 'study_types': {'type': 'string', 'title': 'Study Types', 'default': 'rct,meta-analysis,systematic-review'}, 'fetch_fulltext': {'type': 'boolean', 'title': 'Fetch Fulltext', 'default': True}}}
search_openalex_authors
OpenAlex: search authors; returns profiles with h-index, i10-index, works/citation counts and institutions.
Eingabeschema
{'type': 'object', 'title': 'search_openalex_authorsArguments', 'required': ['query'], 'properties': {'query': {'type': 'string', 'title': 'Query'}, 'start': {'type': 'integer', 'title': 'Start', 'default': 0}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 10}}}
search_openalex_institutions
OpenAlex: search institutions (universities, labs) with ROR id, country, works/citation counts.
Eingabeschema
{'type': 'object', 'title': 'search_openalex_institutionsArguments', 'required': ['query'], 'properties': {'query': {'type': 'string', 'title': 'Query'}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 10}}}
search_openalex_works
OpenAlex: advanced filtered work search. Filters: from_year, to_year, is_oa (open access only), min_citations, institution_id. sort_by: relevance|newest|cited.
Eingabeschema
{'type': 'object', 'title': 'search_openalex_worksArguments', 'properties': {'is_oa': {'type': 'boolean', 'title': 'Is Oa', 'default': False}, 'query': {'type': 'string', 'title': 'Query', 'default': ''}, 'sort_by': {'type': 'string', 'title': 'Sort By', 'default': 'relevance'}, 'to_year': {'type': 'integer', 'title': 'To Year', 'default': 0}, 'from_year': {'type': 'integer', 'title': 'From Year', 'default': 0}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 25}, 'min_citations': {'type': 'integer', 'title': 'Min Citations', 'default': 0}, 'institution_id': {'type': 'string', 'title': 'Institution Id', 'default': ''}}}
search_papers
Search academic papers. Returns normalized hits with a short abstract preview; call get_paper for the full record.
Eingabeschema
{'type': 'object', 'title': 'search_papersArguments', 'required': ['query'], 'properties': {'query': {'type': 'string', 'title': 'Query'}, 'start': {'type': 'integer', 'title': 'Start', 'default': 0}, 'source': {'type': 'string', 'title': 'Source', 'default': 'arxiv'}, 'sort_by': {'type': 'string', 'title': 'Sort By', 'default': 'relevance'}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 10}}}
search_papers_bulk
Semantic Scholar: bulk paper search (up to 1000 hits, sortable e.g. 'citationCount:desc' or 'publicationDate:desc', with a continuation token). Filters: fields_of_study, year (e.g. '2020-2024'), venue, publication_types, open_access_pdf.
Eingabeschema
{'type': 'object', 'title': 'search_papers_bulkArguments', 'required': ['query'], 'properties': {'sort': {'type': 'string', 'title': 'Sort', 'default': ''}, 'year': {'type': 'string', 'title': 'Year', 'default': ''}, 'query': {'type': 'string', 'title': 'Query'}, 'token': {'type': 'string', 'title': 'Token', 'default': ''}, 'venue': {'type': 'string', 'title': 'Venue', 'default': ''}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 100}, 'fields_of_study': {'type': 'string', 'title': 'Fields Of Study', 'default': ''}, 'open_access_pdf': {'type': 'boolean', 'title': 'Open Access Pdf', 'default': False}, 'publication_types': {'type': 'string', 'title': 'Publication Types', 'default': ''}}}
search_snippets
Semantic Scholar: search INSIDE paper full text and return matching text snippets (not just titles/abstracts).
Eingabeschema
{'type': 'object', 'title': 'search_snippetsArguments', 'required': ['query'], 'properties': {'query': {'type': 'string', 'title': 'Query'}, 'max_results': {'type': 'integer', 'title': 'Max Results', 'default': 10}}}
Hinzugefügt
list_openalex_topics
17. September 2026 12:45
Hinzugefügt
get_openalex_trends
17. September 2026 12:45
Hinzugefügt
search_openalex_works
17. September 2026 12:45
Hinzugefügt
search_openalex_institutions
17. September 2026 12:45
Hinzugefügt
search_openalex_authors
17. September 2026 12:45
Hinzugefügt
get_openalex_references
17. September 2026 12:45
Hinzugefügt
get_openalex_citations
17. September 2026 12:45
Hinzugefügt
get_openalex_work
17. September 2026 12:45
Hinzugefügt
get_dataset_diffs
17. September 2026 12:45
Hinzugefügt
get_dataset_download_links
17. September 2026 12:45
Hinzugefügt
get_dataset_release
17. September 2026 12:45
Hinzugefügt
list_dataset_releases
17. September 2026 12:45
Hinzugefügt
recommend_papers_from_examples
17. September 2026 12:45
Hinzugefügt
recommend_papers_for_paper
17. September 2026 12:45
Hinzugefügt
search_snippets
17. September 2026 12:45
Hinzugefügt
get_authors_batch
17. September 2026 12:45
Hinzugefügt
get_author_papers
17. September 2026 12:45
Hinzugefügt
get_author
17. September 2026 12:45
Hinzugefügt
search_authors
17. September 2026 12:45
Hinzugefügt
get_papers_batch
17. September 2026 12:45
Hinzugefügt
search_papers_bulk
17. September 2026 12:45
Hinzugefügt
autocomplete_papers
17. September 2026 12:45
Hinzugefügt
match_paper_title
17. September 2026 12:45
Hinzugefügt
get_paper_authors
17. September 2026 12:45
Hinzugefügt
get_paper_references
17. September 2026 12:45
Hinzugefügt
get_paper_citations
17. September 2026 12:45
Hinzugefügt
extract_pdf_result
17. September 2026 12:45
Hinzugefügt
extract_pdf
17. September 2026 12:45
Hinzugefügt
lint_latex
17. September 2026 12:45
Hinzugefügt
list_ocr_models
17. September 2026 12:45