MCPサーバー

arxiv-mcp-server

io.github.cyanheads/arxiv-mcp-server
科学・工学 検索・リサーチ 公開・接続可能 MCP 2025-11-25

このMCPでできること

Searches arXiv, retrieves paper metadata, and reads full-text scientific papers.

arxiv_get_metadata
Arxiv Get Metadata
Get full metadata for one or more arXiv papers by ID. Use when you have known IDs from citations, prior search results, or memory.
読み取り専用
入力スキーマ
{'type': 'object', '$schema': 'https://json-schema.org/draft/2020-12/schema', 'required': ['paper_ids'], 'properties': {'paper_ids': {'anyOf': [{'type': 'string', 'minLength': 1, 'description': 'Single arXiv paper ID (e.g., "2401.12345" or "2401.12345v2").'}, {'type': 'array', 'items': {'type': 'string', 'minLength': 1}, 'maxItems': 10, 'minItems': 1, 'description': 'Array of up to 10 arXiv paper IDs for batch lookup.'}], 'description': 'arXiv paper ID or array of up to 10 IDs. Format: "2401.12345" or "2401.12345v2" (with version). Also accepts legacy IDs like "hep-th/9901001".'}}, 'additionalProperties': False}
出力スキーマ
{'type': 'object', 'anyOf': [{'not': {'required': ['error']}, 'required': ['papers', 'totalSucceeded']}, {'required': ['error']}], '$schema': 'https://json-schema.org/draft/2020-12/schema', 'properties': {'error': {'type': 'object', 'required': ['code', 'message'], 'properties': {'code': {'type': 'integer', 'maximum': 9007199254740991, 'minimum': -9007199254740991, 'description': 'JSON-RPC error code for this failure.'}, 'data': {'type': 'object', 'properties': {'reason': {'type': 'string', 'examples': ['no_match', 'version_unavailable', 'rate_limited', 'invalid_request'], 'description': 'Machine-readable failure mode. Declared by this tool: `no_match`: None of the requested IDs returned data from arXiv. `version_unavailable`: Every requested ID pinned a version the local mirror does not hold, and live arXiv fallback is disabled. `rate_limited`: arXiv has throttled requests (HTTP 429 or "Rate exceeded." body). `invalid_request`: arXiv rejected the request (HTTP 4xx other than 429), e.g. malformed ID syntax. Other values are possible when a failure originates below the handler.'}, 'recovery': {'type': 'object', 'required': ['hint'], 'properties': {'hint': {'type': 'string'}}, 'description': 'Actionable next step for the caller.', 'additionalProperties': {}}, 'retryable': {'type': 'boolean', 'description': 'Whether retrying may succeed.'}}, 'additionalProperties': {}}, 'message': {'type': 'string', 'description': 'Human-readable description of what went wrong.'}}, 'description': 'Present when the call failed. Absent on success.', 'additionalProperties': {}}, 'papers': {'type': 'array', 'items': {'type': 'object', 'required': ['id', 'title', 'authors', 'abstract', 'primary_category', 'categories', 'published', 'updated', 'pdf_url', 'abstract_url'], 'properties': {'id': {'type': 'string', 'description': 'arXiv paper ID (e.g., "2401.12345v1").'}, 'doi': {'type': 'string', 'description': 'DOI if available.'}, 'title': {'type': 'string', 'description': 'Paper title.'}, 'authors': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Author names.'}, 'comment': {'type': 'string', 'description': 'Author comment (e.g., page count, conference).'}, 'pdf_url': {'type': 'string', 'description': 'Direct PDF download URL.'}, 'updated': {'type': 'string', 'description': 'Last update date (ISO 8601).'}, 'abstract': {'type': 'string', 'description': 'Full abstract text.'}, 'published': {'type': 'string', 'description': 'Original submission date (ISO 8601).'}, 'categories': {'type': 'array', 'items': {'type': 'string'}, 'description': 'All arXiv categories assigned to this paper.'}, 'journal_ref': {'type': 'string', 'description': 'Journal reference if published.'}, 'abstract_url': {'type': 'string', 'description': 'arXiv abstract page URL.'}, 'primary_category': {'type': 'string', 'description': 'Primary arXiv category (e.g., "cs.CL").'}}, 'description': 'arXiv paper metadata â\x80\x94 identifier, title, authors, abstract, categories, and links.', 'additionalProperties': False}, 'description': 'Papers found. May be fewer than requested if some IDs are invalid.'}, 'not_found': {'type': 'array', 'items': {'type': 'object', 'required': ['id', 'reason'], 'properties': {'id': {'type': 'string', 'description': 'arXiv ID that returned no data.'}, 'detail': {'type': 'string', 'description': 'Additional human-readable context, when available'}, 'reason': {'enum': ['not_in_arxiv', 'version_not_in_mirror'], 'type': 'string', 'description': 'Why the paper ID could not be returned.'}}, 'description': 'A requested ID that could not be returned, with the reason it was missed.', 'additionalProperties': False}, 'description': 'Per-input explanations for inputs that could not be returned. Absent when nothing failed.'}, 'totalSucceeded': {'type': 'integer', 'maximum': 9007199254740991, 'minimum': 0, 'description': "Number of successful items in 'papers'"}}, 'additionalProperties': False}
arxiv_list_categories
Arxiv List Categories
List arXiv category codes and names. Useful for discovering valid category filters for arxiv_search. Lists subject classes only; arxiv_search also accepts a bare archive code (the part before the dot, e.g. "astro-ph" or "cs") to search a whole archive at once.
読み取り専用
入力スキーマ
{'type': 'object', '$schema': 'https://json-schema.org/draft/2020-12/schema', 'properties': {'group': {'enum': ['cs', 'econ', 'eess', 'math', 'physics', 'q-bio', 'q-fin', 'stat'], 'type': 'string', 'description': 'Filter by top-level group (e.g., "cs", "math", "physics"). Returns all categories if omitted.'}}, 'additionalProperties': False}
出力スキーマ
{'type': 'object', 'anyOf': [{'not': {'required': ['error']}, 'required': ['categories', 'totalCount']}, {'required': ['error']}], '$schema': 'https://json-schema.org/draft/2020-12/schema', 'properties': {'error': {'type': 'object', 'required': ['code', 'message'], 'properties': {'code': {'type': 'integer', 'maximum': 9007199254740991, 'minimum': -9007199254740991, 'description': 'JSON-RPC error code for this failure.'}, 'data': {'type': 'object', 'properties': {'reason': {'type': 'string', 'description': 'Machine-readable failure mode.'}, 'recovery': {'type': 'object', 'required': ['hint'], 'properties': {'hint': {'type': 'string'}}, 'description': 'Actionable next step for the caller.', 'additionalProperties': {}}, 'retryable': {'type': 'boolean', 'description': 'Whether retrying may succeed.'}}, 'additionalProperties': {}}, 'message': {'type': 'string', 'description': 'Human-readable description of what went wrong.'}}, 'description': 'Present when the call failed. Absent on success.', 'additionalProperties': {}}, 'notice': {'type': 'string', 'description': 'Guidance when the group filter returns no categories.'}, 'categories': {'type': 'array', 'items': {'type': 'object', 'required': ['code', 'name', 'group'], 'properties': {'code': {'type': 'string', 'description': 'Category code (e.g., "cs.AI").'}, 'name': {'type': 'string', 'description': 'Full name (e.g., "Artificial Intelligence").'}, 'group': {'type': 'string', 'description': 'Top-level group (e.g., "cs").'}}, 'description': 'arXiv category â\x80\x94 subject code, full name, and top-level group.', 'additionalProperties': False}, 'description': 'arXiv categories matching the filter.'}, 'totalCount': {'type': 'number', 'description': 'Total number of categories returned.'}}, 'additionalProperties': False}
arxiv_read_paper
Arxiv Read Paper
Fetch the full text of an arXiv paper. Tries arxiv.org/html first, falls back to ar5iv.labs.arxiv.org, and falls back again to text extracted from the PDF when neither has an HTML render — check the source field to know which one answered. Page through long papers with start and max_characters, or pass max_characters null to get the entire body in one call.
読み取り専用
入力スキーマ
{'type': 'object', '$schema': 'https://json-schema.org/draft/2020-12/schema', 'required': ['paper_id'], 'properties': {'start': {'type': 'integer', 'default': 0, 'maximum': 9007199254740991, 'minimum': 0, 'description': 'Character offset into the cleaned body to begin reading from. Defaults to 0. Use with max_characters to page through long papers â\x80\x94 e.g., start=100000 with max_characters=100000 returns chars 100,000â\x80\x93199,999. The total length is reported as body_characters in the response.'}, 'paper_id': {'type': 'string', 'minLength': 1, 'description': 'arXiv paper ID (e.g., "2401.12345" or "2401.12345v2").'}, 'max_characters': {'anyOf': [{'type': 'integer', 'maximum': 9007199254740991, 'minimum': 1}, {'type': 'null'}], 'default': 100000, 'description': 'Maximum characters of paper body to return, counted after boilerplate stripping. Defaults to 100,000; pass null to return the entire body in one call. Whole-paper reads can exceed a client tool-result size cap â\x80\x94 math-heavy bodies run 300KB-1MB+ â\x80\x94 so prefer the default plus start-based paging unless the full text is needed. When truncated, a notice and the total character count are included.'}}, 'additionalProperties': False}
出力スキーマ
{'type': 'object', 'anyOf': [{'not': {'required': ['error']}, 'required': ['paper_id', 'title', 'content', 'source', 'truncated', 'start', 'total_characters', 'body_characters', 'pdf_url', 'abstract_url']}, {'required': ['error']}], '$schema': 'https://json-schema.org/draft/2020-12/schema', 'properties': {'error': {'type': 'object', 'required': ['code', 'message'], 'properties': {'code': {'type': 'integer', 'maximum': 9007199254740991, 'minimum': -9007199254740991, 'description': 'JSON-RPC error code for this failure.'}, 'data': {'type': 'object', 'properties': {'reason': {'type': 'string', 'examples': ['no_match', 'content_unavailable', 'pdf_extraction_failed', 'version_unavailable', 'rate_limited', 'invalid_request'], 'description': 'Machine-readable failure mode. Declared by this tool: `no_match`: Paper ID is not present in the arXiv index. `content_unavailable`: Paper exists but neither arxiv.org/html nor ar5iv has an HTML rendering and arXiv served no PDF either. `pdf_extraction_failed`: Paper has no HTML rendering and its PDF carries no text layer â\x80\x94 an image-only or scanned submission. `version_unavailable`: A version-pinned paper_id was requested, arXiv is unreachable, and the local mirror holds only a different version â\x80\x94 per-version reads require the live API. `rate_limited`: arXiv has throttled requests (HTTP 429 or "Rate exceeded." body). `invalid_request`: arXiv rejected the metadata lookup (HTTP 4xx other than 429), e.g. malformed ID syntax. Other values are possible when a failure originates below the handler.'}, 'recovery': {'type': 'object', 'required': ['hint'], 'properties': {'hint': {'type': 'string'}}, 'description': 'Actionable next step for the caller.', 'additionalProperties': {}}, 'retryable': {'type': 'boolean', 'description': 'Whether retrying may succeed.'}}, 'additionalProperties': {}}, 'message': {'type': 'string', 'description': 'Human-readable description of what went wrong.'}}, 'description': 'Present when the call failed. Absent on success.', 'additionalProperties': {}}, 'start': {'type': 'number', 'description': 'Character offset of the first character in content within the cleaned body.'}, 'title': {'type': 'string', 'description': 'Paper title (from metadata, not parsed from HTML).'}, 'source': {'enum': ['arxiv_html', 'ar5iv', 'pdf_text'], 'type': 'string', 'description': 'Which upstream artifact the body was read from. arxiv_html and ar5iv are HTML renders; pdf_text is text extracted from the PDF, where prose is reliable but math, tables, and heading structure are flattened.'}, 'content': {'type': 'string', 'description': 'Paper body for the requested slice â\x80\x94 cleaned HTML when source is arxiv_html or ar5iv, plain text when source is pdf_text. Empty when start is past body_characters.'}, 'pdf_url': {'type': 'string', 'description': 'Direct PDF download URL.'}, 'paper_id': {'type': 'string', 'description': 'arXiv paper ID.'}, 'truncated': {'type': 'boolean', 'description': 'True when more body content exists past this slice (start + content.length < body_characters).'}, 'abstract_url': {'type': 'string', 'description': 'arXiv abstract page URL for attribution.'}, 'body_characters': {'type': 'number', 'description': 'Character count of the full cleaned body. Use with start and max_characters to page. Typically 3-4Ã\x97 smaller than total_characters for math-heavy HTML papers.'}, 'total_characters': {'type': 'number', 'description': 'Character count of the body before cleaning â\x80\x94 the unprocessed HTML body for arxiv_html and ar5iv, and equal to body_characters for pdf_text, which needs no cleaning.'}}, 'additionalProperties': False}
arxiv_search
Arxiv Search
Search arXiv papers by query with category and sort filters. Returns paper metadata including title, authors, abstract, categories, and links.
読み取り専用
入力スキーマ
{'type': 'object', '$schema': 'https://json-schema.org/draft/2020-12/schema', 'required': ['query'], 'properties': {'query': {'type': 'string', 'pattern': '^[^\\x00-\\x08\\x0B\\x0C\\x0E-\\x1F]*$', 'maxLength': 1000, 'minLength': 1, 'description': 'Search query. Field prefixes: ti: (title), au: (author â\x80\x94 token-based; quote multi-token names like au:"hinton g" or pair with a topical clause to disambiguate common surnames), abs: (abstract), cat: (category â\x80\x94 a leaf code matches exactly, a bare archive code such as cat:astro-ph matches its whole subtree), co: (comment), jr: (journal ref), all: (all fields). Boolean operators: AND, OR, ANDNOT. Examples: "au:bengio AND ti:attention", "all:transformer AND cat:cs.CL".'}, 'start': {'type': 'integer', 'default': 0, 'maximum': 10000, 'minimum': 0, 'description': 'Pagination offset (0-10000). Use with max_results to page through results. E.g., start=10 with max_results=10 returns results 11-20. Matches beyond offset 10000 + max_results are unreachable by paging â\x80\x94 carve the search into submitted_from/submitted_to windows and page within each.'}, 'sort_by': {'enum': ['relevance', 'submitted', 'updated'], 'type': 'string', 'default': 'relevance', 'description': 'Sort criterion. Use "submitted" for newest papers, "relevance" for best query matches.'}, 'category': {'type': 'string', 'description': 'Restrict results to an arXiv category. A leaf code ("cs.CL", "math.AG") matches exactly. A bare archive code ("astro-ph", "cond-mat", "cs", "math") matches the whole archive â\x80\x94 its subject classes plus the legacy flat papers filed before the archive was subdivided. Note "physics" is the general-physics archive (physics.*), not the wider physics group: astro-ph, cond-mat, hep-*, quant-ph and the rest are separate archive codes. Use arxiv_list_categories to discover subject classes.'}, 'sort_order': {'enum': ['ascending', 'descending'], 'type': 'string', 'default': 'descending', 'description': 'Sort direction. "descending" returns newest/most relevant first.'}, 'max_results': {'type': 'integer', 'default': 10, 'maximum': 50, 'minimum': 1, 'description': 'Maximum results to return (1-50). Default 10. Each result includes title, authors, abstract, and metadata â\x80\x94 keep low to limit response size.'}, 'submitted_to': {'type': 'string', 'pattern': '^(\\d{4}-\\d{2}-\\d{2})?$', 'description': 'Latest submission date to include, inclusive, as a UTC YYYY-MM-DD date. Omit for no upper bound. Both bounds are inclusive, so consecutive windows ("2024-01-01".."2024-01-15" then "2024-01-16".."2024-01-31") cover the matches with no gap; a paper submitted at exactly the midnight seam between two windows appears in both, so de-duplicate collected results by paper id. That is the way to reach matches past the start ceiling: split the date range, then page within each window.'}, 'submitted_from': {'type': 'string', 'pattern': '^(\\d{4}-\\d{2}-\\d{2})?$', 'description': 'Earliest submission date to include, inclusive, as a UTC YYYY-MM-DD date. Omit for no lower bound.'}}, 'additionalProperties': False}
出力スキーマ
{'type': 'object', 'anyOf': [{'not': {'required': ['error']}, 'required': ['papers', 'effectiveQuery', 'totalFound', 'pageStart']}, {'required': ['error']}], '$schema': 'https://json-schema.org/draft/2020-12/schema', 'properties': {'cap': {'type': 'number', 'description': 'The max_results limit applied to this page.'}, 'error': {'type': 'object', 'required': ['code', 'message'], 'properties': {'code': {'type': 'integer', 'maximum': 9007199254740991, 'minimum': -9007199254740991, 'description': 'JSON-RPC error code for this failure.'}, 'data': {'type': 'object', 'properties': {'reason': {'type': 'string', 'examples': ['unknown_category', 'rate_limited', 'invalid_request', 'unsupported_query_syntax', 'invalid_date_range'], 'description': 'Machine-readable failure mode. Declared by this tool: `unknown_category`: Provided category code is not part of the arXiv taxonomy. `rate_limited`: arXiv has throttled requests (HTTP 429 or "Rate exceeded." body). `invalid_request`: arXiv rejected the request (HTTP 4xx other than 429), typically malformed query syntax. `unsupported_query_syntax`: Query translates to a mirror FTS5 expression the search engine cannot parse, typically two operands juxtaposed across a parenthesized group without an explicit operator. `invalid_date_range`: submitted_from or submitted_to is not a real UTC calendar date, or the window starts after it ends. Other values are possible when a failure originates below the handler.'}, 'recovery': {'type': 'object', 'required': ['hint'], 'properties': {'hint': {'type': 'string'}}, 'description': 'Actionable next step for the caller.', 'additionalProperties': {}}, 'retryable': {'type': 'boolean', 'description': 'Whether retrying may succeed.'}}, 'additionalProperties': {}}, 'message': {'type': 'string', 'description': 'Human-readable description of what went wrong.'}}, 'description': 'Present when the call failed. Absent on success.', 'additionalProperties': {}}, 'shown': {'type': 'number', 'description': 'Papers returned on this page.'}, 'notice': {'type': 'string', 'description': 'Recovery guidance when results are empty or paging overshot. Absent on successful pages.'}, 'papers': {'type': 'array', 'items': {'type': 'object', 'required': ['id', 'title', 'authors', 'abstract', 'primary_category', 'categories', 'published', 'updated', 'pdf_url', 'abstract_url'], 'properties': {'id': {'type': 'string', 'description': 'arXiv paper ID (e.g., "2401.12345v1").'}, 'doi': {'type': 'string', 'description': 'DOI if available.'}, 'title': {'type': 'string', 'description': 'Paper title.'}, 'authors': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Author names.'}, 'comment': {'type': 'string', 'description': 'Author comment (e.g., page count, conference).'}, 'pdf_url': {'type': 'string', 'description': 'Direct PDF download URL.'}, 'updated': {'type': 'string', 'description': 'Last update date (ISO 8601).'}, 'abstract': {'type': 'string', 'description': 'Full abstract text.'}, 'published': {'type': 'string', 'description': 'Original submission date (ISO 8601).'}, 'categories': {'type': 'array', 'items': {'type': 'string'}, 'description': 'All arXiv categories assigned to this paper.'}, 'journal_ref': {'type': 'string', 'description': 'Journal reference if published.'}, 'abstract_url': {'type': 'string', 'description': 'arXiv abstract page URL.'}, 'primary_category': {'type': 'string', 'description': 'Primary arXiv category (e.g., "cs.CL").'}}, 'description': 'arXiv paper metadata â\x80\x94 identifier, title, authors, abstract, categories, and links.', 'additionalProperties': False}, 'description': 'Matching papers with full metadata.'}, 'pageStart': {'type': 'number', 'description': 'Pagination offset of this result page.'}, 'truncated': {'type': 'boolean', 'description': 'True when more matching papers exist beyond this page (totalFound > start + shown).'}, 'totalFound': {'type': 'number', 'description': 'Total matching papers reported by arXiv (before pagination).'}, 'effectiveQuery': {'type': 'string', 'description': 'The query as actually searched, carrying every filter applied â\x80\x94 the category subtree and submitted-date window folded into arXiv syntax alongside the supplied terms. Replaying it as `query` with no other filters reproduces this exact result set.'}}, 'additionalProperties': False}
追加
arxiv_list_categories
2026年9月17日12:41
追加
arxiv_read_paper
2026年9月17日12:41
追加
arxiv_get_metadata
2026年9月17日12:41
追加
arxiv_search
2026年9月17日12:41