AI Crawler Index
What this MCP does
Identifies web crawlers, verifies crawler IP addresses, provides crawler metadata, and generates robots.txt guidance.
Tools
Input schema
{'type': 'object', 'examples': [{'since': '0'}], 'properties': {'limit': {'type': 'integer', 'maximum': 400, 'minimum': 1, 'description': 'Events per page, default 100.'}, 'since': {'type': 'string', 'description': 'Cursor from the last result, or an ISO date. Omit for all retained.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'examples': [{'user_agent': 'GPTBot/1.2'}], 'required': ['user_agent'], 'properties': {'user_agent': {'type': 'string', 'description': 'Raw User-Agent header value.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'examples': [{}], 'required': [], 'properties': {}, 'additionalProperties': False}
Output schema
{'type': 'object', 'required': ['ran', 'input_came_from', 'what_it_shows', 'answer', 'reproduce', 'this_is_not_a_mock', 'answered_by', 'license'], 'properties': {'ran': {'type': 'object', 'description': 'The tool name and the exact arguments that were run.'}, 'answer': {'type': 'object', 'description': 'The real structuredContent of that call, not a mock.'}, 'license': {'type': 'string'}, 'reproduce': {'type': 'string', 'description': 'A command that reproduces this answer.'}, 'answered_by': {'type': 'object'}, 'what_it_shows': {'type': 'string'}, 'input_came_from': {'type': 'string', 'description': "Where the canned input came from — always this host's own data."}, 'this_is_not_a_mock': {'type': 'string'}}, 'description': "This server's own worked example, executed for real on a canned input from this host's own data.", 'additionalProperties': True}
Input schema
{'type': 'object', 'examples': [{'stance': 'block-ai-training'}], 'properties': {'stance': {'type': 'string', 'description': 'One of allow-all, block-ai-training, block-all-ai, block-datasets, allow-ai-search-only, block-seo-tools, block-disputed, maximum-ai-visibility.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'examples': [{'ip': '20.171.206.5'}], 'required': ['ip'], 'properties': {'ip': {'type': 'string', 'description': 'IPv4 or IPv6 address to check.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'examples': [{'limit': 20, 'category': 'ai-training'}], 'properties': {'q': {'type': 'string', 'description': 'Free text over slug, name and operator.'}, 'limit': {'type': 'integer', 'maximum': 200, 'minimum': 1, 'description': 'Max rows, 1-200 (default 100).'}, 'category': {'type': 'string', 'description': 'ai-training, ai-search, user-fetch, dataset, search, seo, archive, preview or tool.'}, 'operator': {'type': 'string', 'description': 'Operator slug or name.'}, 'verification': {'type': 'string', 'description': 'published-ranges, reverse-dns or none.'}, 'respects_robots_txt': {'type': 'string', 'description': 'documented, disputed, by-design-no or n-a.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'examples': [{'slug': 'claudebot'}], 'required': ['slug'], 'properties': {'slug': {'type': 'string', 'description': 'Crawler slug, name or robots.txt token.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'examples': [{'since': '2026-08-01'}], 'properties': {'since': {'type': 'string', 'description': 'ISO-8601 date or timestamp; omit for the current state.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'examples': [{}], 'required': [], 'properties': {}, 'additionalProperties': False}
Output schema
{'type': 'object', 'required': ['you', 'we_book_you_as', 'we_have_seen_you', 'answered_by', 'this_call_touched', 'caveats', 'license', 'independent'], 'properties': {'you': {'type': 'object', 'description': 'The user-agent you sent and the address you came from.'}, 'caveats': {'type': 'array', 'items': {'type': 'string'}, 'description': 'What this answer does NOT establish — a user-agent is a claim.'}, 'license': {'type': 'string'}, 'answered_by': {'type': 'object', 'description': 'Which server answered, at which endpoint.'}, 'independent': {'type': 'boolean', 'description': 'This host is independent and unaffiliated.'}, 'we_book_you_as': {'type': 'object', 'description': "The class this host's own instrument records for that user-agent."}, 'we_have_seen_you': {'type': 'object', 'description': 'Whether this user-agent appears in the published observation window.'}, 'this_call_touched': {'type': 'object', 'description': 'Exactly which files were read to answer. No third party is contacted.'}}, 'description': 'Facts about the caller, derived only from the headers of this request and from files this host already publishes.', 'additionalProperties': True}
Recent tool changes
Similar MCP servers
osint-terminal
Provides keyless OSINT and reconnaissance tools for domains, DNS, IPs, breach exposure, threat intelligence, and related lookups.
AIMEAT
Provides a self-hosted agent operating system with agent work delegation, access controls, federation, hooks, SSO, security admin…
hyperion
Acts as a paid MCP tool marketplace and utility gateway with server discovery, HTTP and JavaScript tools, research, data conversi…
Vee3
Manages Clerk authentication infrastructure, including users, organizations, domains, sessions, tokens, OAuth, SSO, machines, per…
BorealHost
Provides web hosting and infrastructure management, including site deployment, DNS, domains, containers, compute, backups, cachin…
Proof Holdings
Provides domain verification, identity and delegation proofs, human approval workflows, trusted-contact challenges, and controlle…
GoCreative Agent API
Offers pay-per-call LLM completions and data services for company intelligence, KYB, sanctions screening, threat intelligence, co…
Japan Public Ledgers MCP
Provides agent identity, memory, audit, trust, proxy, temporary email, webhook, CAPTCHA, and alerting capabilities alongside publ…