MCP-Server

Web Content Extract Mcp

io.github.varvararatta/web_content_extract_mcp
Medien & Inhalte Suche & Recherche Öffentlich und erreichbar MCP 2025-11-25

Was dieses MCP kann

Fetches public web pages, extracts article and page text, returns metadata, and enumerates page links.

extract_article
Extract main article content from a news/blog URL. Returns: {title, description, body, author, date}
Eingabeschema
{'type': 'object', 'title': 'extract_articleArguments', 'required': ['url'], 'properties': {'url': {'type': 'string', 'title': 'Url'}}}
fetch_url_content
Fetch and extract clean text content from a public URL. Returns: {title, text, url, word_count}
Eingabeschema
{'type': 'object', 'title': 'fetch_url_contentArguments', 'required': ['url'], 'properties': {'url': {'type': 'string', 'title': 'Url'}, 'max_chars': {'type': 'integer', 'title': 'Max Chars', 'default': 5000}}}
get_page_links
Extract all links from a page. Returns: {links: [href], internal_count, external_count}
Eingabeschema
{'type': 'object', 'title': 'get_page_linksArguments', 'required': ['url'], 'properties': {'url': {'type': 'string', 'title': 'Url'}, 'same_domain_only': {'type': 'boolean', 'title': 'Same Domain Only', 'default': True}}}
get_page_metadata
Get metadata from a page: title, description, og tags, keywords. Returns: {title, description, keywords, og_title, og_image, og_type}
Eingabeschema
{'type': 'object', 'title': 'get_page_metadataArguments', 'required': ['url'], 'properties': {'url': {'type': 'string', 'title': 'Url'}}}
health_check
Server health check.
Eingabeschema
{'type': 'object', 'title': 'health_checkArguments', 'properties': {}}
Hinzugefügt
health_check
17. September 2026 12:53
Hinzugefügt
get_page_metadata
17. September 2026 12:53
Hinzugefügt
get_page_links
17. September 2026 12:53
Hinzugefügt
extract_article
17. September 2026 12:53
Hinzugefügt
fetch_url_content
17. September 2026 12:53