MCP 서버

smart-data-extractor

io.github.lazymac2x/smart-data-extractor
데이터 및 분석 개발자 도구 공개 · 연결 가능 MCP 2026-07-28

이 MCP로 할 수 있는 일

Extracts structured data from URLs, APIs, text, JSON, and multiple sources using inferred or supplied schemas.

auto_schema_learn
Idempotent · 30s timeout · Automatically infer JSON Schema from sample data without extraction. Pass `idempotency_key` to deduplicate within 5 minutes.
입력 스키마
{'type': 'object', 'required': ['sample_data'], 'properties': {'sample_data': {'type': 'string', 'examples': ['{"product_id":123,"title":"Laptop Pro","price":1299.99,"in_stock":true}', '[{"id":1,"email":"alice@example.com","verified":true},{"id":2,"email":"bob@example.com","verified":false}]', '{"name":"Report Q1","sections":[{"title":"Sales","value":50000}]}'], 'description': 'Representative sample data as JSON string (array of objects or single object). Schema is inferred from structure; use first 1-10 rows for array samples. Max 200KB.'}, 'idempotency_key': {'type': 'string', 'description': 'Optional cache key (UUID/string) for 5-minute deduplication. Repeat calls with same key return cached inferred schema instantly.'}}, 'additionalProperties': False}
batch_extract
Idempotent · 30s timeout · Extract data from multiple sources (JSON/JSONL/text) with a single consistent schema. Pass `idempotency_key` to deduplicate within 5 minutes.
입력 스키마
{'type': 'object', 'required': ['sources'], 'properties': {'schema': {'type': 'object', 'description': 'Optional target JSON Schema (draft-07) applied to all sources. If omitted, inferred from first source and reused across remaining sources. Enables consistent field extraction from diverse formats.'}, 'sources': {'type': 'array', 'items': {'type': 'object', 'required': ['type', 'content'], 'properties': {'type': {'enum': ['json', 'jsonl', 'text'], 'type': 'string', 'description': 'Source format: "json" (single object or array), "jsonl" (newline-delimited JSON objects), "text" (key:value pairs delimited by commas/semicolons)'}, 'content': {'type': 'string', 'description': 'Raw source content (max 200KB per source). For jsonl: each line must be valid JSON. For text: format "key1:value1,key2:value2".'}}, 'additionalProperties': False}, 'maxItems': 100, 'minItems': 1, 'description': 'Array of 1-100 data sources to extract from. Each source includes type (format) and content (raw data).'}, 'idempotency_key': {'type': 'string', 'description': 'Optional deduplication key (UUID/string) for 5-minute cache. Identical batch calls (same sources + schema + key) return cached results instantly.'}}, 'additionalProperties': False}
extract_from_api
Idempotent · 30s timeout · Extract structured data from API response JSON with schema adaptation. Pass `idempotency_key` to deduplicate within 5 minutes.
입력 스키마
{'type': 'object', 'required': ['content'], 'properties': {'schema': {'type': 'object', 'description': 'Optional target JSON Schema (draft-07) for field extraction. If omitted, inferred from content structure. Enforces consistent field extraction across multiple API responses.'}, 'content': {'type': 'string', 'examples': ['{"id":1,"name":"Alice","role":"admin"}', '[{"id":1,"name":"Alice"},{"id":2,"name":"Bob"}]', '[{"user":{"id":1,"name":"Alice"},"status":"active"}]'], 'description': 'API response body as raw JSON string (max 200KB). Can be single object, array of objects, or array of primitives. Automatically parsed and validated.'}, 'idempotency_key': {'type': 'string', 'description': 'Optional deduplication key (UUID or unique string) for 5-minute cache. Identical calls return cached result instantly.'}}, 'additionalProperties': False}
extract_from_url
Idempotent · 30s timeout · Extract structured data from URL content with auto schema learning. Pass `idempotency_key` to deduplicate identical calls within 5 minutes.
입력 스키마
{'type': 'object', 'required': ['url'], 'properties': {'url': {'type': 'string', 'examples': ['https://api.github.com/users/octocat', 'https://example.com/page.html', '{"name":"John Doe","age":30,"email":"john@example.com"}'], 'description': 'HTTP(S) URL to fetch (will auto-download and parse), or raw content string (up to 200KB). Max 200KB after fetch.'}, 'schema': {'type': 'object', 'description': 'Optional pre-defined JSON Schema (draft-07). If omitted, schema is auto-inferred from content. Provide to enforce strict field extraction and type coercion.'}, 'idempotency_key': {'type': 'string', 'description': 'Optional UUID or unique identifier for 5-minute deduplication cache. Same key + tool = cached result in <5ms, zero re-fetching.'}}, 'additionalProperties': False}
추가됨
batch_extract
2026년 9월 17일 12:43 PM
추가됨
auto_schema_learn
2026년 9월 17일 12:43 PM
추가됨
extract_from_api
2026년 9월 17일 12:43 PM
추가됨
extract_from_url
2026년 9월 17일 12:43 PM