MCP-Server

Misata Studio: verified synthetic data

studio.misata/misata
Daten & Analytik Entwicklertools Öffentlich und erreichbar MCP 2025-11-25

Was dieses MCP kann

Plans, generates, validates, queries, certifies, and exports verified relational synthetic datasets in multiple formats.

blueprint_guide
Blueprint reference
The reference for the engine's full design language (the `blueprint` argument of generate_dataset, start_generation and validate_blueprint): roles, distributions, formulas that read parent columns and draw noise, per-group sequences (seq, lag, cumsum, ar1 drift, random walk), aggregates, event windows, cause-and-effect, lifecycles, with patterns. Read it before writing a blueprint.
Nur Lesen Idempotent
Eingabeschema
{'type': 'object', 'title': 'blueprint_guideArguments', 'properties': {}}
Ausgabeschema
{'type': 'object', 'title': 'blueprint_guideOutput', 'required': ['result'], 'properties': {'result': {'type': 'string', 'title': 'Result'}}}
cancel_generation
Cancel a generation
Stop a running start_generation job. It stops at the next stage boundary and keeps nothing.
Idempotent
Eingabeschema
{'type': 'object', 'title': 'cancel_generationArguments', 'required': ['job_id'], 'properties': {'job_id': {'type': 'string', 'title': 'Job Id'}}}
export_dataset
Export a dataset
Export a generated dataset as a file. Returns a `download_url` the person can open (or you can fetch, e.g. with curl) for as long as the dataset is held, about 2 hours. Give the person the link rather than pasting file contents into the chat. Args: dataset_id: from a prior generate_dataset call. format: data: csv, parquet, jsonl, json, avro, xlsx, feather, orc, sqlite, duckdb, sql. code and docs: dbt, notebook, dictionary, dbml, mermaid, prisma, sqlalchemy, typescript, jsonschema, expectations, django, openapi, mockapi, demo. `sql` is schema.sql (DDL with keys) + data.sql (COPY/INSERT) — the way to seed a real database: run the returned SQL through your own database connection, since this server never holds a database credential itself. dialect: for `sql` only: postgres, mysql, sqlite, mssql, oracle, bigquery, snowflake. inline: also return the file itself as `base64` (only for files under a few MB). Use it when you must write the file yourself and cannot fetch a URL. Returns: filename, content_type, bytes, download_url, expires_at (unix seconds), and `base64` when `inline` and small enough.
Idempotent
Eingabeschema
{'type': 'object', 'title': 'export_datasetArguments', 'required': ['dataset_id'], 'properties': {'format': {'type': 'string', 'title': 'Format', 'default': 'csv'}, 'inline': {'type': 'boolean', 'title': 'Inline', 'default': False}, 'dialect': {'type': 'string', 'title': 'Dialect', 'default': 'postgres'}, 'dataset_id': {'type': 'string', 'title': 'Dataset Id'}}}
find_ready_dataset
Find a ready-made dataset
The ready-made datasets Misata publishes: free sample databases (direct download, public domain) and premium datasets with a full answer key (a free preview, then a one-off price). Check this first when someone wants sample, demo, practice or teaching data for a common scenario (retail, e-commerce, SaaS, manufacturing SPC, predictive maintenance, insurance claims, fraud/AML, clinical, network security): handing over a dataset that already exists is instant. If none fits their tables, generate one instead. Args: query: what the person needs, in their words. Only orders the list (closest first); every dataset is still returned, so judge the fit yourself from the tables and summary. Returns: datasets: [{slug, kind (free|premium), title, summary, rows, tables, page_url, and download_url (free) or free_preview_url + buy_url + price_usd (premium)}].
Nur Lesen Externer Zugriff Idempotent
Eingabeschema
{'type': 'object', 'title': 'find_ready_datasetArguments', 'properties': {'query': {'type': 'string', 'title': 'Query', 'default': ''}}}
generate_dataset
Generate a verified dataset
Make a verified relational dataset and wait for it: every foreign key checked, dates correctly ordered, declared aggregates and rates exact, before anything is returned. Deterministic for a seed. Use this when you give a `schema` or `ddl` (seconds). For a plain-English `request` the engine designs the tables with a model, which can take several minutes: call start_generation instead and poll get_status, so the call does not sit open and time out. Args: request: Plain-English description (needs an LLM key unless `schema`/`ddl` is also given). schema: A Misata schema dict for exact structural control. No key needed for structure. ddl: CREATE TABLE statements. No key needed for structure. seed: Reproducibility seed — the same request/schema and seed always produce the same rows. research: Ground realistic numbers (prices, growth rates) in real published facts via a web search. Off by default (an anonymous caller's request should not trigger external calls unless asked for); needs a key regardless of `schema`/`ddl`. blueprint: The engine's full design language (blueprint_guide has the reference): readings around a parent's nominal, drift, autocorrelation, cause-and-effect, event windows, exact aggregates. Use it whenever the data must behave like the real process. No key needed; run validate_blueprint on it first. Returns: dataset_id: pass this to get_certificate / query_dataset / export_dataset. passed: whether every check held. A dataset that did not pass is still returned, with `certificate.findings` saying what failed — inspect before trusting it. verification: a short, plain-language account of what was checked and what held — show it to the person as the proof, in place of asking them to take the data on trust. tables: {name: {columns, preview_rows (first 20), total_rows}} — the preview only; every row is in the stored dataset, reachable by query_dataset/export_dataset. certificate: the short form (claims, requirements, findings). get_certificate returns the rest (patterns, realism scorecard, every proof chart).
Externer Zugriff Idempotent
Eingabeschema
{'type': 'object', 'title': 'generate_datasetArguments', 'properties': {'ddl': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'title': 'Ddl', 'default': None}, 'seed': {'type': 'integer', 'title': 'Seed', 'default': 42}, 'schema': {'anyOf': [{'type': 'object', 'additionalProperties': True}, {'type': 'null'}], 'title': 'Schema', 'default': None}, 'request': {'type': 'string', 'title': 'Request', 'default': ''}, 'research': {'type': 'boolean', 'title': 'Research', 'default': False}, 'blueprint': {'anyOf': [{'type': 'object', 'additionalProperties': True}, {'type': 'null'}], 'title': 'Blueprint', 'default': None}}}
get_certificate
Dataset certificate
The full certificate for a dataset made by generate_dataset: every claim (stated vs actual), every requirement's status and evidence, realism findings, planted defects/anomalies if any, and the pattern charts. This is the whole answer key, not the trimmed form generate_dataset returns.
Nur Lesen Idempotent
Eingabeschema
{'type': 'object', 'title': 'get_certificateArguments', 'required': ['dataset_id'], 'properties': {'dataset_id': {'type': 'string', 'title': 'Dataset Id'}}}
get_status
Generation status
Where a start_generation job is. While running: the stage it is in and how long it has run. When done: the same answer generate_dataset returns (dataset_id, verification, tables, certificate).
Nur Lesen Idempotent
Eingabeschema
{'type': 'object', 'title': 'get_statusArguments', 'required': ['job_id'], 'properties': {'job_id': {'type': 'string', 'title': 'Job Id'}}}
plan_dataset
Plan a dataset
See the tables, sizes and relationships the engine would build, before any rows exist. Free (no rows are made), so it is worth calling before generate_dataset on anything non-trivial: review what it understood and assumed, then adjust your request or schema before spending a real call. Args: request: Plain-English description (needs an LLM key, unless `schema`/`ddl` is also given — then it is still used to ground realism, e.g. locale and what columns mean). schema: A Misata schema dict (see the server instructions for the format). No key needed for structure. ddl: CREATE TABLE statements. No key needed for structure. Returns: route: "chat" (not a dataset request — see `reply`), "design" (the model designed the tables) or "pack" (matched a built-in shape). tables: name, estimated rows, columns, foreign keys for each table the engine would build. understanding: what the engine read the request as (business, archetype, assumptions). requirements: every specific thing the request asked for, so you can see what was understood.
Nur Lesen Idempotent
Eingabeschema
{'type': 'object', 'title': 'plan_datasetArguments', 'properties': {'ddl': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'title': 'Ddl', 'default': None}, 'schema': {'anyOf': [{'type': 'object', 'additionalProperties': True}, {'type': 'null'}], 'title': 'Schema', 'default': None}, 'request': {'type': 'string', 'title': 'Request', 'default': ''}}}
query_dataset
Query a dataset with SQL
Read-only SQL over a dataset's own tables (one SELECT/WITH statement; every table name is a view over that dataset's own files — nothing else on the server is reachable this way). Args: dataset_id: from a prior generate_dataset call. sql: a single SELECT (or WITH ... SELECT) statement. No semicolons, file paths, or statements that write or read outside the dataset (checked before running). limit: rows returned (capped at 5000). Returns: columns, rows, truncated (whether more rows existed than `limit`).
Nur Lesen Idempotent
Eingabeschema
{'type': 'object', 'title': 'query_datasetArguments', 'required': ['dataset_id', 'sql'], 'properties': {'sql': {'type': 'string', 'title': 'Sql'}, 'limit': {'type': 'integer', 'title': 'Limit', 'default': 1000}, 'dataset_id': {'type': 'string', 'title': 'Dataset Id'}}}
start_generation
Start a generation (background)
Start a generation in the background and return a job_id at once. Use it for any plain-English `request` (a model designs the tables, which takes minutes) — then call get_status(job_id) every 20-30 seconds until it says done. Arguments are the same as generate_dataset. Returns: job_id, status "running". get_status gives the stage and, when finished, the dataset.
Externer Zugriff Idempotent
Eingabeschema
{'type': 'object', 'title': 'start_generationArguments', 'properties': {'ddl': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'title': 'Ddl', 'default': None}, 'seed': {'type': 'integer', 'title': 'Seed', 'default': 42}, 'schema': {'anyOf': [{'type': 'object', 'additionalProperties': True}, {'type': 'null'}], 'title': 'Schema', 'default': None}, 'request': {'type': 'string', 'title': 'Request', 'default': ''}, 'research': {'type': 'boolean', 'title': 'Research', 'default': False}, 'blueprint': {'anyOf': [{'type': 'object', 'additionalProperties': True}, {'type': 'null'}], 'title': 'Blueprint', 'default': None}}}
validate_blueprint
Validate a blueprint
Check a blueprint before generating it: every error said as what to change, the design traps it falls into (a column that will round away, a copy of a value not made yet, a window that counts nothing), its size against your row cap, and a small preview run (a few thousand rows) with the verifier's findings and sample rows, so you can see the data behave before the real call. Free. Returns: valid, errors (fix these), warnings (read these), estimated_rows, row_cap, preview: {passed, findings, tables: {name: first rows}} when `preview`.
Nur Lesen Idempotent
Eingabeschema
{'type': 'object', 'title': 'validate_blueprintArguments', 'required': ['blueprint'], 'properties': {'preview': {'type': 'boolean', 'title': 'Preview', 'default': True}, 'blueprint': {'type': 'object', 'title': 'Blueprint', 'additionalProperties': True}}}
whoami
Who am I connected as
Which account this connection is, and what it may do: signed in with a Studio MCP key or anonymous, the row cap, and whether a model key is available for plain-English requests.
Nur Lesen Idempotent
Eingabeschema
{'type': 'object', 'title': 'whoamiArguments', 'properties': {}}
Hinzugefügt
find_ready_dataset
2. October 2026 02:40
Hinzugefügt
export_dataset
2. October 2026 02:40
Hinzugefügt
query_dataset
2. October 2026 02:40
Hinzugefügt
get_certificate
2. October 2026 02:40
Hinzugefügt
cancel_generation
2. October 2026 02:40
Hinzugefügt
get_status
2. October 2026 02:40
Hinzugefügt
start_generation
2. October 2026 02:40
Hinzugefügt
generate_dataset
2. October 2026 02:40
Hinzugefügt
validate_blueprint
2. October 2026 02:40
Hinzugefügt
blueprint_guide
2. October 2026 02:40
Hinzugefügt
plan_dataset
2. October 2026 02:40
Hinzugefügt
whoami
2. October 2026 02:40