mcp-server
What this MCP does
Onboards data sources, generates processing specifications, runs data jobs and SQL queries, and manages SFTP or S3 integrations and schedules.
Tools
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['path', 'action'], 'properties': {'path': {'type': 'string', 'description': 'API path, e.g. "/auth/billing" (leading slash, no query string).'}, 'action': {'type': 'string', 'description': 'The "action" field this endpoint routes on, e.g. "get-balance".'}, 'params': {'type': 'object', 'description': 'Additional action-specific fields to merge into the request body alongside action/workspaceId.', 'additionalProperties': {}}, 'workspaceId': {'type': 'string', 'description': 'Include for workspace-scoped actions. Omit entirely for account-level actions.'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['reason', 'message', 'email'], 'properties': {'name': {'type': 'string'}, 'email': {'type': 'string', 'format': 'email', 'description': "The sender's email address, so DPF can reply."}, 'reason': {'enum': ['request-a-demo', 'understand-licensing', 'report-an-issue', 'request-a-feature'], 'type': 'string'}, 'message': {'type': 'string', 'maxLength': 5000, 'minLength': 1}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['name'], 'properties': {'name': {'type': 'string', 'description': 'Workspace name'}, 'description': {'type': 'string', 'description': 'Optional workspace description.'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['workspaceId', 'name', 'createdBy', 'createdAt', 'userPermission', 'permission', 'permissions'], 'properties': {'name': {'type': 'string'}, 'createdAt': {'type': 'string'}, 'createdBy': {'type': 'string'}, 'permission': {'type': 'string'}, 'description': {'type': 'string'}, 'permissions': {'type': 'array', 'items': {'type': 'object', 'required': ['userId', 'email', 'permission'], 'properties': {'email': {'type': 'string'}, 'userId': {'type': 'string'}, 'lastName': {'type': 'string'}, 'firstName': {'type': 'string'}, 'permission': {'type': 'string'}}, 'additionalProperties': False}}, 'workspaceId': {'type': 'string'}, 'userPermission': {'type': 'string'}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['specName'], 'properties': {'specName': {'type': 'string', 'description': 'Name of the data spec to delete.'}, 'workspaceId': {'type': 'string', 'description': 'Workspace to act on. Defaults to your only workspace if you have exactly one.'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['specId', 'specName', 'message'], 'properties': {'errors': {'type': 'array', 'items': {'type': 'object', 'required': ['error'], 'properties': {'error': {'type': 'string'}, 'jobId': {'type': 'string'}, 'phase': {'type': 'string'}, 'tableName': {'type': 'string'}}, 'additionalProperties': False}}, 'specId': {'type': 'string'}, 'message': {'type': 'string'}, 'specName': {'type': 'string'}, 'retainedJobs': {'type': 'number'}, 'deletedS3Data': {'type': 'number'}, 'alreadyDeleted': {'type': 'boolean'}, 'deletedS3Specs': {'type': 'number'}, 'deletedTargetTables': {'type': 'number'}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['jobId', 'specName'], 'properties': {'jobId': {'type': 'string', 'description': 'jobId returned by run_data_job.'}, 'specName': {'type': 'string', 'description': 'Name of the data spec this job belongs to.'}, 'workspaceId': {'type': 'string', 'description': 'Workspace to act on. Defaults to your only workspace if you have exactly one.'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['message'], 'properties': {'error': {'type': 'string'}, 'jobId': {'type': 'string'}, 'specId': {'type': 'string'}, 'status': {'type': 'string', 'description': '"processing" | "completed" | "failed"'}, 'endTime': {'type': 'string'}, 'jobSize': {'type': 'string'}, 'message': {'type': 'string'}, 'metrics': {'type': 'object', 'additionalProperties': {}}, 'progress': {'type': 'number'}, 'timedOut': {'type': 'boolean'}, 'startTime': {'type': 'string'}, 'durationMs': {'type': 'number'}, 'logLocation': {'type': 'string'}, 'workspaceId': {'type': 'string'}, 'errorDetails': {'type': 'string'}, 'currentLambda': {'type': 'string'}, 'statusMessage': {'type': 'string'}, 'creditsCharged': {'type': 'number'}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['specId', 'specName'], 'properties': {'specId': {'type': 'string', 'description': 'specId returned by onboard_data_source.'}, 'specName': {'type': 'string', 'description': 'Name of the data spec being onboarded.'}, 'workspaceId': {'type': 'string', 'description': 'Workspace to act on. Defaults to your only workspace if you have exactly one.'}, 'loadSampleData': {'type': 'boolean', 'description': 'Whether to load the sample file and trigger the data-load job once analysis finishes (default true).'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['message'], 'properties': {'error': {'type': ['string', 'null']}, 'specId': {'type': 'string'}, 'status': {'type': 'string', 'description': '"processing" | "ready" | "failed"'}, 'message': {'type': 'string'}, 'progress': {'type': 'number'}, 'timedOut': {'type': 'boolean'}, 'lastJobId': {'type': ['string', 'null']}, 'workspaceId': {'type': 'string'}, 'errorDetails': {'type': ['string', 'null']}, 'statusMessage': {'type': 'string'}, 'analysisEndTime': {'type': ['string', 'null']}, 'analysisStartTime': {'type': ['string', 'null']}, 'analysisDurationMs': {'type': ['number', 'null']}, 'hasTransformationConfig': {'type': 'boolean'}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['specId', 'specName'], 'properties': {'specId': {'type': 'string', 'description': 'specId returned by update_data_spec.'}, 'specName': {'type': 'string', 'description': 'Name of the data spec being updated.'}, 'runAnalysis': {'type': 'boolean', 'description': 'Default true — set false to skip analysis and just confirm the upload.'}, 'workspaceId': {'type': 'string', 'description': 'Workspace to act on. Defaults to your only workspace if you have exactly one.'}, 'loadSampleData': {'type': 'boolean', 'description': 'Whether analysis should also trigger the data-load job (default true).'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['message'], 'properties': {'error': {'type': ['string', 'null']}, 'specId': {'type': 'string'}, 'status': {'type': 'string', 'description': '"processing" | "ready" | "failed"'}, 'message': {'type': 'string'}, 'progress': {'type': 'number'}, 'timedOut': {'type': 'boolean'}, 'lastJobId': {'type': ['string', 'null']}, 'workspaceId': {'type': 'string'}, 'errorDetails': {'type': ['string', 'null']}, 'statusMessage': {'type': 'string'}, 'analysisEndTime': {'type': ['string', 'null']}, 'analysisStartTime': {'type': ['string', 'null']}, 'analysisDurationMs': {'type': ['number', 'null']}, 'hasTransformationConfig': {'type': 'boolean'}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'jobId': {'type': 'string', 'description': "Poll a data-load job's status. Pass exactly one of specId or jobId."}, 'specId': {'type': 'string', 'description': "Poll a data spec's analysis status. Pass exactly one of specId or jobId."}, 'workspaceId': {'type': 'string', 'description': 'Workspace to act on. Defaults to your only workspace if you have exactly one.'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'error': {'type': ['string', 'null']}, 'jobId': {'type': 'string'}, 'specId': {'type': 'string'}, 'status': {'type': 'string', 'description': 'e.g. "processing" | "ready" | "failed" for a spec; "processing" | "completed" | "failed" for a job'}, 'endTime': {'type': 'string'}, 'jobSize': {'type': 'string'}, 'metrics': {'type': 'object', 'description': 'Present once a job completes: recordsRead, recordsWritten, filesProcessed, etc.', 'additionalProperties': {}}, 'progress': {'type': 'number'}, 'lastJobId': {'type': ['string', 'null'], 'description': 'spec poll only: the data-load job triggered once analysis reaches "ready".'}, 'startTime': {'type': 'string'}, 'durationMs': {'type': 'number'}, 'logLocation': {'type': 'string'}, 'workspaceId': {'type': 'string'}, 'errorDetails': {'type': ['string', 'null']}, 'currentLambda': {'type': 'string'}, 'statusMessage': {'type': 'string'}, 'creditsCharged': {'type': 'number'}, 'analysisEndTime': {'type': ['string', 'null']}, 'analysisStartTime': {'type': ['string', 'null']}, 'analysisDurationMs': {'type': ['number', 'null']}, 'hasTransformationConfig': {'type': 'boolean'}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['resource'], 'properties': {'cursor': {'type': 'string', 'description': 'Opaque `nextCursor` from a prior page (omit for the first page).'}, 'pageSize': {'type': 'integer', 'maximum': 100, 'minimum': 1, 'description': 'Records per page (default 25).'}, 'resource': {'enum': ['specs', 'jobs'], 'type': 'string', 'description': 'Which kind of resource to list'}, 'workspaceId': {'type': 'string', 'description': 'Workspace to act on. Defaults to your only workspace if you have exactly one.'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['count'], 'properties': {'jobs': {'type': 'array', 'items': {'type': 'object', 'required': ['jobId', 'specId'], 'properties': {'error': {'type': 'string'}, 'jobId': {'type': 'string'}, 'specId': {'type': 'string'}, 'status': {'type': 'string'}, 'endTime': {'type': 'string'}, 'jobSize': {'type': 'string'}, 'progress': {'type': 'number'}, 'specName': {'type': 'string'}, 'createdAt': {'type': 'string'}, 'createdBy': {'type': 'string'}, 'startTime': {'type': 'string'}, 'durationMs': {'type': 'number'}, 'workspaceId': {'type': 'string'}, 'errorDetails': {'type': 'string'}, 'statusMessage': {'type': 'string'}, 'creditsCharged': {'type': 'number'}}, 'additionalProperties': False}, 'description': 'Present when resource is "jobs"'}, 'count': {'type': 'number', 'description': 'Number of items in this page, not the workspace total'}, 'specs': {'type': 'array', 'items': {'type': 'object', 'required': ['specId', 'specName'], 'properties': {'error': {'type': ['string', 'null']}, 'merge': {'type': 'boolean'}, 'specId': {'type': 'string'}, 'status': {'type': 'string'}, 'progress': {'type': 'number'}, 'specName': {'type': 'string'}, 'createdAt': {'type': 'string'}, 'createdBy': {'type': 'string'}, 'lastJobId': {'type': ['string', 'null']}, 'updatedAt': {'type': 'string'}, 'computeSize': {'enum': ['small', 'large'], 'type': 'string'}, 'description': {'type': 'string'}, 'errorDetails': {'type': ['string', 'null']}, 'targetOption': {'enum': ['existing-tables', 'target-schema-file', 'auto-infer'], 'type': 'string'}, 'targetTables': {'type': 'array', 'items': {'type': 'string'}}, 'statusMessage': {'type': 'string'}, 'targetSchemaFileName': {'type': 'string'}, 'hasTransformationConfig': {'type': 'boolean'}}, 'additionalProperties': False}, 'description': 'Present when resource is "specs"'}, 'pageSize': {'type': 'number'}, 'nextCursor': {'type': ['string', 'null'], 'description': 'Opaque cursor for the next page; null when exhausted'}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {}}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['workspaces', 'count'], 'properties': {'count': {'type': 'number'}, 'workspaces': {'type': 'array', 'items': {'type': 'object', 'required': ['workspaceId', 'name', 'createdBy', 'createdAt', 'userPermission', 'permissions'], 'properties': {'name': {'type': 'string'}, 'createdAt': {'type': 'string'}, 'createdBy': {'type': 'string'}, 'description': {'type': 'string'}, 'permissions': {'type': 'array', 'items': {'type': 'object', 'required': ['userId', 'email', 'permission'], 'properties': {'email': {'type': 'string'}, 'userId': {'type': 'string'}, 'lastName': {'type': 'string'}, 'firstName': {'type': 'string'}, 'permission': {'type': 'string'}}, 'additionalProperties': False}}, 'workspaceId': {'type': 'string'}, 'userPermission': {'type': 'string'}}, 'additionalProperties': False}}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['action', 'email'], 'properties': {'otp': {'type': 'string', 'pattern': '^\\d{6}$', 'description': 'action "verify" and "reset-password" only. The 6-digit code from the email DPF sent.'}, 'email': {'type': 'string', 'format': 'email'}, 'action': {'enum': ['register', 'verify', 'resend', 'forgot-password', 'reset-password'], 'type': 'string'}, 'lastName': {'type': 'string', 'description': 'action "register" only'}, 'firstName': {'type': 'string', 'description': 'action "register" only'}, 'termsAccepted': {'type': 'boolean', 'description': 'action "register" only. Must be true.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['action'], 'properties': {'type': {'enum': ['sftp', 'aws_s3'], 'type': 'string', 'description': 'Connection type. Required for create; defaults to "sftp".'}, 'action': {'enum': ['create', 'list', 'test', 'delete'], 'type': 'string', 'description': 'Which operation to perform.'}, 'roleArn': {'type': 'string', 'description': 'aws_s3 only. The IAM role the customer will create/update. Required for create.'}, 'hostname': {'type': 'string', 'description': 'sftp only. Remote server hostname. Required for create.'}, 'username': {'type': 'string', 'description': 'sftp only. Remote username. Optional for create; defaults to "sftpuser".'}, 'workspaceId': {'type': 'string', 'description': 'Workspace to act on. Defaults to your only workspace if you have exactly one.'}, 'connectionId': {'type': 'string', 'description': 'Existing connection to test or delete. Required for test/delete.'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'type': {'enum': ['sftp', 'aws_s3'], 'type': 'string'}, 'count': {'type': 'number', 'description': 'action "list" only'}, 'account': {'type': 'string', 'description': 'action "test", type "aws_s3" only'}, 'message': {'type': 'string'}, 'roleArn': {'type': 'string'}, 'success': {'type': 'boolean', 'description': 'action "test" only'}, 'hostname': {'type': 'string'}, 'username': {'type': 'string'}, 'createdAt': {'type': 'number'}, 'createdBy': {'type': 'string'}, 'fileCount': {'type': 'number', 'description': 'action "test", type "sftp" only'}, 'publicKey': {'type': 'string'}, 'updatedAt': {'type': 'number'}, 'entryCount': {'type': 'number', 'description': 'action "test", type "sftp" only'}, 'externalId': {'type': 'string'}, 'connections': {'type': 'array', 'items': {'type': 'object', 'additionalProperties': {}}, 'description': 'action "list" only'}, 'trustPolicy': {'type': 'object', 'additionalProperties': {}}, 'workspaceId': {'type': 'string'}, 'connectionId': {'type': 'string'}, 'lastTestedAt': {'type': 'number'}, 'assumedRoleArn': {'type': 'string', 'description': 'action "test", type "aws_s3" only'}, 'lastTestStatus': {'enum': ['passed', 'failed'], 'type': 'string'}, 'dpfPrincipalArn': {'type': 'string'}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['action'], 'properties': {'type': {'enum': ['sftp', 'aws_s3', 'spec_success', 'schedule'], 'type': 'string', 'description': 'Trigger type. Optional for create (defaults to "sftp"). For sftp/aws_s3 must match the connection\'s type.'}, 'action': {'enum': ['create', 'list', 'update', 'delete', 'run-now', 'clear-processed-files', 'run-history'], 'type': 'string', 'description': 'Which operation to perform.'}, 'cursor': {'type': 'string', 'description': 'run-history: opaque `nextCursor` from a prior page (omit for the first page).'}, 'dedupe': {'type': 'boolean', 'description': "sftp/aws_s3 only. Required for create — ask the user rather than assuming a value; do not default it silently. Whether repeat pulls should skip files already loaded into this spec, matched by file name. Has real consequences: with dedupe true, a file that reappears under the same name (e.g. re-uploaded with corrected data) will be silently skipped; with dedupe false, an unchanged file left on the server will be reloaded every run. Omit only for update, where omitting leaves the trigger's existing setting unchanged."}, 'specId': {'type': 'string', 'description': 'run-history: filter to runs of triggers feeding this spec.'}, 'enabled': {'type': 'boolean', 'description': 'Whether the trigger is active. Defaults to true on create.'}, 'endTime': {'type': 'string', 'description': 'run-history: ISO 8601 upper bound (inclusive) on when the run started.'}, 'pageSize': {'type': 'integer', 'maximum': 100, 'minimum': 1, 'description': 'run-history: records per page (default 25).'}, 'preRules': {'type': 'string', 'description': 'sftp/aws_s3 only. Natural language: which files to pick up (e.g. "only *.csv under /outbound").'}, 's3Bucket': {'type': 'string', 'description': 'aws_s3 only. Bucket to poll. Required for create when type is "aws_s3", or to change it on update. Each run lists at most 5000 objects from the bucket/prefix (oldest key first) — past that, new files can be missed. On create, a successful response includes a `warnings` array with this note; relay it to the user and suggest an S3 lifecycle rule to expire/transition old objects.'}, 's3Prefix': {'type': 'string', 'description': 'aws_s3 only. Optional key prefix; defaults to the whole bucket.'}, 'specName': {'type': 'string', 'description': 'The spec this trigger fires. Required for create.'}, 'frequency': {'type': 'object', 'required': ['unit'], 'properties': {'unit': {'enum': ['hourly', 'daily', 'monthly'], 'type': 'string', 'description': 'Schedule cadence.'}, 'hourOfDay': {'type': 'integer', 'maximum': 23, 'minimum': 0, 'description': 'Required for daily/monthly (UTC).'}, 'dayOfMonth': {'type': 'integer', 'maximum': 31, 'minimum': 1, 'description': 'Required for monthly.'}}, 'description': 'Required for create when type is "sftp", "aws_s3", or "schedule"; optional on update to change the schedule. Not applicable to spec_success.', 'additionalProperties': False}, 'postRules': {'type': 'string', 'description': 'sftp/aws_s3 only. Natural language: what to do after a file loads (e.g. "rename with .done suffix").'}, 'startTime': {'type': 'string', 'description': 'run-history: ISO 8601 lower bound (inclusive) on when the run started.'}, 'triggerId': {'type': 'string', 'description': 'Existing trigger. Required for update/delete/run-now/clear-processed-files.'}, 'workspaceId': {'type': 'string', 'description': 'Workspace to act on. Defaults to your only workspace if you have exactly one.'}, 'connectionId': {'type': 'string', 'description': 'sftp/aws_s3 only. Connection to pull from. Required for create when type is "sftp"/"aws_s3". Also usable as a run-history filter.'}, 'upstreamSpecName': {'type': 'string', 'description': 'spec_success only. The spec whose successful job completion fires this trigger. Required for create when type is "spec_success".'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'runs': {'type': 'array', 'items': {'type': 'object', 'required': ['triggerId', 'startedAt', 'status', 'filesPulled', 'message'], 'properties': {'jobId': {'type': 'string'}, 'specId': {'type': 'string'}, 'status': {'enum': ['running', 'success', 'failed', 'no-files'], 'type': 'string'}, 'message': {'type': 'string'}, 'specName': {'type': 'string'}, 'startedAt': {'type': 'string'}, 'triggerId': {'type': 'string'}, 'finishedAt': {'type': 'string'}, 'filesPulled': {'type': 'number'}, 'workspaceId': {'type': 'string'}, 'connectionId': {'type': 'string'}}, 'additionalProperties': False}, 'description': 'action "run-history" only'}, 'type': {'enum': ['sftp', 'aws_s3', 'spec_success', 'schedule'], 'type': 'string'}, 'count': {'type': 'number', 'description': 'action "list" only'}, 'dedupe': {'type': 'boolean'}, 'specId': {'type': 'string'}, 'deleted': {'type': 'number', 'description': 'action "clear-processed-files" only'}, 'enabled': {'type': 'boolean'}, 'message': {'type': 'string'}, 'pageSize': {'type': 'number', 'description': 'action "run-history" only'}, 'preRules': {'type': 'string'}, 's3Bucket': {'type': 'string'}, 's3Prefix': {'type': 'string'}, 'specName': {'type': 'string'}, 'triggers': {'type': 'array', 'items': {'type': 'object', 'additionalProperties': {}}, 'description': 'action "list" only'}, 'warnings': {'type': 'array', 'items': {'type': 'string'}, 'description': 'action "create", type "aws_s3" only. Advisory notes, e.g. the 5000-object S3 listing cap — relay to the user.'}, 'createdAt': {'type': 'number'}, 'createdBy': {'type': 'string'}, 'frequency': {'type': 'object', 'required': ['unit'], 'properties': {'unit': {'enum': ['hourly', 'daily', 'monthly'], 'type': 'string'}, 'hourOfDay': {'type': 'number'}, 'dayOfMonth': {'type': 'number'}}, 'additionalProperties': False}, 'lastJobId': {'type': 'string'}, 'lastRunAt': {'type': 'number'}, 'postRules': {'type': 'string'}, 'triggerId': {'type': 'string'}, 'updatedAt': {'type': 'number'}, 'nextCursor': {'type': ['string', 'null'], 'description': 'action "run-history" only'}, 'workspaceId': {'type': 'string'}, 'connectionId': {'type': 'string'}, 'lastRunStatus': {'type': 'string'}, 'upstreamSpecId': {'type': 'string', 'description': 'type "spec_success" only'}, 'upstreamSpecName': {'type': 'string', 'description': 'type "spec_success" only'}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['specName'], 'properties': {'merge': {'type': 'boolean', 'description': 'Upsert instead of plain append when true (default false). For sourceType "tables": generates a MERGE statement instead of an INSERT — use true for a running aggregate/summary that updates existing rows. For sourceType "file" with targetOption "existing-tables": upserts loaded rows by the target table\'s inferred key instead of always inserting — use true whenever the request implies re-loading the same rows shouldn\'t create duplicates (e.g. "upsert on id", syncing/backfilling into a table that already has overlapping rows). Already automatic, no need to request it via additionalPrompt: target columns with no corresponding source column are null on newly inserted rows, and on a match keep their existing value rather than being nulled out.'}, 'specName': {'type': 'string', 'description': 'Name for the new data spec.'}, 'sourceType': {'enum': ['file', 'tables', 'compaction'], 'type': 'string', 'description': 'Defaults to "file" (upload a sample file). Use "tables" to query existing workspace table(s) — see sourceTables — instead of loading a new file. Use "compaction" to bin-pack the small data files of existing tables: it moves no data and produces no new table, so it takes NO target of any kind, needs no analysis, and is ready to run the moment it is created.'}, 'autoRefresh': {'enum': ['spec_success', 'schedule', 'none'], 'type': 'string', 'description': 'sourceType "tables" only. Required for it — ask the user rather than assuming, and do not infer this from other jobs/triggers already in the workspace (a similar existing pipeline is not the user\'s answer for this one). "spec_success" re-runs this spec whenever autoRefreshUpstreamSpecName finishes loading; "schedule" re-runs it on autoRefreshFrequency; "none" leaves it manual-only (re-run later with run_data_job).'}, 'description': {'type': 'string', 'description': 'Optional description of the data spec.'}, 'workspaceId': {'type': 'string', 'description': 'Workspace to act on. Defaults to your only workspace if you have exactly one.'}, 'sourceTables': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Names of existing workspace tables. Required for sourceType "tables" (the tables the generated query reads from) and for sourceType "compaction" (the tables to compact).'}, 'targetOption': {'enum': ['auto-infer', 'existing-tables', 'target-schema-file'], 'type': 'string', 'description': 'Where transformed data should land — works the same for both sourceType values: "auto-infer" (default) lets the AI design the target table (for sourceType "tables", it designs the schema and the query together in one pass), "existing-tables" uses a table already in the workspace (requires targetTables), "target-schema-file" creates the table from a provided schema file (requires targetSchemaFileName).'}, 'targetTables': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Names of existing workspace tables to target — exactly one entry for sourceType "tables" (the generated query has a single target), one or more for sourceType "file". Required when targetOption is "existing-tables". Optional otherwise: for "target-schema-file"/"auto-infer" the target table (and its name) is derived automatically — from the schema file, or AI-designed — unless you want to pin the name yourself, in which case pass exactly one entry.'}, 'formatFileName': {'type': 'string', 'description': 'sourceType "file" only. File name of an optional format spec file.'}, 'sampleFileName': {'type': 'string', 'description': 'sourceType "file" only (and required for it). File name of the sample data file (e.g. "customers.csv") — used to derive content-type, not read from disk.'}, 'additionalPrompt': {'type': 'string', 'description': 'Instructions for the AI. For sourceType "tables", describe what the query should compute from the source table(s) (e.g. "count signups per day per region"). This is stored on the spec verbatim and reused on every future re-analysis, so keep it to instructions that actually change behavior — do not restate default platform behavior (e.g. that unmapped target columns are null/preserved, see merge above) just to document it, since a note that\'s only true for one case (like new rows) can read as a standing instruction later and cause confusion on updates.'}, 'autoRefreshFrequency': {'type': 'object', 'required': ['unit'], 'properties': {'unit': {'enum': ['hourly', 'daily', 'monthly'], 'type': 'string', 'description': 'Schedule cadence.'}, 'hourOfDay': {'type': 'integer', 'maximum': 23, 'minimum': 0, 'description': 'Required for daily/monthly (UTC).'}, 'dayOfMonth': {'type': 'integer', 'maximum': 31, 'minimum': 1, 'description': 'Required for monthly.'}}, 'description': 'Required when autoRefresh is "schedule".', 'additionalProperties': False}, 'expirePriorSnapshots': {'type': 'boolean', 'description': 'sourceType "compaction" only, default false. When false the job commits the compacted files and changes nothing else — prior snapshots still reference the replaced files, so no storage is freed. When true it also expires every snapshot older than its own commit and deletes the replaced files in the same run, which frees storage but ends the ability to roll back to before the compaction.'}, 'targetSchemaFileName': {'type': 'string', 'description': 'File name of a target schema file. Required when targetOption is "target-schema-file", for either sourceType.'}, 'autoRefreshUpstreamSpecName': {'type': 'string', 'description': 'Required when autoRefresh is "spec_success". The spec whose successful job completion should re-run this one.'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['message', 'specId', 'specName', 'files', 'nextStep'], 'properties': {'files': {'type': 'array', 'items': {'type': 'object', 'required': ['fileName', 'url', 'contentType'], 'properties': {'url': {'type': 'string'}, 'fileName': {'type': 'string'}, 'contentType': {'type': 'string'}}, 'additionalProperties': False}}, 'specId': {'type': 'string'}, 'message': {'type': 'string'}, 'nextStep': {'type': 'string', 'description': 'The finish_data_source_onboarding call to make (once upload(s) are done, or immediately for sourceType "tables").'}, 'specName': {'type': 'string'}, 'triggerId': {'type': 'string', 'description': 'sourceType "tables" only, when autoRefresh was not "none": the auto-refresh trigger created alongside the spec.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['specName', 'fileNames'], 'properties': {'specName': {'type': 'string', 'description': 'Name of the already-configured data spec to process files through.'}, 'fileNames': {'type': 'array', 'items': {'type': 'string'}, 'minItems': 1, 'description': 'File names of the data files to process (e.g. ["jan.csv", "feb.csv"])'}, 'workspaceId': {'type': 'string', 'description': 'Workspace to act on. Defaults to your only workspace if you have exactly one.'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['message', 'jobId', 'specName', 'files', 'nextStep'], 'properties': {'files': {'type': 'array', 'items': {'type': 'object', 'required': ['fileName', 'url', 'contentType'], 'properties': {'url': {'type': 'string'}, 'fileName': {'type': 'string'}, 'contentType': {'type': 'string'}}, 'additionalProperties': False}}, 'jobId': {'type': 'string'}, 'message': {'type': 'string'}, 'nextStep': {'type': 'string', 'description': 'The finish_data_job call to make once upload(s) are done.'}, 'specName': {'type': 'string'}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['specName', 'frequency'], 'properties': {'dedupe': {'type': 'boolean', 'description': 'Required — ask the user rather than assuming a value; omitting it fails the call. Whether repeat pulls should skip files already loaded into this spec, matched by file name. Has real consequences: with dedupe true, a file that reappears under the same name (e.g. re-uploaded with corrected data) will be silently skipped; with dedupe false, an unchanged file left on the server will be reloaded every run.'}, 'roleArn': {'type': 'string', 'description': 'aws_s3: the IAM role the customer will create/update.'}, 'hostname': {'type': 'string', 'description': 'sftp: SFTP server hostname to pull from.'}, 'preRules': {'type': 'string', 'description': 'Natural language: which files to pick up (e.g. "only *.csv under /outbound")'}, 's3Bucket': {'type': 'string', 'description': 'aws_s3: bucket to poll. Required when roleArn is given.'}, 's3Prefix': {'type': 'string', 'description': 'aws_s3 only. Optional key prefix; defaults to the whole bucket.'}, 'specName': {'type': 'string', 'description': 'Already-analyzed data spec to load files into (see onboard_data_source)'}, 'username': {'type': 'string', 'description': 'sftp only. Defaults to "sftpuser".'}, 'frequency': {'type': 'object', 'required': ['unit'], 'properties': {'unit': {'enum': ['hourly', 'daily', 'monthly'], 'type': 'string', 'description': 'Schedule cadence.'}, 'hourOfDay': {'type': 'integer', 'maximum': 23, 'minimum': 0, 'description': 'Required for daily/monthly (UTC).'}, 'dayOfMonth': {'type': 'integer', 'maximum': 31, 'minimum': 1, 'description': 'Required for monthly.'}}, 'description': 'Pull schedule.', 'additionalProperties': False}, 'postRules': {'type': 'string', 'description': 'Natural language: what to do after a file loads (e.g. "rename with .done suffix")'}, 'workspaceId': {'type': 'string', 'description': 'Workspace to act on. Defaults to your only workspace if you have exactly one.'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['message'], 'properties': {'dedupe': {'type': 'boolean'}, 'specId': {'type': 'string'}, 'enabled': {'type': 'boolean'}, 'message': {'type': 'string'}, 's3Bucket': {'type': 'string'}, 's3Prefix': {'type': 'string'}, 'specName': {'type': 'string'}, 'warnings': {'type': 'array', 'items': {'type': 'string'}, 'description': 'aws_s3 only. Advisory notes about the created trigger, e.g. the 5000-object S3 listing cap.'}, 'frequency': {'type': 'object', 'required': ['unit'], 'properties': {'unit': {'enum': ['hourly', 'daily', 'monthly'], 'type': 'string'}, 'hourOfDay': {'type': 'number'}, 'dayOfMonth': {'type': 'number'}}, 'additionalProperties': False}, 'triggerId': {'type': 'string', 'description': 'Present only once the connection test succeeded and a trigger was created.'}, 'connectionId': {'type': 'string'}, 'testSucceeded': {'type': 'boolean'}, 'connectionType': {'enum': ['sftp', 'aws_s3'], 'type': 'string'}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['sql'], 'properties': {'sql': {'type': 'string', 'description': 'SQL query, e.g. SELECT * FROM customers LIMIT 10'}, 'workspaceId': {'type': 'string', 'description': 'Workspace to act on. Defaults to your only workspace if you have exactly one.'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['schema', 'rows'], 'properties': {'rows': {'type': 'array', 'items': {'type': 'object', 'additionalProperties': {}}}, 'schema': {'type': 'array', 'items': {'type': 'object', 'required': ['name', 'type'], 'properties': {'name': {'type': 'string'}, 'type': {'type': 'string'}}, 'additionalProperties': False}, 'description': 'Column name -> DuckDB type'}, 'rowCount': {'type': 'number'}, 'executionTimeMs': {'type': 'number'}}, 'additionalProperties': False}
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['specName'], 'properties': {'merge': {'type': 'boolean', 'description': 'Whether new data should merge/upsert into existing rows rather than append. For sourceType "tables" also changes the generated SQL between MERGE and INSERT.'}, 'specName': {'type': 'string', 'description': 'Name of the existing data spec to update'}, 'computeSize': {'enum': ['small', 'large'], 'type': 'string', 'description': 'Compute size for analysis/processing. Omit to keep the current setting.'}, 'description': {'type': 'string', 'description': 'New description for the spec. Omit to keep the current value.'}, 'runAnalysis': {'type': 'boolean', 'description': 'Whether to run analysis and wait for it after saving the changes (default true). Only applies to the synchronous (no-file-change) path.'}, 'workspaceId': {'type': 'string', 'description': 'Workspace to act on. Defaults to your only workspace if you have exactly one.'}, 'sourceTables': {'type': 'array', 'items': {'type': 'string'}, 'description': 'sourceType "tables" specs only: replacement list of source tables the generated query reads from.'}, 'targetOption': {'enum': ['auto-infer', 'existing-tables', 'target-schema-file'], 'type': 'string', 'description': 'Change where transformed data lands. Omit to keep the current setting.'}, 'targetTables': {'type': 'array', 'items': {'type': 'string'}, 'description': 'sourceType "file" specs: new list of existing workspace tables to load into. Required when setting targetOption to "existing-tables". sourceType "tables" specs: the query\'s single target table name — pass a one-element array to rename the target (its schema is re-resolved per the spec\'s targetOption).'}, 'formatFileName': {'type': 'string', 'description': 'sourceType "file" specs only. File name of a replacement format spec file, if replacing it.'}, 'loadSampleData': {'type': 'boolean', 'description': 'Whether re-analysis should also trigger the data-load job (default true). Only used when runAnalysis is true.'}, 'sampleFileName': {'type': 'string', 'description': 'sourceType "file" specs only. File name of a replacement sample data file, if replacing it.'}, 'additionalPrompt': {'type': 'string', 'description': "Extra natural-language guidance for the AI schema inference/mapping. Replaces the previously stored value when given (omit to keep it as-is), and is reused on every future re-analysis — keep it to instructions that actually change behavior. Don't restate default platform behavior (e.g. that unmapped target columns are null on insert and preserved on merge match) just to document it; a note only true for one case (like new rows) can read as a standing instruction later and confuse updates."}, 'targetSchemaFileName': {'type': 'string', 'description': 'File name of a replacement target schema file. Required when setting targetOption to "target-schema-file".'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['message'], 'properties': {'error': {'type': ['string', 'null']}, 'files': {'type': 'array', 'items': {'type': 'object', 'required': ['fileName', 'url', 'contentType'], 'properties': {'url': {'type': 'string'}, 'fileName': {'type': 'string'}, 'contentType': {'type': 'string'}}, 'additionalProperties': False}, 'description': 'Present only when replacement file(s) were given — upload these, then call finish_data_spec_update.'}, 'specId': {'type': 'string'}, 'status': {'type': 'string'}, 'message': {'type': 'string'}, 'nextStep': {'type': 'string', 'description': 'The finish_data_spec_update call to make once upload(s) are done. Only present alongside files.'}, 'progress': {'type': 'number'}, 'specName': {'type': 'string'}, 'timedOut': {'type': 'boolean'}, 'lastJobId': {'type': ['string', 'null']}, 'errorDetails': {'type': ['string', 'null']}, 'statusMessage': {'type': 'string'}, 'hasTransformationConfig': {'type': 'boolean'}}, 'additionalProperties': False}
Recent tool changes
Similar MCP servers
Scalix Cloud
Provides managed databases, SQL tools, container builds, persistent Linux machines, domains, scheduled functions, storage, and pr…
RationalBloks
Creates, deploys, manages, and searches schema-based REST API projects and Neo4j knowledge graphs across staging and production e…
RationalBloks
Creates, deploys, manages, and queries REST API and Neo4j graph projects from JSON schemas, including graph data modeling, versio…
Supabase
Manages Supabase projects, PostgreSQL databases, migrations, branches, storage-related infrastructure, Edge Functions, logs, keys…
Databricks
Connects to Databricks workspaces to browse Unity Catalog metadata, execute SQL, inspect tables and warehouses, and manage long-r…
mcp
Manages Google Cloud Bigtable instances, tables, logical views, and hot-tablet monitoring.
mcp
Manages Google Cloud Datastream resources, including streams, connection profiles, stream objects, static IPs, and long-running o…
site
Provides searchable and statistical access to an RMM software comparison dataset and submits consent-based enquiries to human pro…