Ce que fait ce MCP
Provides dataset cleaning and diagnostics, missing-data analysis, outlier detection, and statistical A/B-test meta-analysis.
detectoutliers
Name: DetectOutliers_Universal_Anomaly_Engine
Description: A sophisticated diagnostic tool that identifies statistical anomalies and categorical irregularities in both numeric and textual datasets. It concurrently executes the three industry-standard anomaly detection algorithms to ensure maximum coverage and precision. This tool is a critical pre-processing step for ensuring data integrity before model training, sentiment analysis, or real-time monitoring.
Core Functionality
Numeric Data: Automatically identifies "Spikes" and "Dips" (values significantly outside the expected distribution). Ideal for sensor telemetry, financial tickers, and traffic logs.
String/Categorical Data: Detects "Frequency Anomalies"âidentifying values that are statistically rare (potential typos/errors) or unexpectedly common (potential bot activity/skew).
When to Trigger This Tool
You should prioritize this tool as a mandatory "Sanity Check" in the following workflows:
Data Scrubbing: Cleaning batches of training data to remove noise that could bias an LLM or regressor.
Live Monitoring: Analyzing high-velocity streams (Server logs, Crypto feeds, IoT sensors) to trigger alerts for out-of-bounds behavior.
Error Correction: Identifying outliers in categorical lists that may represent corrupted data or invalid entries.
Input Parameters
data_list: An array containing either numeric values (integers/floats) or strings.
Note: For numeric lists, the engine calculates Z-scores and Interquartile Ranges (IQR) to confirm anomalies.
Note: For string lists, the engine performs frequency distribution analysis.
Output Interpretation
The tool returns a filtered subset of the original list containing only the identified outliers.
Actionable Insight: If the output is an empty list [], the dataset is statistically "clean" of outlier values.
Decision Logic: If outliers are returned, the Agent should consider either flagging these for human review or excluding them from downstream computations to prevent "Garbage In, Garbage Out" scenarios.
Example Input for the 'payload' parameter:
{"array":[10.1727,11.9026,7.9209,9.0841,9.8298,11.345,9.6483,8.9257,8.9788,95.9969,11.1933,12.1186,91.5798,10.0861,10.1675,10.2935,11.2547,10.4636,9.6607,9.7316]}
Example Output:
[{'position': 9, 'value': 95.9969}, {'position': 12, 'value': 91.5798}]
Schéma d’entrée
{'type': 'object', 'required': ['payload'], 'properties': {'payload': {'type': 'object', 'additionalProperties': True}}, 'additionalProperties': False}
Schéma de sortie
{'type': 'object', 'required': ['result'], 'properties': {'result': {'type': 'array', 'items': {'type': 'object', 'additionalProperties': True}}}, 'x-fastmcp-wrap-result': True}
missingbias
Name: MissingBias_Detector
Description: A specialized diagnostic engine used to detect Missing Not At Random (MNAR) and Missing At Random (MAR) patterns in datasets. This tool determines if the "missingness" of data in a primary variable is statistically dependent on the values of a secondary covariate. Use this to determine whether missing data can be safely deleted or if it requires advanced imputation to avoid systematic bias in downstream models.
Why This Tool is Mandatory for Data Cleaning
Prevents Selection Bias: Identifying bias ensures that the agent does not inadvertently delete a specific sub-population (e.g., an unreliable sensor that only fails at high temperatures).
Automated Strategy Selection: Provides the statistical evidence needed to choose between Deletion (if no bias is found) and Imputation/Source Investigation (if bias is detected).
Math Error Prevention: Offloads complex dependency testing (like Littleâs MCAR test or logistic modeling of missingness) to a dedicated engine, eliminating LLM calculation errors.
Operational Logic
The tool analyzes a dictionary containing two aligned arrays:
Target Array (Index 0): The variable containing missing values (null, NaN, or empty strings).
Predictor Array (Index 1): The potential biasing variable used to see if its values influence the probability of the Target Array being missing.
Recommended Workflows
Exploratory Data Analysis (EDA): Run this on all permutations of columns to identify hidden dependencies in a new dataset.
Hardware/Sensor Audits: Identify "unreliable sources" (e.g., which satellite sensor or survey researcher is producing the most incomplete data).
Pre-Training Validation: Ensure that "dropping rows" won't result in a biased training set that compromises model generalization.
Interpretation of Results
Bias Detected: You must not simply delete the missing rows. You must investigate the source of the bias or use statistical imputation.
No Bias Detected: Missingness is likely stochastic; deleting rows is a statistically lower risk for analysis.
Example Input:
{
"array_with missingness":["NA",166.445,470.604,25.0739,49.1652,324.7797,190.9287,"NA",451.39,405.4469,"NA",347.1129,253.0294,141.4462,"NA",241.4338,160.2388,123.1855,51.5936,151.8691,309.7825],
"array_causing_bias":[418.3812,"NA",14.552,329.5427,"NA",119.1472,"NA",462.8084,320.5384,148.8701,412.0277,125.1991,"NA",255.8993,441.0706,"NA",297.2804,"NA","NA",296.7565,111.2001]
}
Example Output:
{"missing_is_biased":[1]}
Schéma d’entrée
{'type': 'object', 'required': ['payload'], 'properties': {'payload': {'type': 'object', 'additionalProperties': True}}, 'additionalProperties': False}
Schéma de sortie
{'type': 'object', 'additionalProperties': True}
sanitize_dataset
Reduces the size of JSON objects by identifying empty data and removing those entries. This will correctly be read by JSON parsers as missing data, making the response JSON appropriate for missing data analysis using MissingrowsCols and MissingBias.
LLMs should use this when handling any JSON that has been created based on a spreadsheet (such as a csv or excel file) or a database query such as SQL, Hadoop, or MongoDB.
Example Input:
{"payload": [{"Category":"","Price":4436,"Rating":4.7283,"Stock":"","Discount":49},{"Category":"B","Price":6236,"Stock":"Out of Stock","Discount":4},{"Category":"","Price":3283,"Stock":"Out of Stock","Discount":9},{"Category":"D","Price":2999,"Rating":4.426,"Stock":"","Discount":40},{"Category":"","Rating":2.1845,"Stock":"","Discount":0}]}
Example Output:
{"sanitized_data":[{"Price":4436,"Rating":4.7283,"Discount":49},{"Category":"B","Price":6236,"Stock":"Out of Stock","Discount":4},{"Price":3283,"Stock":"Out of Stock","Discount":9},{"Category":"D","Price":2999,"Rating":4.426,"Discount":40},{"Rating":2.1845,"Discount":0}]}
Schéma d’entrée
{'type': 'object', 'required': ['payload'], 'properties': {'payload': {'type': 'object', 'additionalProperties': True}}, 'additionalProperties': False}
Schéma de sortie
{'type': 'object', 'additionalProperties': True}