CompletionKit
Qué hace este MCP
Runs prompt evaluation workflows over datasets using deterministic checks and LLM judges, with metrics, scoring runs, agreements, and versioned prompts.
Herramientas
Esquema de entrada
{'type': 'object', 'required': ['run_id', 'response_id', 'metric_id', 'verdict'], 'properties': {'note': {'type': 'string'}, 'run_id': {'type': 'integer'}, 'verdict': {'enum': ['agree', 'disagree', 'borderline'], 'type': 'string'}, 'metric_id': {'type': 'integer'}, 'created_by': {'type': 'string'}, 'response_id': {'type': 'integer'}, 'corrected_score': {'type': 'number'}}}
Esquema de entrada
{'type': 'object', 'required': [], 'properties': {'run_id': {'type': 'integer'}, 'metric_id': {'type': 'integer'}, 'created_by': {'type': 'string'}, 'response_id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['name', 'csv_data'], 'properties': {'name': {'type': 'string'}, 'csv_data': {'type': 'string'}, 'tag_names': {'type': 'array', 'items': {'type': 'string'}}}}
Esquema de entrada
{'type': 'object', 'required': ['name', 'url'], 'properties': {'url': {'type': 'string', 'description': 'Public http(s) URL of the CSV file to download.'}, 'name': {'type': 'string'}, 'tag_names': {'type': 'array', 'items': {'type': 'string'}}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': [], 'properties': {}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}, 'name': {'type': 'string'}, 'csv_data': {'type': 'string'}, 'tag_names': {'type': 'array', 'items': {'type': 'string'}}}}
Esquema de entrada
{'type': 'object', 'required': ['metric_id', 'metric_version_a_id', 'metric_version_b_id'], 'properties': {'metric_id': {'type': 'integer'}, 'metric_version_a_id': {'type': 'integer'}, 'metric_version_b_id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['name', 'metric_id', 'dataset_id', 'judge_model'], 'properties': {'name': {'type': 'string'}, 'metric_id': {'type': 'integer'}, 'dataset_id': {'type': 'integer'}, 'judge_model': {'type': 'string'}, 'output_column': {'type': 'string', 'description': 'Dataset column with the existing outputs to grade. Defaults to actual_output.'}}}
Esquema de entrada
{'type': 'object', 'required': ['name'], 'properties': {'name': {'type': 'string'}, 'tag_names': {'type': 'array', 'items': {'type': 'string'}}, 'metric_ids': {'type': 'array', 'items': {'type': 'integer'}}, 'description': {'type': 'string'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': [], 'properties': {}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}, 'name': {'type': 'string'}, 'tag_names': {'type': 'array', 'items': {'type': 'string'}}, 'metric_ids': {'type': 'array', 'items': {'type': 'integer'}}, 'description': {'type': 'string'}}}
Esquema de entrada
{'type': 'object', 'required': ['name'], 'properties': {'name': {'type': 'string'}, 'tag_names': {'type': 'array', 'items': {'type': 'string'}}, 'instruction': {'type': 'string'}, 'metric_type': {'enum': ['llm_judge', 'check'], 'type': 'string'}, 'check_config': {'type': 'object', 'properties': {'max': {'type': 'integer'}, 'min': {'type': 'integer'}, 'trim': {'type': 'boolean'}, 'value': {'type': 'string'}, 'target': {'enum': ['response_text', 'input_data', 'json_path'], 'type': 'string'}, 'pattern': {'type': 'string'}, 'expected': {}, 'json_path': {'type': 'string'}, 'multiline': {'type': 'boolean'}, 'check_kind': {'enum': ['contains', 'not_contains', 'equals', 'regex', 'valid_json', 'json_path_equals', 'length_bounds'], 'type': 'string'}, 'compare_to': {'enum': ['constant', 'expected'], 'type': 'string'}, 'target_path': {'type': 'string'}, 'expected_path': {'type': 'string'}, 'case_sensitive': {'type': 'boolean'}}}, 'rubric_bands': {'type': 'array', 'items': {'type': 'object', 'properties': {'stars': {'type': 'integer'}, 'description': {'type': 'string'}}}}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': [], 'properties': {}}
Esquema de entrada
{'type': 'object', 'required': ['metric_id'], 'properties': {'count': {'type': 'integer', 'description': 'How many variants to request (default 1, max 3). One focused rewrite beats five reworded copies.'}, 'model': {'type': 'string', 'description': 'Override the model used to generate variants. Defaults to the configured judge model or an available judging model.'}, 'metric_id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}, 'name': {'type': 'string'}, 'tag_names': {'type': 'array', 'items': {'type': 'string'}}, 'instruction': {'type': 'string'}, 'metric_type': {'enum': ['llm_judge', 'check'], 'type': 'string'}, 'check_config': {'type': 'object', 'properties': {'max': {'type': 'integer'}, 'min': {'type': 'integer'}, 'trim': {'type': 'boolean'}, 'value': {'type': 'string'}, 'target': {'enum': ['response_text', 'input_data', 'json_path'], 'type': 'string'}, 'pattern': {'type': 'string'}, 'expected': {}, 'json_path': {'type': 'string'}, 'multiline': {'type': 'boolean'}, 'check_kind': {'enum': ['contains', 'not_contains', 'equals', 'regex', 'valid_json', 'json_path_equals', 'length_bounds'], 'type': 'string'}, 'compare_to': {'enum': ['constant', 'expected'], 'type': 'string'}, 'target_path': {'type': 'string'}, 'expected_path': {'type': 'string'}, 'case_sensitive': {'type': 'boolean'}}}, 'rubric_bands': {'type': 'array', 'items': {'type': 'object', 'properties': {'stars': {'type': 'integer'}, 'description': {'type': 'string'}}}}}}
Esquema de entrada
{'type': 'object', 'required': ['metric_version_id'], 'properties': {'metric_version_id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['metric_id'], 'properties': {'metric_id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['metric_version_id'], 'properties': {'metric_version_id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['config'], 'properties': {'config': {'type': 'string', 'description': 'The full promptfooconfig.yaml contents.'}}}
Esquema de entrada
{'type': 'object', 'required': ['name', 'template', 'llm_model'], 'properties': {'name': {'type': 'string'}, 'template': {'type': 'string'}, 'llm_model': {'type': 'string'}, 'tag_names': {'type': 'array', 'items': {'type': 'string'}}, 'description': {'type': 'string'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer', 'description': 'Prompt ID'}}}
Esquema de entrada
{'type': 'object', 'required': [], 'properties': {}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['run_id'], 'properties': {'run_id': {'type': 'integer', 'description': 'The run whose results ground the improvement.'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}, 'name': {'type': 'string'}, 'template': {'type': 'string'}, 'llm_model': {'type': 'string'}, 'tag_names': {'type': 'array', 'items': {'type': 'string'}}, 'description': {'type': 'string'}}}
Esquema de entrada
{'type': 'object', 'required': ['provider', 'api_key'], 'properties': {'api_key': {'type': 'string'}, 'provider': {'enum': ['openai', 'anthropic', 'ollama', 'openrouter', 'azure_foundry'], 'type': 'string'}, 'api_version': {'type': 'string'}, 'api_endpoint': {'type': 'string'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': [], 'properties': {}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}, 'api_key': {'type': 'string'}, 'provider': {'type': 'string'}, 'api_version': {'type': 'string'}, 'api_endpoint': {'type': 'string'}}}
Esquema de entrada
{'type': 'object', 'required': ['run_id', 'id'], 'properties': {'id': {'type': 'integer'}, 'run_id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['run_id'], 'properties': {'sort': {'enum': ['id', 'score_asc', 'score_desc'], 'type': 'string', 'description': 'Row order; defaults to "id".'}, 'limit': {'type': 'integer', 'description': 'Rows to return; defaults to 50, capped at 500.'}, 'fields': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Only return these keys, keeping the payload small. Response keys: id, run_id, input_data, response_text, expected_output, created_at, score, reviewed, reviews, status, attempts, row_index, error. Prefix with "reviews." to trim each review, e.g. ["score", "reviews.metric_name", "reviews.ai_score"]. id is always included.'}, 'offset': {'type': 'integer', 'description': 'Rows to skip before returning results.'}, 'run_id': {'type': 'integer'}, 'status': {'type': 'string', 'description': 'Filter by row status: pending, retrying, succeeded or failed.'}, 'max_score': {'type': 'number', 'description': 'Only rows whose average judge score is at most this. Use with sort "score_asc" for failure-mode analysis.'}, 'min_score': {'type': 'number', 'description': 'Only rows whose average judge score is at least this.'}}}
Esquema de entrada
{'type': 'object', 'required': ['name'], 'properties': {'name': {'type': 'string'}, 'prompt_id': {'type': 'integer'}, 'tag_names': {'type': 'array', 'items': {'type': 'string'}}, 'dataset_id': {'type': 'integer'}, 'max_tokens': {'type': 'integer', 'description': "Cap on generated tokens per row. Leave unset to use the provider client's default, which is what silently truncates long outputs and makes the judge score malformed JSON. Set it to whatever the prompt uses in production so the eval matches."}, 'metric_ids': {'type': 'array', 'items': {'type': 'integer'}}, 'judge_model': {'type': 'string'}, 'temperature': {'type': 'number', 'description': 'Sampling temperature for generation, 0 to 1. Leave it unset, which is the default, and no temperature is sent at all, so the model applies its own. Most current frontier models refuse the parameter outright; set it only when you are targeting a model that honours it, such as anything served locally through Ollama. A refused value is re-sent without one and the run is flagged temperature_ignored.'}, 'output_column': {'type': 'string', 'description': 'Dataset column to grade when prompt_id is omitted; defaults to "actual_output".'}, 'expected_column': {'type': 'string', 'description': 'Dataset column holding each row\'s answer key / ground truth, graded by checks with compare_to "expected" and passed to the judge; defaults to "expected_output".'}, 'metric_group_id': {'type': 'integer', 'description': 'Attach the metrics belonging to this metric group (its current metric_ids). Ignored when metric_ids is also given.'}, 'judge_temperature': {'type': 'number', 'description': "Sampling temperature for the judge, 0 to 1. Defaults to 0 so re-judging the same output gives the same score. Raise it only to measure judge variance on purpose; any value above 0 makes the run's scores irreproducible."}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': [], 'properties': {}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}, 'only': {'type': 'array', 'items': {'type': 'integer'}}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}, 'name': {'type': 'string'}, 'tag_names': {'type': 'array', 'items': {'type': 'string'}}, 'dataset_id': {'type': 'integer'}, 'max_tokens': {'type': 'integer', 'description': "Cap on generated tokens per row. Leave unset to use the provider client's default, which is what silently truncates long outputs and makes the judge score malformed JSON. Set it to whatever the prompt uses in production so the eval matches."}, 'metric_ids': {'type': 'array', 'items': {'type': 'integer'}}, 'judge_model': {'type': 'string'}, 'temperature': {'type': 'number', 'description': 'Sampling temperature for generation, 0 to 1. Leave it unset, which is the default, and no temperature is sent at all, so the model applies its own. Most current frontier models refuse the parameter outright; set it only when you are targeting a model that honours it, such as anything served locally through Ollama. A refused value is re-sent without one and the run is flagged temperature_ignored.'}, 'output_column': {'type': 'string'}, 'expected_column': {'type': 'string'}, 'metric_group_id': {'type': 'integer', 'description': "Replace the run's metrics with those belonging to this metric group. Ignored when metric_ids is also given."}, 'judge_temperature': {'type': 'number', 'description': "Sampling temperature for the judge, 0 to 1. Defaults to 0 so re-judging the same output gives the same score. Raise it only to measure judge variance on purpose; any value above 0 makes the run's scores irreproducible."}}}
Esquema de entrada
{'type': 'object', 'required': ['name'], 'properties': {'name': {'type': 'string'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}}}
Esquema de entrada
{'type': 'object', 'required': [], 'properties': {}}
Esquema de entrada
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'integer'}, 'name': {'type': 'string'}}}
Esquema de entrada
{'type': 'object', 'required': [], 'properties': {}}
Cambios recientes en herramientas
Servidores MCP similares
Andreax
Offers pay-per-call AI services for inference, agent and workflow design, OCR and transcription, code generation and review, clas…
IA-QA — 130+ QA & Dev Tools for AI Agents
Provides deterministic QA, evaluation, testing, code analysis, prompt and RAG checks, model comparison, and web security diagnost…
Replicate
Provides access to Replicate models, versions, collections, hardware, predictions, and deployments for running and managing hoste…
Huggingface
Provides access to Hugging Face model, dataset, and Space metadata, alongside broader structured research and data-routing tools.
Ai Model Experiments
Runs prompts across multiple AI models and compares their outputs, costs, latency, token usage, and errors.
rubrkit
Manages AI evaluation artifacts, rubric audits, eval runs, golden cases, proof reports, and drift monitoring for LLM outputs.
Hipocampo MCP
Provides bilingual semantic memory, embeddings, graph links, code indexing and search, context preload, deduplication, compressio…
Axint
Supports Apple development by generating Swift features, compiling intent definitions, validating and repairing Swift, coordinati…