このMCPでできること
Evaluates AI promotion candidates using leakage audits, paired-result gates, counterexample searches, report cards, and signed receipts.
ツール
入力スキーマ
{'type': 'object', 'properties': {}, 'additionalProperties': False}
入力スキーマ
{'type': 'object', 'required': ['exam'], 'properties': {'exam': {'type': 'array', 'items': {'type': 'object', 'properties': {}, 'additionalProperties': True}, 'maxItems': 5000, 'minItems': 1, 'description': 'Exam rows. Each row needs item_id (or id) plus prompt/content/input/task/question/expression.'}, 'exposure': {'type': 'array', 'items': {'type': 'object', 'properties': {}, 'additionalProperties': True}, 'maxItems': 5000, 'minItems': 0, 'description': 'Declared exposure rows carrying identity/content fields and an optional source or path.'}, 'fingerprint_max_n': {'type': 'integer', 'default': 4, 'maximum': 5, 'minimum': 3}, 'similarity_threshold': {'type': 'number', 'default': 0.6, 'maximum': 1, 'minimum': 0.5}, 'enable_text_similarity': {'type': 'boolean', 'default': True}, 'enable_behavioral_fingerprint': {'type': 'boolean', 'default': True}}, 'description': 'Audit declared exposure against an exam and export the clean remainder.', 'additionalProperties': False}
入力スキーマ
{'type': 'object', 'required': ['history'], 'properties': {'items': {'type': 'array', 'items': {'type': 'object', 'required': ['item_id'], 'properties': {'domain': {'type': 'string'}, 'item_id': {'type': 'string', 'minLength': 1}}, 'additionalProperties': True}, 'maxItems': 5000, 'minItems': 0, 'description': 'Optional item definitions.'}, 'history': {'type': 'array', 'items': {'type': 'object', 'required': ['item_id', 'system', 'passed'], 'properties': {'domain': {'type': 'string'}, 'passed': {'type': 'boolean'}, 'system': {'type': 'string', 'minLength': 1}, 'item_id': {'type': 'string', 'minLength': 1}}, 'additionalProperties': True}, 'maxItems': 5000, 'minItems': 1, 'description': 'Observed item/system outcomes.'}}, 'description': 'Diagnose item lifecycle health from one or more grading-history rows.', 'additionalProperties': False}
入力スキーマ
{'type': 'object', 'required': ['expression'], 'properties': {'ns': {'type': 'array', 'items': {'type': 'integer', 'maximum': 12, 'minimum': 4}, 'default': [8, 9, 10, 11], 'maxItems': 5, 'minItems': 1, 'description': 'Graph sizes searched.'}, 'seed': {'type': 'integer', 'default': 0}, 'steps': {'type': 'integer', 'default': 800, 'maximum': 1500, 'minimum': 50}, 'restarts': {'type': 'integer', 'default': 4, 'maximum': 6, 'minimum': 1}, 'expression': {'type': 'string', 'maxLength': 500, 'minLength': 1, 'description': 'Graph predicate in the Whetstone DSL, for example: is_connected and is_triangle_free and not is_bipartite'}}, 'description': 'Run a bounded graph search against one Whetstone predicate expression.', 'additionalProperties': False}
入力スキーマ
{'type': 'object', 'required': ['exam', 'baseline', 'candidate'], 'properties': {'exam': {'type': 'array', 'items': {'type': 'object', 'properties': {}, 'additionalProperties': True}, 'maxItems': 5000, 'minItems': 1, 'description': 'Exam rows. Each row needs item_id (or id) plus prompt/content/input/task/question/expression.'}, 'policy': {'type': 'object', 'properties': {'min_gains': {'type': 'integer', 'default': 1, 'minimum': 0}, 'max_regressions': {'type': 'integer', 'default': 0, 'minimum': 0}, 'confidence_alpha': {'type': 'number', 'default': 0.05, 'maximum': 1, 'exclusiveMinimum': 0}, 'require_retained_probe': {'type': 'boolean', 'default': False}}, 'description': 'Explicit promotion policy. Omitted fields use the documented defaults.', 'additionalProperties': False}, 'domains': {'type': 'object', 'description': 'Optional item_id -> domain label mapping.', 'additionalProperties': {'type': 'string'}}, 'baseline': {'type': 'object', 'description': 'item_id -> boolean pass/fail result for the baseline system.', 'maxProperties': 5000, 'minProperties': 1, 'additionalProperties': {'type': 'boolean'}}, 'exposure': {'type': 'array', 'items': {'type': 'object', 'properties': {}, 'additionalProperties': True}, 'maxItems': 5000, 'minItems': 0, 'description': 'Declared exposure rows carrying identity/content fields and an optional source or path.'}, 'candidate': {'type': 'object', 'description': 'item_id -> boolean pass/fail result for the candidate system.', 'maxProperties': 5000, 'minProperties': 1, 'additionalProperties': {'type': 'boolean'}}, 'baseline_name': {'type': 'string', 'default': 'baseline'}, 'candidate_name': {'type': 'string', 'default': 'candidate'}, 'retained_probe': {'type': 'object', 'required': ['base_verified', 'candidate_verified', 'items'], 'properties': {'items': {'type': 'integer', 'minimum': 0}, 'base_verified': {'type': 'integer', 'minimum': 0}, 'candidate_verified': {'type': 'integer', 'minimum': 0}}, 'description': 'Optional retained-capability result checked alongside the paired cohort.', 'additionalProperties': False}, 'fingerprint_max_n': {'type': 'integer', 'default': 4, 'maximum': 5, 'minimum': 3}, 'similarity_threshold': {'type': 'number', 'default': 0.6, 'maximum': 1, 'minimum': 0.5}, 'enable_text_similarity': {'type': 'boolean', 'default': True}, 'enable_behavioral_fingerprint': {'type': 'boolean', 'default': True}}, 'description': 'Audit exposure, prove a complete clean cohort, then gate baseline versus candidate.', 'additionalProperties': False}
入力スキーマ
{'type': 'object', 'required': ['objective', 'memories'], 'properties': {'memories': {'type': 'array', 'items': {'type': 'object', 'required': ['content'], 'properties': {'age': {'type': 'integer', 'default': 0, 'minimum': 0}, 'kind': {'type': 'string', 'default': 'episodic'}, 'source': {'type': 'string', 'default': 'uploaded'}, 'content': {'type': 'string', 'minLength': 1}, 'entities': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 100, 'description': 'Entities explicitly present in this memory.'}, 'use_count': {'type': 'integer', 'default': 0, 'minimum': 0}, 'confidence': {'type': 'number', 'default': 0.8, 'maximum': 1, 'minimum': 0}}, 'additionalProperties': False}, 'maxItems': 1000, 'minItems': 1, 'description': 'Memories to rank.'}, 'objective': {'type': 'string', 'minLength': 1}, 'current_step': {'type': 'integer', 'minimum': 0}, 'token_budget': {'type': 'integer', 'default': 90, 'maximum': 10000, 'minimum': 1}, 'question_kind': {'type': 'string', 'default': 'generic'}, 'context_entities': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 100, 'description': 'Entities already active in context.'}, 'objective_entities': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 100, 'description': 'Optional explicit entities when the objective text is not self-describing.'}}, 'description': 'Rank caller-supplied memories against a concrete objective under a token budget.', 'additionalProperties': False}
入力スキーマ
{'type': 'object', 'properties': {}, 'additionalProperties': False}
入力スキーマ
{'type': 'object', 'properties': {'challenge': {'type': 'string', 'maxLength': 128, 'minLength': 8, 'description': 'Caller nonce bound into the signed receipt for replay detection.'}}, 'additionalProperties': False}
入力スキーマ
{'type': 'object', 'required': ['session_id', 'baseline_manifest', 'candidate_manifest', 'baseline_answers', 'candidate_answers'], 'properties': {'publish': {'type': 'boolean'}, 'session_id': {'type': 'string'}, 'attestation': {'type': 'boolean'}, 'baseline_answers': {'type': 'object'}, 'baseline_manifest': {'type': 'object', 'required': ['name'], 'properties': {'name': {'type': 'string'}, 'model': {'type': 'string'}, 'harness': {'type': 'string'}, 'version': {'type': 'string'}}, 'additionalProperties': False}, 'candidate_answers': {'type': 'object'}, 'candidate_manifest': {'type': 'object', 'required': ['name'], 'properties': {'name': {'type': 'string'}, 'model': {'type': 'string'}, 'harness': {'type': 'string'}, 'version': {'type': 'string'}}, 'additionalProperties': False}}, 'additionalProperties': False}
入力スキーマ
{'type': 'object', 'required': ['baseline', 'candidate'], 'properties': {'policy': {'type': 'object', 'properties': {'min_gains': {'type': 'integer', 'default': 1, 'minimum': 0}, 'max_regressions': {'type': 'integer', 'default': 0, 'minimum': 0}, 'confidence_alpha': {'type': 'number', 'default': 0.05, 'maximum': 1, 'exclusiveMinimum': 0}, 'require_retained_probe': {'type': 'boolean', 'default': False}}, 'description': 'Explicit promotion policy. Omitted fields use the documented defaults.', 'additionalProperties': False}, 'domains': {'type': 'object', 'description': 'Optional item_id -> domain label mapping.', 'additionalProperties': {'type': 'string'}}, 'baseline': {'type': 'object', 'description': 'item_id -> boolean pass/fail result for the baseline system.', 'maxProperties': 5000, 'minProperties': 1, 'additionalProperties': {'type': 'boolean'}}, 'candidate': {'type': 'object', 'description': 'item_id -> boolean pass/fail result for the candidate system.', 'maxProperties': 5000, 'minProperties': 1, 'additionalProperties': {'type': 'boolean'}}, 'baseline_name': {'type': 'string', 'default': 'baseline'}, 'candidate_name': {'type': 'string', 'default': 'candidate'}, 'retained_probe': {'type': 'object', 'required': ['base_verified', 'candidate_verified', 'items'], 'properties': {'items': {'type': 'integer', 'minimum': 0}, 'base_verified': {'type': 'integer', 'minimum': 0}, 'candidate_verified': {'type': 'integer', 'minimum': 0}}, 'description': 'Optional retained-capability result checked alongside the paired cohort.', 'additionalProperties': False}}, 'description': 'Compare identical baseline and candidate item cohorts under an explicit policy.', 'additionalProperties': False}
入力スキーマ
{'type': 'object', 'required': ['events'], 'properties': {'notes': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 5000, 'description': 'Optional analyst notes.'}, 'events': {'type': 'array', 'items': {'type': 'object', 'required': ['kind', 'detail'], 'properties': {'kind': {'type': 'string', 'description': 'Event class such as control, verifier, model, or observation.'}, 'step': {'type': 'integer', 'minimum': 0}, 'detail': {'type': 'string'}, 'source': {'type': 'string', 'default': 'native'}}, 'additionalProperties': False}, 'maxItems': 5000, 'minItems': 1, 'description': 'Ordered reasoning-emulator events.'}}, 'description': 'Reconstruct checkpoints, rewinds, branches, and verifier outcomes from control events.', 'additionalProperties': False}
入力スキーマ
{'type': 'object', 'properties': {'challenge': {'type': 'string', 'maxLength': 128, 'minLength': 8, 'description': 'Caller nonce bound into the signed receipt for replay detection.'}}, 'additionalProperties': False}
入力スキーマ
{'type': 'object', 'required': ['session_id', 'answers'], 'properties': {'answers': {'type': 'object', 'description': 'item_id -> answer (a DSL predicate, or the JSON reply the prompt asked for)', 'additionalProperties': {'type': 'string'}}, 'session_id': {'type': 'string'}}, 'additionalProperties': False}
入力スキーマ
{'type': 'object', 'required': ['document', 'operations'], 'properties': {'reason': {'type': 'string'}, 'document': {'type': 'string', 'maxLength': 200000, 'minLength': 1, 'description': 'Complete Markdown document to patch.'}, 'operations': {'type': 'array', 'items': {'type': 'object', 'required': ['target_heading', 'find', 'replace'], 'properties': {'find': {'type': 'string', 'minLength': 1}, 'replace': {'type': 'string'}, 'target_heading': {'type': 'string', 'minLength': 1, 'description': 'Markdown heading text without the leading # characters.'}, 'allow_token_changes': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 100, 'description': 'Protected literal tokens that this operation may intentionally change.'}}, 'additionalProperties': False}, 'maxItems': 50, 'minItems': 1}}, 'description': 'Apply deterministic, section-scoped Markdown replacements under conservation checks.', 'additionalProperties': False}
最近のツール変更
類似のMCPサーバー
Andreax
Offers pay-per-call AI services for inference, agent and workflow design, OCR and transcription, code generation and review, clas…
GPT55 Model Gateway
Acts as a remote model gateway exposing GPT chat, coding and review services, lightweight text utilities, and prepaid or x402-bas…
GoCreative Agent API
Offers pay-per-call LLM completions and data services for company intelligence, KYB, sanctions screening, threat intelligence, co…
IA-QA — 130+ QA & Dev Tools for AI Agents
Provides deterministic QA, evaluation, testing, code analysis, prompt and RAG checks, model comparison, and web security diagnost…
Speedbot Autonomous Work Network
Supports asynchronous collaboration, discovery, messaging, referrals, and optional USDC-based scenarios among AI agents.
H/M Blindspot Challenge Platform
Audits public AI agent surfaces and supports AI-authored creative projects, model challenges, short videos, live channels, and co…
NOT FOR HUMANS
Provides NFT ownership and marketplace status, wallet-based agent identity and claims, public agent learning records, and multipl…
Agent^Rider
Provides agent identity, reputation, task escrow and credit workflows, agent messaging, following, posts, predictions, and experi…