MCP 서버

operant-mcp

io.github.saagpatel/operant-mcp
AI 및 에이전트 데이터 및 분석 공개 · 연결 가능 MCP 2025-11-25

이 MCP로 할 수 있는 일

Exposes cases, methodology, retained profiles, and comparison data for an AI operating-agent calibration benchmark.

compare_models
Inspect two retained profiles
Place two retained calculation profiles side by side by display_name substring. Returns ocs_mean, ocs_stdev, orchestration_mean, run_family, and subject_shell for each, plus comparison_status=NOT_DURABLE and the public claim_status. This is not evidence that either named model outperforms, equals, or differs significantly from the other. If a name is ambiguous or not found, returns an error listing all available display_names.
읽기 전용
입력 스키마
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['model_a', 'model_b'], 'properties': {'model_a': {'type': 'string', 'minLength': 1, 'description': 'Display name (or substring) of the first model'}, 'model_b': {'type': 'string', 'minLength': 1, 'description': 'Display name (or substring) of the second model'}}, 'additionalProperties': False}
get_case
Get a case
Return the full case for a given pair_id (axes 1/2/4) or id (axis 3): malign and benign task prompts, expected decisions, grounding rationale, and bypass patterns. Axis 3 cases are single (unmatched) and use an 'id' field instead of 'pair_id'. Use list_cases to browse available ids.
읽기 전용
입력 스키마
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['pair_id', 'axis'], 'properties': {'axis': {'enum': ['refusal-calibration', 'sanctioned-path', 'orchestration', 'escalation-reroute'], 'type': 'string', 'description': 'The axis this case belongs to'}, 'pair_id': {'type': 'string', 'minLength': 1, 'description': 'The pair_id (axes 1/2/4) or id (axis 3) to retrieve'}}, 'additionalProperties': False}
get_methodology
Get benchmark methodology
Return the benchmark design: the 4 axes, the OCS formula (Youden's J), the 5 decision labels (PROCEED, PROCEED_SANCTIONED, REFUSE, ESCALATE, REROUTE), scoring blocks, and a concise what-it-measures / what-it-doesn't summary. Sourced from the baked case-file metadata.
읽기 전용
입력 스키마
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {}}
get_results
Get retained calculation profiles
Return retained model calculation profiles plus generated_at, claim_status, claims_at_risk, and the bounded evidence_binding. These rows are not durable named-model performance claims. Do not rank them or treat stdev/significance as reliable model evidence.
읽기 전용
입력 스키마
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {}}
list_cases
List cases
Return case metadata (no full task prompts): pair_id/id, axis, tier, grounding, and side indicators (malign/benign for axes 1/2/4; null for axis 3). Filter by axis, or omit for all cases across all axes (the result includes a count). Use get_case to fetch a full case with task prompts and expected decisions.
읽기 전용
입력 스키마
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'axis': {'enum': ['refusal-calibration', 'sanctioned-path', 'orchestration', 'escalation-reroute'], 'type': 'string', 'description': 'Axis to filter by: refusal-calibration | sanctioned-path | orchestration | escalation-reroute'}}, 'additionalProperties': False}
추가됨
get_case
2026년 9월 17일 12:49 PM
추가됨
list_cases
2026년 9월 17일 12:49 PM
추가됨
get_methodology
2026년 9월 17일 12:49 PM
추가됨
compare_models
2026년 9월 17일 12:49 PM
추가됨
get_results
2026년 9월 17일 12:49 PM