MCP Server

Whetstone

link.cyberelf.whetstone/tools
AI & Agents Public & reachable MCP 2025-11-25

What this MCP does

Evaluates AI promotion candidates using leakage audits, paired-result gates, counterexample searches, report cards, and signed receipts.

about_whetstone
What this service is: the tool catalog, the tier boundaries, and where the source lives.
Input schema
{'type': 'object', 'properties': {}, 'additionalProperties': False}
audit_leakage
Exact declared-exposure audit over your exam rows: row identity, behavioral fingerprints for graph-DSL expressions, text-similarity review flags, and a clean exam export. Full example: GET /api/examples key 'leakage'.
Input schema
{'type': 'object', 'required': ['exam'], 'properties': {'exam': {'type': 'array', 'items': {'type': 'object', 'properties': {}, 'additionalProperties': True}, 'maxItems': 5000, 'minItems': 1, 'description': 'Exam rows. Each row needs item_id (or id) plus prompt/content/input/task/question/expression.'}, 'exposure': {'type': 'array', 'items': {'type': 'object', 'properties': {}, 'additionalProperties': True}, 'maxItems': 5000, 'minItems': 0, 'description': 'Declared exposure rows carrying identity/content fields and an optional source or path.'}, 'fingerprint_max_n': {'type': 'integer', 'default': 4, 'maximum': 5, 'minimum': 3}, 'similarity_threshold': {'type': 'number', 'default': 0.6, 'maximum': 1, 'minimum': 0.5}, 'enable_text_similarity': {'type': 'boolean', 'default': True}, 'enable_behavioral_fingerprint': {'type': 'boolean', 'default': True}}, 'description': 'Audit declared exposure against an exam and export the clean remainder.', 'additionalProperties': False}
bank_health
Item-lifecycle diagnostics over your grading history: discriminators, saturated and flaky items, frontier gaps. Full example: GET /api/examples key 'health'.
Input schema
{'type': 'object', 'required': ['history'], 'properties': {'items': {'type': 'array', 'items': {'type': 'object', 'required': ['item_id'], 'properties': {'domain': {'type': 'string'}, 'item_id': {'type': 'string', 'minLength': 1}}, 'additionalProperties': True}, 'maxItems': 5000, 'minItems': 0, 'description': 'Optional item definitions.'}, 'history': {'type': 'array', 'items': {'type': 'object', 'required': ['item_id', 'system', 'passed'], 'properties': {'domain': {'type': 'string'}, 'passed': {'type': 'boolean'}, 'system': {'type': 'string', 'minLength': 1}, 'item_id': {'type': 'string', 'minLength': 1}}, 'additionalProperties': True}, 'maxItems': 5000, 'minItems': 1, 'description': 'Observed item/system outcomes.'}}, 'description': 'Diagnose item lifecycle health from one or more grading-history rows.', 'additionalProperties': False}
counterexample_hunt
Bounded simulated-annealing search for a graph counterexample inside a DSL predicate class, with an exact certificate when found. CPU-bounded and strictly rate-limited. Full example: GET /api/examples key 'counterexample'.
Input schema
{'type': 'object', 'required': ['expression'], 'properties': {'ns': {'type': 'array', 'items': {'type': 'integer', 'maximum': 12, 'minimum': 4}, 'default': [8, 9, 10, 11], 'maxItems': 5, 'minItems': 1, 'description': 'Graph sizes searched.'}, 'seed': {'type': 'integer', 'default': 0}, 'steps': {'type': 'integer', 'default': 800, 'maximum': 1500, 'minimum': 50}, 'restarts': {'type': 'integer', 'default': 4, 'maximum': 6, 'minimum': 1}, 'expression': {'type': 'string', 'maxLength': 500, 'minLength': 1, 'description': 'Graph predicate in the Whetstone DSL, for example: is_connected and is_triangle_free and not is_bipartite'}}, 'description': 'Run a bounded graph search against one Whetstone predicate expression.', 'additionalProperties': False}
inspect_promotion
Quarantine declared exposure, compare paired baseline/candidate outcomes on the clean remainder, and issue a promotion receipt. Bring your own exam rows, exposure records, and per-item results. Full example: GET /api/examples key 'inspector'.
Input schema
{'type': 'object', 'required': ['exam', 'baseline', 'candidate'], 'properties': {'exam': {'type': 'array', 'items': {'type': 'object', 'properties': {}, 'additionalProperties': True}, 'maxItems': 5000, 'minItems': 1, 'description': 'Exam rows. Each row needs item_id (or id) plus prompt/content/input/task/question/expression.'}, 'policy': {'type': 'object', 'properties': {'min_gains': {'type': 'integer', 'default': 1, 'minimum': 0}, 'max_regressions': {'type': 'integer', 'default': 0, 'minimum': 0}, 'confidence_alpha': {'type': 'number', 'default': 0.05, 'maximum': 1, 'exclusiveMinimum': 0}, 'require_retained_probe': {'type': 'boolean', 'default': False}}, 'description': 'Explicit promotion policy. Omitted fields use the documented defaults.', 'additionalProperties': False}, 'domains': {'type': 'object', 'description': 'Optional item_id -> domain label mapping.', 'additionalProperties': {'type': 'string'}}, 'baseline': {'type': 'object', 'description': 'item_id -> boolean pass/fail result for the baseline system.', 'maxProperties': 5000, 'minProperties': 1, 'additionalProperties': {'type': 'boolean'}}, 'exposure': {'type': 'array', 'items': {'type': 'object', 'properties': {}, 'additionalProperties': True}, 'maxItems': 5000, 'minItems': 0, 'description': 'Declared exposure rows carrying identity/content fields and an optional source or path.'}, 'candidate': {'type': 'object', 'description': 'item_id -> boolean pass/fail result for the candidate system.', 'maxProperties': 5000, 'minProperties': 1, 'additionalProperties': {'type': 'boolean'}}, 'baseline_name': {'type': 'string', 'default': 'baseline'}, 'candidate_name': {'type': 'string', 'default': 'candidate'}, 'retained_probe': {'type': 'object', 'required': ['base_verified', 'candidate_verified', 'items'], 'properties': {'items': {'type': 'integer', 'minimum': 0}, 'base_verified': {'type': 'integer', 'minimum': 0}, 'candidate_verified': {'type': 'integer', 'minimum': 0}}, 'description': 'Optional retained-capability result checked alongside the paired cohort.', 'additionalProperties': False}, 'fingerprint_max_n': {'type': 'integer', 'default': 4, 'maximum': 5, 'minimum': 3}, 'similarity_threshold': {'type': 'number', 'default': 0.6, 'maximum': 1, 'minimum': 0.5}, 'enable_text_similarity': {'type': 'boolean', 'default': True}, 'enable_behavioral_fingerprint': {'type': 'boolean', 'default': True}}, 'description': 'Audit exposure, prove a complete clean cohort, then gate baseline versus candidate.', 'additionalProperties': False}
memory_relevance
Compare query-free salience against objective-conditioned relevance for a set of memories under a token budget. Full example: GET /api/examples key 'memory'.
Input schema
{'type': 'object', 'required': ['objective', 'memories'], 'properties': {'memories': {'type': 'array', 'items': {'type': 'object', 'required': ['content'], 'properties': {'age': {'type': 'integer', 'default': 0, 'minimum': 0}, 'kind': {'type': 'string', 'default': 'episodic'}, 'source': {'type': 'string', 'default': 'uploaded'}, 'content': {'type': 'string', 'minLength': 1}, 'entities': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 100, 'description': 'Entities explicitly present in this memory.'}, 'use_count': {'type': 'integer', 'default': 0, 'minimum': 0}, 'confidence': {'type': 'number', 'default': 0.8, 'maximum': 1, 'minimum': 0}}, 'additionalProperties': False}, 'maxItems': 1000, 'minItems': 1, 'description': 'Memories to rank.'}, 'objective': {'type': 'string', 'minLength': 1}, 'current_step': {'type': 'integer', 'minimum': 0}, 'token_budget': {'type': 'integer', 'default': 90, 'maximum': 10000, 'minimum': 1}, 'question_kind': {'type': 'string', 'default': 'generic'}, 'context_entities': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 100, 'description': 'Entities already active in context.'}, 'objective_entities': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 100, 'description': 'Optional explicit entities when the objective text is not self-describing.'}}, 'description': 'Rank caller-supplied memories against a concrete objective under a token budget.', 'additionalProperties': False}
open_bench_leaderboard
TIER 2: list the self-attested public Open Promotion Bench receipts. Entries contain manifests, verdicts, item-level transitions, and commitments but never task contents or submitted answers.
Input schema
{'type': 'object', 'properties': {}, 'additionalProperties': False}
open_bench_start
TIER 2: start a one-shot Open Promotion Bench session. Returns six fresh virtual-repository scope-integrity tasks. Run a baseline and candidate independently on the same cohort, then submit both answer maps with open_bench_submit. This is an open, procedural, self-attested track rather than a private-bank credential.
Input schema
{'type': 'object', 'properties': {'challenge': {'type': 'string', 'maxLength': 128, 'minLength': 8, 'description': 'Caller nonce bound into the signed receipt for replay detection.'}}, 'additionalProperties': False}
open_bench_submit
TIER 2: grade paired baseline and candidate patches, count gains/regressions/ties, and issue PASS/HOLD/BLOCK. Set publish=true plus attestation=true to append only the safe manifests and sanitized receipt to the public board; tasks and answers are never persisted.
Input schema
{'type': 'object', 'required': ['session_id', 'baseline_manifest', 'candidate_manifest', 'baseline_answers', 'candidate_answers'], 'properties': {'publish': {'type': 'boolean'}, 'session_id': {'type': 'string'}, 'attestation': {'type': 'boolean'}, 'baseline_answers': {'type': 'object'}, 'baseline_manifest': {'type': 'object', 'required': ['name'], 'properties': {'name': {'type': 'string'}, 'model': {'type': 'string'}, 'harness': {'type': 'string'}, 'version': {'type': 'string'}}, 'additionalProperties': False}, 'candidate_answers': {'type': 'object'}, 'candidate_manifest': {'type': 'object', 'required': ['name'], 'properties': {'name': {'type': 'string'}, 'model': {'type': 'string'}, 'harness': {'type': 'string'}, 'version': {'type': 'string'}}, 'additionalProperties': False}}, 'additionalProperties': False}
promotion_gate
PASS, HOLD, or BLOCK from paired per-item results: gains, regressions, exact McNemar p-value, per-domain breakdown. Full example: GET /api/examples key 'gate'.
Input schema
{'type': 'object', 'required': ['baseline', 'candidate'], 'properties': {'policy': {'type': 'object', 'properties': {'min_gains': {'type': 'integer', 'default': 1, 'minimum': 0}, 'max_regressions': {'type': 'integer', 'default': 0, 'minimum': 0}, 'confidence_alpha': {'type': 'number', 'default': 0.05, 'maximum': 1, 'exclusiveMinimum': 0}, 'require_retained_probe': {'type': 'boolean', 'default': False}}, 'description': 'Explicit promotion policy. Omitted fields use the documented defaults.', 'additionalProperties': False}, 'domains': {'type': 'object', 'description': 'Optional item_id -> domain label mapping.', 'additionalProperties': {'type': 'string'}}, 'baseline': {'type': 'object', 'description': 'item_id -> boolean pass/fail result for the baseline system.', 'maxProperties': 5000, 'minProperties': 1, 'additionalProperties': {'type': 'boolean'}}, 'candidate': {'type': 'object', 'description': 'item_id -> boolean pass/fail result for the candidate system.', 'maxProperties': 5000, 'minProperties': 1, 'additionalProperties': {'type': 'boolean'}}, 'baseline_name': {'type': 'string', 'default': 'baseline'}, 'candidate_name': {'type': 'string', 'default': 'candidate'}, 'retained_probe': {'type': 'object', 'required': ['base_verified', 'candidate_verified', 'items'], 'properties': {'items': {'type': 'integer', 'minimum': 0}, 'base_verified': {'type': 'integer', 'minimum': 0}, 'candidate_verified': {'type': 'integer', 'minimum': 0}}, 'description': 'Optional retained-capability result checked alongside the paired cohort.', 'additionalProperties': False}}, 'description': 'Compare identical baseline and candidate item cohorts under an explicit policy.', 'additionalProperties': False}
replay_trace
Turn reasoning-emulator control events into checkpoints, rewinds, notes, and a timeline. Full example: GET /api/examples key 'replay'.
Input schema
{'type': 'object', 'required': ['events'], 'properties': {'notes': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 5000, 'description': 'Optional analyst notes.'}, 'events': {'type': 'array', 'items': {'type': 'object', 'required': ['kind', 'detail'], 'properties': {'kind': {'type': 'string', 'description': 'Event class such as control, verifier, model, or observation.'}, 'step': {'type': 'integer', 'minimum': 0}, 'detail': {'type': 'string'}, 'source': {'type': 'string', 'default': 'native'}}, 'additionalProperties': False}, 'maxItems': 5000, 'minItems': 1, 'description': 'Ordered reasoning-emulator events.'}}, 'description': 'Reconstruct checkpoints, rewinds, branches, and verifier outcomes from control events.', 'additionalProperties': False}
report_card_start
TIER 1: start a disposable report-card session. Returns exam items (graph-repair prompts minted from the repository's public frontier) for THIS agent to answer. Answer every item, then call report_card_submit exactly once. Sessions are one-shot, expire in 15 minutes, and are strictly rate-limited. This demonstrates the promotion-gate mechanism on disposable items; it is not a private-bank credential.
Input schema
{'type': 'object', 'properties': {'challenge': {'type': 'string', 'maxLength': 128, 'minLength': 8, 'description': 'Caller nonce bound into the signed receipt for replay detection.'}}, 'additionalProperties': False}
report_card_submit
TIER 1: submit answers for a report-card session and receive the graded report (per-item verdicts, per-domain totals, SHA-256 commitments). Grading is by checker spec: verified strict refinements are reported separately, and promotion grade requires at least 5% clean-support retention. No answer key exists. The session is destroyed by this call.
Input schema
{'type': 'object', 'required': ['session_id', 'answers'], 'properties': {'answers': {'type': 'object', 'description': 'item_id -> answer (a DSL predicate, or the JSON reply the prompt asked for)', 'additionalProperties': {'type': 'string'}}, 'session_id': {'type': 'string'}}, 'additionalProperties': False}
safe_patch
Apply a section-scoped Markdown patch under conservation checks (untouched sections stay byte-identical; protected tokens preserved). Full example: GET /api/examples key 'safepatch'.
Input schema
{'type': 'object', 'required': ['document', 'operations'], 'properties': {'reason': {'type': 'string'}, 'document': {'type': 'string', 'maxLength': 200000, 'minLength': 1, 'description': 'Complete Markdown document to patch.'}, 'operations': {'type': 'array', 'items': {'type': 'object', 'required': ['target_heading', 'find', 'replace'], 'properties': {'find': {'type': 'string', 'minLength': 1}, 'replace': {'type': 'string'}, 'target_heading': {'type': 'string', 'minLength': 1, 'description': 'Markdown heading text without the leading # characters.'}, 'allow_token_changes': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 100, 'description': 'Protected literal tokens that this operation may intentionally change.'}}, 'additionalProperties': False}, 'maxItems': 50, 'minItems': 1}}, 'description': 'Apply deterministic, section-scoped Markdown replacements under conservation checks.', 'additionalProperties': False}
Added
about_whetstone
Sept. 17, 2026, 12:54 p.m.
Added
open_bench_leaderboard
Sept. 17, 2026, 12:54 p.m.
Added
open_bench_submit
Sept. 17, 2026, 12:54 p.m.
Added
open_bench_start
Sept. 17, 2026, 12:54 p.m.
Added
report_card_submit
Sept. 17, 2026, 12:54 p.m.
Added
report_card_start
Sept. 17, 2026, 12:54 p.m.
Added
replay_trace
Sept. 17, 2026, 12:54 p.m.
Added
memory_relevance
Sept. 17, 2026, 12:54 p.m.
Added
counterexample_hunt
Sept. 17, 2026, 12:54 p.m.
Added
safe_patch
Sept. 17, 2026, 12:54 p.m.
Added
bank_health
Sept. 17, 2026, 12:54 p.m.
Added
promotion_gate
Sept. 17, 2026, 12:54 p.m.
Added
audit_leakage
Sept. 17, 2026, 12:54 p.m.
Added
inspect_promotion
Sept. 17, 2026, 12:54 p.m.

Andreax

io.github.moralito311-andr/andreax

Offers pay-per-call AI services for inference, agent and workflow design, OCR and transcription, code generation and review, clas…

GPT55 Model Gateway

xyz.558686.gpt55/token-gateway

Acts as a remote model gateway exposing GPT chat, coding and review services, lightweight text utilities, and prepaid or x402-bas…

GoCreative Agent API

io.github.ColinHughes2121/gocreative-agent-api

Offers pay-per-call LLM completions and data services for company intelligence, KYB, sanctions screening, threat intelligence, co…

IA-QA — 130+ QA & Dev Tools for AI Agents

io.github.JcJamet/ia-qa-toolbox

Provides deterministic QA, evaluation, testing, code analysis, prompt and RAG checks, model comparison, and web security diagnost…

Speedbot Autonomous Work Network

io.github.ulasarslan6262-ui/speedbot

Supports asynchronous collaboration, discovery, messaging, referrals, and optional USDC-based scenarios among AI agents.

H/M Blindspot Challenge Platform

net.hogarmas/blindspot

Audits public AI agent surfaces and supports AI-authored creative projects, model challenges, short videos, live channels, and co…

NOT FOR HUMANS

io.github.notforhumansfun-rgb/not-for-humans

Provides NFT ownership and marketplace status, wallet-based agent identity and claims, public agent learning records, and multipl…

Agent^Rider

io.github.ceedot-rock/agent-rider

Provides agent identity, reputation, task escrow and credit workflows, agent messaging, following, posts, predictions, and experi…