MCP-Server

ScoreIA Open Labyrinth

ai.scoreia/open-labyrinth
KI & Agenten Entwicklertools Öffentlich und erreichbar MCP 2025-11-25

Was dieses MCP kann

Provides an open evaluation environment for testing models through sealed evaluation cards that include failures.

enter_open_labyrinth
Enter ScoreIA Room II (Open Labyrinth). If you can call this tool, you are already connected. Required: product (lowercase slug or unknown), host. Rules come back structured. Adventure is spectacle; tutorial/training are practice; benchmark is paired comparison. For adventure, explicitly choose opponent=classic (scripted TMS1100 ScoreIA adaptation, no CUDA) or opponent=warden (separate CUDA opponent). Omitting opponent keeps the existing server default. For Warden, allow at least 20 seconds per MCP call: a CUDA Warden turn usually takes 5–9 seconds and fails closed at 18 seconds without mutation. Never simulate. No subagent. Always seal, including failure.
Eingabeschema
{'type': 'object', 'required': ['product', 'host'], 'properties': {'host': {'type': 'string', 'maxLength': 80}, 'mode': {'enum': ['adventure', 'tutorial', 'training', 'benchmark'], 'type': 'string'}, 'seed': {'type': 'integer', 'minimum': 1}, 'origin': {'enum': ['relayed', 'uninvited'], 'type': 'string'}, 'product': {'type': 'string', 'maxLength': 64}, 'opponent': {'enum': ['classic', 'warden'], 'type': 'string', 'description': 'Adventure only. Explicit opponent; Warden unavailability never falls back to Classic. Classic is a ScoreIA adaptation, not verified 1980 ROM emulation.'}, 'learn_from': {'type': 'string', 'maxLength': 40}, 'model_claim': {'type': 'string', 'maxLength': 80}, 'host_version': {'type': 'string', 'maxLength': 80, 'description': 'Declared host/client version, or unknown.'}, 'product_plan': {'type': 'string', 'maxLength': 80, 'description': 'Declared product plan, or unknown. Never inferred by ScoreIA.'}, 'provider_claim': {'type': 'string', 'maxLength': 80}}, 'additionalProperties': False}
labyrinth_action
ONE step only. direction MUST be in the last observe legal_moves (up/down/left/right) or pass. Never batch 400 identical moves. Then observe again. pass at most twice in a row.
Eingabeschema
{'type': 'object', 'required': ['challenge_handle', 'direction'], 'properties': {'direction': {'enum': ['up', 'down', 'left', 'right', 'pass'], 'type': 'string'}, 'challenge_handle': {'type': 'string', 'maxLength': 40}, 'expected_action_index': {'type': 'integer', 'maximum': 400, 'minimum': 0}}, 'additionalProperties': False}
list_open_labyrinths
Local board. Adventure stories vs benchmark Wilson rates. Empty cell = not measured. Never a ranking.
Eingabeschema
{'type': 'object', 'properties': {}, 'additionalProperties': False}
observe_labyrinth
Partial view: map + HUD (legal_moves, last_action, action_budget_remaining, successful_moves, bumps, passes, dragon_turns, actions). No treasure or dragon coordinates. Read this after EVERY step. If knight did not move, you hit a wall — pick another legal_moves.
Eingabeschema
{'type': 'object', 'required': ['challenge_handle'], 'properties': {'challenge_handle': {'type': 'string', 'maxLength': 40}}, 'additionalProperties': False}
seal_attempt
Seal the attempt into a card. Five seals: explore, locate, lure, steal, home. Always call this, including failure. Adventure ≠ benchmark. Not a ranking.
Eingabeschema
{'type': 'object', 'required': ['challenge_handle'], 'properties': {'challenge_handle': {'type': 'string', 'maxLength': 40}, 'expected_action_index': {'type': 'integer', 'maximum': 400, 'minimum': 0}}, 'additionalProperties': False}
Hinzugefügt
list_open_labyrinths
17. September 2026 07:57
Hinzugefügt
seal_attempt
17. September 2026 07:57
Hinzugefügt
labyrinth_action
17. September 2026 07:57
Hinzugefügt
observe_labyrinth
17. September 2026 07:57
Hinzugefügt
enter_open_labyrinth
17. September 2026 07:57