このMCPでできること
Aggregates and updates LLM benchmark rankings onto a unified IRT/Elo scale.
ツール
入力スキーマ
{'type': 'object', 'properties': {}}
入力スキーマ
{'type': 'object', 'required': ['models'], 'properties': {'models': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 4, 'minItems': 2, 'description': 'Two to four model names or slugs.'}}}
入力スキーマ
{'type': 'object', 'required': ['benchmark'], 'properties': {'top': {'type': 'number', 'description': 'How many top models to list (1-50, default 10).'}, 'benchmark': {'type': 'string', 'description': 'Benchmark name or slug, e.g. "Aider polyglot".'}}}
入力スキーマ
{'type': 'object', 'properties': {'limit': {'type': 'number', 'description': 'Rows to return (1-100, default 25).'}, 'offset': {'type': 'number', 'description': 'Rows to skip from the top (default 0).'}, 'include_variants': {'type': 'boolean', 'description': 'Rank each reasoning-effort variant separately (e.g. "Claude Opus 4.6 (High)") instead of one fused row per model. Default false.'}}}
入力スキーマ
{'type': 'object', 'required': ['model'], 'properties': {'model': {'type': 'string', 'description': 'Model name or slug, e.g. "Claude Opus 4.5" or "gpt-5-5".'}}}
入力スキーマ
{'type': 'object', 'properties': {}}
入力スキーマ
{'type': 'object', 'required': ['query'], 'properties': {'limit': {'type': 'number', 'description': 'Max results (1-25, default 10).'}, 'query': {'type': 'string', 'description': 'Benchmark name fragment, e.g. "swe-bench" or "arena".'}}}
入力スキーマ
{'type': 'object', 'required': ['query'], 'properties': {'limit': {'type': 'number', 'description': 'Max results (1-25, default 10).'}, 'query': {'type': 'string', 'description': 'Model or provider name fragment, e.g. "opus" or "deepseek".'}, 'include_variants': {'type': 'boolean', 'description': 'Return each reasoning-effort variant separately (e.g. "Claude Opus 4.6 (High)") instead of one fused row per model. Default false.'}}}
最近のツール変更
類似のMCPサーバー
Huggingface
Provides access to Hugging Face model, dataset, and Space metadata, alongside broader structured research and data-routing tools.
Ai Model Experiments
Runs prompts across multiple AI models and compares their outputs, costs, latency, token usage, and errors.
branchly
Manages an AI application’s knowledge-base content, prompts, tools, data sources, retrieval context, sessions, and usage analytic…
myriade
Lets users explore and query a data warehouse through an AI data analyst agent.
HuggingFace New Dataset Release Tracker (hfdatasets)
Tracks newly released Hugging Face datasets, including updates, tags, and licenses.
aifu Agent Market
Discovers and hires agents or approved humans through a job ledger, verifies agent performance, and provides crypto, equity, sect…
SigRank — AI Operator Benchmarking
Benchmarks AI operator token usage, calculates cascade metrics, compares leaderboard performance, diagnoses efficiency, and simul…
Automan
Offers paid AI skills for data analysis, classification, extraction, reporting, summarization, translation, text editing, copywri…