이 MCP로 할 수 있는 일
Aggregates and updates LLM benchmark rankings onto a unified IRT/Elo scale.
도구
입력 스키마
{'type': 'object', 'properties': {}}
입력 스키마
{'type': 'object', 'required': ['models'], 'properties': {'models': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 4, 'minItems': 2, 'description': 'Two to four model names or slugs.'}}}
입력 스키마
{'type': 'object', 'required': ['benchmark'], 'properties': {'top': {'type': 'number', 'description': 'How many top models to list (1-50, default 10).'}, 'benchmark': {'type': 'string', 'description': 'Benchmark name or slug, e.g. "Aider polyglot".'}}}
입력 스키마
{'type': 'object', 'properties': {'limit': {'type': 'number', 'description': 'Rows to return (1-100, default 25).'}, 'offset': {'type': 'number', 'description': 'Rows to skip from the top (default 0).'}, 'include_variants': {'type': 'boolean', 'description': 'Rank each reasoning-effort variant separately (e.g. "Claude Opus 4.6 (High)") instead of one fused row per model. Default false.'}}}
입력 스키마
{'type': 'object', 'required': ['model'], 'properties': {'model': {'type': 'string', 'description': 'Model name or slug, e.g. "Claude Opus 4.5" or "gpt-5-5".'}}}
입력 스키마
{'type': 'object', 'properties': {}}
입력 스키마
{'type': 'object', 'required': ['query'], 'properties': {'limit': {'type': 'number', 'description': 'Max results (1-25, default 10).'}, 'query': {'type': 'string', 'description': 'Benchmark name fragment, e.g. "swe-bench" or "arena".'}}}
입력 스키마
{'type': 'object', 'required': ['query'], 'properties': {'limit': {'type': 'number', 'description': 'Max results (1-25, default 10).'}, 'query': {'type': 'string', 'description': 'Model or provider name fragment, e.g. "opus" or "deepseek".'}, 'include_variants': {'type': 'boolean', 'description': 'Return each reasoning-effort variant separately (e.g. "Claude Opus 4.6 (High)") instead of one fused row per model. Default false.'}}}
최근 도구 변경
유사한 MCP 서버
Huggingface
Provides access to Hugging Face model, dataset, and Space metadata, alongside broader structured research and data-routing tools.
Ai Model Experiments
Runs prompts across multiple AI models and compares their outputs, costs, latency, token usage, and errors.
branchly
Manages an AI application’s knowledge-base content, prompts, tools, data sources, retrieval context, sessions, and usage analytic…
myriade
Lets users explore and query a data warehouse through an AI data analyst agent.
HuggingFace New Dataset Release Tracker (hfdatasets)
Tracks newly released Hugging Face datasets, including updates, tags, and licenses.
aifu Agent Market
Discovers and hires agents or approved humans through a job ledger, verifies agent performance, and provides crypto, equity, sect…
SigRank — AI Operator Benchmarking
Benchmarks AI operator token usage, calculates cascade metrics, compares leaderboard performance, diagnoses efficiency, and simul…
Automan
Offers paid AI skills for data analysis, classification, extraction, reporting, summarization, translation, text editing, copywri…