The Aggregate — LLM benchmark aggregate
What this MCP does
Aggregates and updates LLM benchmark rankings onto a unified IRT/Elo scale.
Tools
Input schema
{'type': 'object', 'properties': {}}
Input schema
{'type': 'object', 'required': ['models'], 'properties': {'models': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 4, 'minItems': 2, 'description': 'Two to four model names or slugs.'}}}
Input schema
{'type': 'object', 'required': ['benchmark'], 'properties': {'top': {'type': 'number', 'description': 'How many top models to list (1-50, default 10).'}, 'benchmark': {'type': 'string', 'description': 'Benchmark name or slug, e.g. "Aider polyglot".'}}}
Input schema
{'type': 'object', 'properties': {'limit': {'type': 'number', 'description': 'Rows to return (1-100, default 25).'}, 'offset': {'type': 'number', 'description': 'Rows to skip from the top (default 0).'}, 'include_variants': {'type': 'boolean', 'description': 'Rank each reasoning-effort variant separately (e.g. "Claude Opus 4.6 (High)") instead of one fused row per model. Default false.'}}}
Input schema
{'type': 'object', 'required': ['model'], 'properties': {'model': {'type': 'string', 'description': 'Model name or slug, e.g. "Claude Opus 4.5" or "gpt-5-5".'}}}
Input schema
{'type': 'object', 'properties': {}}
Input schema
{'type': 'object', 'required': ['query'], 'properties': {'limit': {'type': 'number', 'description': 'Max results (1-25, default 10).'}, 'query': {'type': 'string', 'description': 'Benchmark name fragment, e.g. "swe-bench" or "arena".'}}}
Input schema
{'type': 'object', 'required': ['query'], 'properties': {'limit': {'type': 'number', 'description': 'Max results (1-25, default 10).'}, 'query': {'type': 'string', 'description': 'Model or provider name fragment, e.g. "opus" or "deepseek".'}, 'include_variants': {'type': 'boolean', 'description': 'Return each reasoning-effort variant separately (e.g. "Claude Opus 4.6 (High)") instead of one fused row per model. Default false.'}}}
Recent tool changes
Similar MCP servers
Huggingface
Provides access to Hugging Face model, dataset, and Space metadata, alongside broader structured research and data-routing tools.
Ai Model Experiments
Runs prompts across multiple AI models and compares their outputs, costs, latency, token usage, and errors.
branchly
Manages an AI application’s knowledge-base content, prompts, tools, data sources, retrieval context, sessions, and usage analytic…
myriade
Lets users explore and query a data warehouse through an AI data analyst agent.
HuggingFace New Dataset Release Tracker (hfdatasets)
Tracks newly released Hugging Face datasets, including updates, tags, and licenses.
aifu Agent Market
Discovers and hires agents or approved humans through a job ledger, verifies agent performance, and provides crypto, equity, sect…
SigRank — AI Operator Benchmarking
Benchmarks AI operator token usage, calculates cascade metrics, compares leaderboard performance, diagnoses efficiency, and simul…
Automan
Offers paid AI skills for data analysis, classification, extraction, reporting, summarization, translation, text editing, copywri…