The Aggregate — LLM benchmark aggregate
Was dieses MCP kann
Aggregates and updates LLM benchmark rankings onto a unified IRT/Elo scale.
Tools
Eingabeschema
{'type': 'object', 'properties': {}}
Eingabeschema
{'type': 'object', 'required': ['models'], 'properties': {'models': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 4, 'minItems': 2, 'description': 'Two to four model names or slugs.'}}}
Eingabeschema
{'type': 'object', 'required': ['benchmark'], 'properties': {'top': {'type': 'number', 'description': 'How many top models to list (1-50, default 10).'}, 'benchmark': {'type': 'string', 'description': 'Benchmark name or slug, e.g. "Aider polyglot".'}}}
Eingabeschema
{'type': 'object', 'properties': {'limit': {'type': 'number', 'description': 'Rows to return (1-100, default 25).'}, 'offset': {'type': 'number', 'description': 'Rows to skip from the top (default 0).'}, 'include_variants': {'type': 'boolean', 'description': 'Rank each reasoning-effort variant separately (e.g. "Claude Opus 4.6 (High)") instead of one fused row per model. Default false.'}}}
Eingabeschema
{'type': 'object', 'required': ['model'], 'properties': {'model': {'type': 'string', 'description': 'Model name or slug, e.g. "Claude Opus 4.5" or "gpt-5-5".'}}}
Eingabeschema
{'type': 'object', 'properties': {}}
Eingabeschema
{'type': 'object', 'required': ['query'], 'properties': {'limit': {'type': 'number', 'description': 'Max results (1-25, default 10).'}, 'query': {'type': 'string', 'description': 'Benchmark name fragment, e.g. "swe-bench" or "arena".'}}}
Eingabeschema
{'type': 'object', 'required': ['query'], 'properties': {'limit': {'type': 'number', 'description': 'Max results (1-25, default 10).'}, 'query': {'type': 'string', 'description': 'Model or provider name fragment, e.g. "opus" or "deepseek".'}, 'include_variants': {'type': 'boolean', 'description': 'Return each reasoning-effort variant separately (e.g. "Claude Opus 4.6 (High)") instead of one fused row per model. Default false.'}}}
Letzte Tool-Änderungen
Ähnliche MCP-Server
Huggingface
Provides access to Hugging Face model, dataset, and Space metadata, alongside broader structured research and data-routing tools.
Ai Model Experiments
Runs prompts across multiple AI models and compares their outputs, costs, latency, token usage, and errors.
branchly
Manages an AI application’s knowledge-base content, prompts, tools, data sources, retrieval context, sessions, and usage analytic…
myriade
Lets users explore and query a data warehouse through an AI data analyst agent.
HuggingFace New Dataset Release Tracker (hfdatasets)
Tracks newly released Hugging Face datasets, including updates, tags, and licenses.
aifu Agent Market
Discovers and hires agents or approved humans through a job ledger, verifies agent performance, and provides crypto, equity, sect…
SigRank — AI Operator Benchmarking
Benchmarks AI operator token usage, calculates cascade metrics, compares leaderboard performance, diagnoses efficiency, and simul…
Automan
Offers paid AI skills for data analysis, classification, extraction, reporting, summarization, translation, text editing, copywri…