LLM Configurator
Ce que fait ce MCP
Checks local language-model hardware compatibility, estimated speed, memory requirements, model specifications, and available models for selected machines.
Outils
Schéma d’entrée
{'type': 'object', 'required': ['hardware_id', 'model_id'], 'properties': {'model_id': {'type': 'string', 'pattern': '^[a-z0-9][a-z0-9._-]*$', 'maxLength': 80, 'minLength': 1, 'description': 'Model id from search_catalog.'}, 'hardware_id': {'type': 'string', 'pattern': '^[a-z0-9][a-z0-9._-]*$', 'maxLength': 80, 'minLength': 1, 'description': 'Hardware id from search_catalog.'}, 'quantisation': {'enum': ['Q2_K', 'Q3_K_M', 'Q4_K_M', 'Q5_K_M', 'Q6_K', 'Q8_0', 'F16'], 'type': 'string', 'default': 'Q4_K_M', 'description': 'Weight quantisation. One of Q2_K, Q3_K_M, Q4_K_M, Q5_K_M, Q6_K, Q8_0, F16. Default Q4_K_M.'}, 'context_length': {'type': 'integer', 'default': 4096, 'maximum': 262144, 'minimum': 512, 'description': 'Context window in tokens that the KV cache is sized for. 512 to 262144. Default 4096.'}}}
Schéma de sortie
{'type': 'object', 'required': ['data_as_of', 'summary', 'verdict', 'quantisation', 'context_length', 'model', 'hardware', 'memory', 'speed', 'notes', 'sources', 'url'], 'properties': {'url': {'type': 'string', 'pattern': '^https:\\/\\/llmconfigurator\\.com\\/en\\/'}, 'model': {'type': 'object', 'required': ['id', 'name', 'total_params_b', 'active_params_b', 'is_moe'], 'properties': {'id': {'type': 'string', 'pattern': '^[a-z0-9][a-z0-9._-]*$', 'maxLength': 80, 'minLength': 1}, 'name': {'type': 'string', 'maxLength': 120}, 'is_moe': {'type': 'boolean'}, 'total_params_b': {'type': 'number', 'description': 'Total parameters in billions. Sets memory residency.', 'exclusiveMinimum': 0}, 'active_params_b': {'type': 'number', 'description': 'Parameters read per token in billions. Equals total for dense models. Sets speed, never memory.', 'exclusiveMinimum': 0}}, 'additionalProperties': False}, 'notes': {'type': 'array', 'items': {'type': 'string', 'maxLength': 300}, 'maxItems': 8}, 'speed': {'anyOf': [{'type': 'object', 'required': ['low_tok_s', 'high_tok_s', 'confidence', 'basis', 'sample_count', 'last_verified'], 'properties': {'basis': {'anyOf': [{'type': 'string', 'maxLength': 200}, {'type': 'null'}], 'description': 'Where the figure comes from, e.g. a benchmark label. null for a modelled estimate.'}, 'low_tok_s': {'type': 'number', 'minimum': 0}, 'confidence': {'enum': ['measured', 'estimated', 'community'], 'type': 'string'}, 'high_tok_s': {'type': 'number', 'minimum': 0}, 'sample_count': {'anyOf': [{'type': 'integer', 'maximum': 9007199254740991, 'exclusiveMinimum': 0}, {'type': 'null'}], 'description': 'Number of measured runs behind a measured or community figure.'}, 'last_verified': {'anyOf': [{'type': 'string', 'pattern': '^\\d{4}-\\d{2}-\\d{2}$'}, {'type': 'null'}]}}, 'description': 'Decode speed in tokens per second, always a range with a confidence label.', 'additionalProperties': False}, {'type': 'null'}], 'description': 'null when the model does not fit, because a speed for hardware that cannot load the weights would be invented.'}, 'memory': {'type': 'object', 'required': ['weights_gb', 'kv_cache_gb', 'overhead_gb', 'needed_gb', 'available_gb', 'headroom_gb', 'offloaded_gb', 'confidence', 'kv_basis'], 'properties': {'kv_basis': {'enum': ['published_architecture', 'inferred_architecture'], 'type': 'string', 'description': 'Whether the KV cache is computed from a published config or inferred from the parameter count.'}, 'needed_gb': {'type': 'number', 'minimum': 0, 'description': 'weights + KV cache + overhead.'}, 'confidence': {'enum': ['measured', 'estimated', 'community'], 'type': 'string'}, 'weights_gb': {'type': 'number', 'minimum': 0, 'description': 'All parameters quantised. For a Mixture-of-Experts model this counts every expert.'}, 'headroom_gb': {'type': 'number', 'description': 'available - needed. Negative when the model does not fit.'}, 'kv_cache_gb': {'type': 'number', 'minimum': 0}, 'overhead_gb': {'type': 'number', 'minimum': 0, 'description': 'Fixed framework overhead.'}, 'available_gb': {'type': 'number', 'minimum': 0, 'description': 'Memory usable for the model after any operating-system reserve.'}, 'offloaded_gb': {'type': 'number', 'minimum': 0, 'description': 'Weights that would live in system RAM. 0 when fully resident.'}}, 'description': 'Memory arithmetic for the requested quantisation and context length.', 'additionalProperties': False}, 'sources': {'type': 'array', 'items': {'type': 'string', 'pattern': '^https:\\/\\/'}, 'maxItems': 8, 'description': 'Primary sources for the model and hardware records. May be empty.'}, 'summary': {'type': 'string', 'maxLength': 600, 'description': 'One or two plain-language sentences stating the answer.'}, 'verdict': {'enum': ['fits_comfortably', 'fits_tight', 'needs_offload', 'does_not_fit'], 'type': 'string'}, 'hardware': {'type': 'object', 'required': ['id', 'name', 'memory_gb', 'memory_kind', 'bandwidth_gb_s', 'status'], 'properties': {'id': {'type': 'string', 'pattern': '^[a-z0-9][a-z0-9._-]*$', 'maxLength': 80, 'minLength': 1}, 'name': {'type': 'string', 'maxLength': 120}, 'status': {'enum': ['current', 'eol', 'used-market', 'announced'], 'type': 'string', 'description': 'announced means not yet on sale; eol means no longer made.'}, 'memory_gb': {'type': 'number', 'exclusiveMinimum': 0}, 'memory_kind': {'enum': ['dedicated', 'unified', 'shared-system'], 'type': 'string'}, 'bandwidth_gb_s': {'type': 'number', 'description': 'Theoretical peak memory bandwidth.', 'exclusiveMinimum': 0}}, 'additionalProperties': False}, 'data_as_of': {'type': 'string', 'pattern': '^\\d{4}-\\d{2}-\\d{2}$', 'description': 'Date the bundled catalogue was last updated. Not the time of this call.'}, 'quantisation': {'enum': ['Q2_K', 'Q3_K_M', 'Q4_K_M', 'Q5_K_M', 'Q6_K', 'Q8_0', 'F16'], 'type': 'string'}, 'context_length': {'type': 'integer', 'maximum': 9007199254740991, 'minimum': -9007199254740991}}, 'additionalProperties': False}
Schéma d’entrée
{'type': 'object', 'required': ['hardware_ids'], 'properties': {'hardware_ids': {'type': 'array', 'items': {'type': 'string', 'pattern': '^[a-z0-9][a-z0-9._-]*$', 'maxLength': 80, 'minLength': 1}, 'maxItems': 4, 'minItems': 2, 'description': 'Two to four hardware ids from search_catalog.'}}}
Schéma de sortie
{'type': 'object', 'required': ['data_as_of', 'summary', 'hardware'], 'properties': {'summary': {'type': 'string', 'maxLength': 600, 'description': 'One or two plain-language sentences stating the answer.'}, 'hardware': {'type': 'array', 'items': {'type': 'object', 'required': ['id', 'name', 'memory_gb', 'memory_kind', 'bandwidth_gb_s', 'status', 'class', 'architecture', 'release_year', 'system_ram_gb', 'price', 'notes', 'sources', 'url'], 'properties': {'id': {'type': 'string', 'pattern': '^[a-z0-9][a-z0-9._-]*$', 'maxLength': 80, 'minLength': 1}, 'url': {'type': 'string', 'pattern': '^https:\\/\\/llmconfigurator\\.com\\/en\\/'}, 'name': {'type': 'string', 'maxLength': 120}, 'class': {'type': 'string', 'maxLength': 40}, 'notes': {'anyOf': [{'type': 'string', 'maxLength': 400}, {'type': 'null'}]}, 'price': {'anyOf': [{'type': 'object', 'required': ['usd', 'as_of', 'market', 'stale', 'source'], 'properties': {'usd': {'type': 'number', 'exclusiveMinimum': 0}, 'as_of': {'type': 'string', 'pattern': '^\\d{4}-\\d{2}-\\d{2}$', 'description': 'Date this price was checked.'}, 'stale': {'type': 'boolean', 'description': 'True when the price is older than 90 days.'}, 'market': {'enum': ['new', 'used', 'refurb'], 'type': 'string'}, 'source': {'type': 'string', 'maxLength': 200}}, 'additionalProperties': False}, {'type': 'null'}], 'description': 'null when this catalogue holds no dated price. Never 0.'}, 'status': {'enum': ['current', 'eol', 'used-market', 'announced'], 'type': 'string', 'description': 'announced means not yet on sale; eol means no longer made.'}, 'sources': {'type': 'array', 'items': {'type': 'string', 'pattern': '^https:\\/\\/'}, 'maxItems': 8}, 'memory_gb': {'type': 'number', 'exclusiveMinimum': 0}, 'memory_kind': {'enum': ['dedicated', 'unified', 'shared-system'], 'type': 'string'}, 'architecture': {'type': 'string', 'maxLength': 80}, 'release_year': {'anyOf': [{'type': 'integer', 'maximum': 9007199254740991, 'minimum': -9007199254740991}, {'type': 'null'}]}, 'system_ram_gb': {'type': 'number', 'minimum': 0, 'description': 'System RAM available for offload. 0 for unified-memory machines.'}, 'bandwidth_gb_s': {'type': 'number', 'description': 'Theoretical peak memory bandwidth.', 'exclusiveMinimum': 0}}, 'additionalProperties': False}, 'maxItems': 4, 'minItems': 2}, 'data_as_of': {'type': 'string', 'pattern': '^\\d{4}-\\d{2}-\\d{2}$', 'description': 'Date the bundled catalogue was last updated. Not the time of this call.'}}, 'additionalProperties': False}
Schéma d’entrée
{'type': 'object', 'required': ['model_id'], 'properties': {'model_id': {'type': 'string', 'pattern': '^[a-z0-9][a-z0-9._-]*$', 'maxLength': 80, 'minLength': 1, 'description': 'Model id from search_catalog.'}}}
Schéma de sortie
{'type': 'object', 'required': ['data_as_of', 'summary', 'model', 'quantisations', 'sources', 'url'], 'properties': {'url': {'type': 'string', 'pattern': '^https:\\/\\/llmconfigurator\\.com\\/en\\/'}, 'model': {'type': 'object', 'required': ['id', 'name', 'total_params_b', 'active_params_b', 'is_moe', 'family', 'context_length', 'license', 'release_date', 'status'], 'properties': {'id': {'type': 'string', 'pattern': '^[a-z0-9][a-z0-9._-]*$', 'maxLength': 80, 'minLength': 1}, 'name': {'type': 'string', 'maxLength': 120}, 'family': {'type': 'string', 'maxLength': 120}, 'is_moe': {'type': 'boolean'}, 'status': {'enum': ['current', 'legacy', 'deprecated'], 'type': 'string'}, 'license': {'anyOf': [{'type': 'string', 'maxLength': 80}, {'type': 'null'}]}, 'release_date': {'anyOf': [{'type': 'string', 'pattern': '^\\d{4}-\\d{2}-\\d{2}$'}, {'type': 'null'}]}, 'context_length': {'anyOf': [{'type': 'integer', 'maximum': 9007199254740991, 'exclusiveMinimum': 0}, {'type': 'null'}], 'description': "The model's default context window in tokens."}, 'total_params_b': {'type': 'number', 'description': 'Total parameters in billions. Sets memory residency.', 'exclusiveMinimum': 0}, 'active_params_b': {'type': 'number', 'description': 'Parameters read per token in billions. Equals total for dense models. Sets speed, never memory.', 'exclusiveMinimum': 0}}, 'additionalProperties': False}, 'sources': {'type': 'array', 'items': {'type': 'string', 'pattern': '^https:\\/\\/'}, 'maxItems': 8}, 'summary': {'type': 'string', 'maxLength': 600, 'description': 'One or two plain-language sentences stating the answer.'}, 'data_as_of': {'type': 'string', 'pattern': '^\\d{4}-\\d{2}-\\d{2}$', 'description': 'Date the bundled catalogue was last updated. Not the time of this call.'}, 'quantisations': {'type': 'array', 'items': {'type': 'object', 'required': ['quantisation', 'bits_per_weight', 'weights_gb', 'confidence'], 'properties': {'confidence': {'enum': ['measured', 'estimated', 'community'], 'type': 'string'}, 'weights_gb': {'type': 'number', 'minimum': 0}, 'quantisation': {'enum': ['Q2_K', 'Q3_K_M', 'Q4_K_M', 'Q5_K_M', 'Q6_K', 'Q8_0', 'F16'], 'type': 'string'}, 'bits_per_weight': {'type': 'number', 'exclusiveMinimum': 0}}, 'additionalProperties': False}, 'description': 'Weight size per quantisation. Excludes KV cache and overhead, which depend on context length.'}}, 'additionalProperties': False}
Schéma d’entrée
{'type': 'object', 'required': ['hardware_id'], 'properties': {'limit': {'type': 'integer', 'default': 10, 'maximum': 20, 'minimum': 1, 'description': 'Maximum models, 1 to 20. Default 10.'}, 'use_case': {'enum': ['coding', 'chat', 'reasoning', 'agents', 'vision'], 'type': 'string', 'description': 'Only list models tagged for this use. One of coding, chat, reasoning, agents, vision. Omit for all models.'}, 'hardware_id': {'type': 'string', 'pattern': '^[a-z0-9][a-z0-9._-]*$', 'maxLength': 80, 'minLength': 1, 'description': 'Hardware id from search_catalog.'}, 'quantisation': {'enum': ['Q2_K', 'Q3_K_M', 'Q4_K_M', 'Q5_K_M', 'Q6_K', 'Q8_0', 'F16'], 'type': 'string', 'default': 'Q4_K_M', 'description': 'Weight quantisation. One of Q2_K, Q3_K_M, Q4_K_M, Q5_K_M, Q6_K, Q8_0, F16. Default Q4_K_M.'}, 'context_length': {'type': 'integer', 'default': 4096, 'maximum': 262144, 'minimum': 512, 'description': 'Context window in tokens that the KV cache is sized for. 512 to 262144. Default 4096.'}}}
Schéma de sortie
{'type': 'object', 'required': ['data_as_of', 'summary', 'quantisation', 'context_length', 'use_case', 'hardware', 'models', 'url'], 'properties': {'url': {'type': 'string', 'pattern': '^https:\\/\\/llmconfigurator\\.com\\/en\\/'}, 'models': {'type': 'array', 'items': {'type': 'object', 'required': ['rank', 'model_id', 'name', 'verdict', 'needed_gb', 'speed', 'url'], 'properties': {'url': {'type': 'string', 'pattern': '^https:\\/\\/llmconfigurator\\.com\\/en\\/'}, 'name': {'type': 'string', 'maxLength': 120}, 'rank': {'type': 'integer', 'maximum': 9007199254740991, 'exclusiveMinimum': 0}, 'speed': {'anyOf': [{'type': 'object', 'required': ['low_tok_s', 'high_tok_s', 'confidence', 'basis', 'sample_count', 'last_verified'], 'properties': {'basis': {'anyOf': [{'type': 'string', 'maxLength': 200}, {'type': 'null'}], 'description': 'Where the figure comes from, e.g. a benchmark label. null for a modelled estimate.'}, 'low_tok_s': {'type': 'number', 'minimum': 0}, 'confidence': {'enum': ['measured', 'estimated', 'community'], 'type': 'string'}, 'high_tok_s': {'type': 'number', 'minimum': 0}, 'sample_count': {'anyOf': [{'type': 'integer', 'maximum': 9007199254740991, 'exclusiveMinimum': 0}, {'type': 'null'}], 'description': 'Number of measured runs behind a measured or community figure.'}, 'last_verified': {'anyOf': [{'type': 'string', 'pattern': '^\\d{4}-\\d{2}-\\d{2}$'}, {'type': 'null'}]}}, 'description': 'Decode speed in tokens per second, always a range with a confidence label.', 'additionalProperties': False}, {'type': 'null'}]}, 'verdict': {'enum': ['fits_comfortably', 'fits_tight', 'needs_offload', 'does_not_fit'], 'type': 'string'}, 'model_id': {'type': 'string', 'pattern': '^[a-z0-9][a-z0-9._-]*$', 'maxLength': 80, 'minLength': 1}, 'needed_gb': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}}, 'summary': {'type': 'string', 'maxLength': 600, 'description': 'One or two plain-language sentences stating the answer.'}, 'hardware': {'type': 'object', 'required': ['id', 'name', 'memory_gb', 'memory_kind', 'bandwidth_gb_s', 'status'], 'properties': {'id': {'type': 'string', 'pattern': '^[a-z0-9][a-z0-9._-]*$', 'maxLength': 80, 'minLength': 1}, 'name': {'type': 'string', 'maxLength': 120}, 'status': {'enum': ['current', 'eol', 'used-market', 'announced'], 'type': 'string', 'description': 'announced means not yet on sale; eol means no longer made.'}, 'memory_gb': {'type': 'number', 'exclusiveMinimum': 0}, 'memory_kind': {'enum': ['dedicated', 'unified', 'shared-system'], 'type': 'string'}, 'bandwidth_gb_s': {'type': 'number', 'description': 'Theoretical peak memory bandwidth.', 'exclusiveMinimum': 0}}, 'additionalProperties': False}, 'use_case': {'anyOf': [{'enum': ['coding', 'chat', 'reasoning', 'agents', 'vision'], 'type': 'string'}, {'type': 'null'}], 'description': 'The use-case filter applied, or null.'}, 'data_as_of': {'type': 'string', 'pattern': '^\\d{4}-\\d{2}-\\d{2}$', 'description': 'Date the bundled catalogue was last updated. Not the time of this call.'}, 'quantisation': {'enum': ['Q2_K', 'Q3_K_M', 'Q4_K_M', 'Q5_K_M', 'Q6_K', 'Q8_0', 'F16'], 'type': 'string'}, 'context_length': {'type': 'integer', 'maximum': 9007199254740991, 'minimum': -9007199254740991}}, 'additionalProperties': False}
Schéma d’entrée
{'type': 'object', 'required': ['query', 'kind'], 'properties': {'kind': {'enum': ['hardware', 'model'], 'type': 'string', 'description': 'Search hardware or models. Required.'}, 'limit': {'type': 'integer', 'default': 5, 'maximum': 10, 'minimum': 1, 'description': 'Maximum matches, 1 to 10. Default 5.'}, 'query': {'type': 'string', 'pattern': '^[^\\u0000-\\u001f\\u007f]*$', 'maxLength': 80, 'minLength': 1, 'description': 'Hardware or model name as the user wrote it, e.g. "4090", "m4 max 64", "llama 3.3 70b".'}}}
Schéma de sortie
{'type': 'object', 'required': ['data_as_of', 'summary', 'kind', 'matches'], 'properties': {'kind': {'enum': ['hardware', 'model'], 'type': 'string'}, 'matches': {'type': 'array', 'items': {'type': 'object', 'required': ['id', 'name', 'kind', 'detail', 'url'], 'properties': {'id': {'type': 'string', 'pattern': '^[a-z0-9][a-z0-9._-]*$', 'maxLength': 80, 'minLength': 1, 'description': 'Pass this id to the other tools.'}, 'url': {'type': 'string', 'pattern': '^https:\\/\\/llmconfigurator\\.com\\/en\\/'}, 'kind': {'enum': ['hardware', 'model'], 'type': 'string'}, 'name': {'type': 'string', 'maxLength': 120}, 'detail': {'type': 'string', 'maxLength': 160, 'description': 'A short factual line, e.g. "24 GB, 1008 GB/s" or "8B dense".'}}, 'additionalProperties': False}}, 'summary': {'type': 'string', 'maxLength': 600, 'description': 'One or two plain-language sentences stating the answer.'}, 'data_as_of': {'type': 'string', 'pattern': '^\\d{4}-\\d{2}-\\d{2}$', 'description': 'Date the bundled catalogue was last updated. Not the time of this call.'}}, 'additionalProperties': False}
Modifications récentes des outils
Serveurs MCP similaires
Andreax
Offers pay-per-call AI services for inference, agent and workflow design, OCR and transcription, code generation and review, clas…
IA-QA — 130+ QA & Dev Tools for AI Agents
Provides deterministic QA, evaluation, testing, code analysis, prompt and RAG checks, model comparison, and web security diagnost…
CompletionKit
Runs prompt evaluation workflows over datasets using deterministic checks and LLM judges, with metrics, scoring runs, agreements,…
Replicate
Provides access to Replicate models, versions, collections, hardware, predictions, and deployments for running and managing hoste…
Huggingface
Provides access to Hugging Face model, dataset, and Space metadata, alongside broader structured research and data-routing tools.
Ai Model Experiments
Runs prompts across multiple AI models and compares their outputs, costs, latency, token usage, and errors.
rubrkit
Manages AI evaluation artifacts, rubric audits, eval runs, golden cases, proof reports, and drift monitoring for LLM outputs.
Hipocampo MCP
Provides bilingual semantic memory, embeddings, graph links, code indexing and search, context preload, deduplication, compressio…