Servidor MCP

Nodegrove VRAM: can I run it?

io.nodegrove/vram-mcp
IA y agentes Herramientas para desarrolladores Público y accesible MCP 2025-11-25

Qué hace este MCP

Estimates LLM VRAM requirements and determines which open-weight models fit specified GPUs, quantizations, and context lengths.

can_i_run
Can I run it?
Can this GPU run this open-weight LLM? Returns fits, tight or no, the memory split (weights, KV cache, overhead), a decode-speed ceiling, the longest context that fits and, on a no, every change that would make it fit: quantisation, KV cache, context, another card or a smaller model. Model: a name or id from list_models, any Hugging Face repo id, or its architecture. GPU: a name or id from list_gpus, or vram_gb for any other card.
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', '$schema': 'https://json-schema.org/draft/2020-12/schema', 'properties': {'gpu': {'type': 'string', 'maxLength': 100, 'minLength': 1, 'description': 'A GPU from list_gpus (id or name, e.g. "rtx-4090", "4090" or "M4 Max").'}, 'model': {'type': 'string', 'maxLength': 200, 'minLength': 1, 'description': 'A model from list_models (id or name, e.g. "llama-3.3-70b" or "Llama 3.3 70B"), or any Hugging Face repo id (e.g. "Qwen/Qwen3-8B"), read live from its config.json.'}, 'quant': {'enum': ['fp16', 'q8', 'q6', 'q5', 'q4', 'q3'], 'type': 'string', 'default': 'q4', 'description': 'Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.'}, 'context': {'type': 'integer', 'default': 8192, 'maximum': 10000000, 'minimum': 1, 'description': 'Tokens held in context: prompt plus conversation.'}, 'vram_gb': {'type': 'number', 'maximum': 4096, 'description': 'Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.', 'exclusiveMinimum': 0}, 'kv_cache': {'enum': ['fp16', 'q8'], 'type': 'string', 'default': 'fp16', 'description': 'KV cache precision. fp16 is what most runtimes use; q8 halves the cache.'}, 'architecture': {'type': 'object', 'required': ['params_b', 'layers', 'kv_heads', 'head_dim'], 'properties': {'layers': {'type': 'integer', 'maximum': 1000, 'description': 'num_hidden_layers', 'exclusiveMinimum': 0}, 'head_dim': {'type': 'integer', 'maximum': 4096, 'description': 'head_dim, or hidden_size ÷ num_attention_heads', 'exclusiveMinimum': 0}, 'kv_heads': {'type': 'integer', 'maximum': 1024, 'description': 'num_key_value_heads', 'exclusiveMinimum': 0}, 'params_b': {'type': 'number', 'maximum': 10000, 'description': 'Total parameters, billions; all experts for a mixture-of-experts model.', 'exclusiveMinimum': 0}, 'kv_groups': {'type': 'array', 'items': {'type': 'object', 'required': ['layers', 'values_per_token'], 'properties': {'layers': {'type': 'integer', 'maximum': 9007199254740991, 'exclusiveMinimum': 0}, 'window_tokens': {'type': 'integer', 'maximum': 9007199254740991, 'description': 'Sliding window: these layers keep only this many tokens.', 'exclusiveMinimum': 0}, 'values_per_token': {'type': 'number', 'description': 'Values each layer caches per token: 2 Ã\x97 KV heads Ã\x97 head dim, or the latent width for MLA.', 'exclusiveMinimum': 0}}}, 'maxItems': 8, 'description': 'Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers Ã\x97 kv_heads Ã\x97 head_dim.'}, 'fixed_state_gb': {'type': 'number', 'maximum': 100, 'minimum': 0, 'description': 'Fixed recurrent state of linear-attention or Mamba layers, GB.'}, 'native_context': {'type': 'integer', 'maximum': 9007199254740991, 'description': 'The context window the model supports, tokens.', 'exclusiveMinimum': 0}, 'active_params_b': {'type': 'number', 'maximum': 10000, 'description': 'Parameters read per token, billions (mixture-of-experts only).', 'exclusiveMinimum': 0}}, 'description': 'A model described by its config.json values instead of a name.'}, 'apple_silicon': {'type': 'boolean', 'description': 'vram_gb is Apple unified memory; the GPU can use about 75% of it by default.'}, 'bandwidth_gb_s': {'type': 'number', 'maximum': 100000, 'description': "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.", 'exclusiveMinimum': 0}, 'active_params_b': {'type': 'number', 'maximum': 10000, 'description': 'Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.', 'exclusiveMinimum': 0}}}
estimate_from_hf_repo
Estimate VRAM from a Hugging Face repo
Reads any Hugging Face model repo's config.json and parameter count and estimates its memory: the attention layout found (standard, sliding-window, hybrid or latent), how much each 1,000 tokens of context costs, and weights + KV cache + overhead at every quantisation. For models nodegrove.io has not reviewed; anything the reader cannot model is listed in warnings.
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', '$schema': 'https://json-schema.org/draft/2020-12/schema', 'required': ['repo'], 'properties': {'repo': {'type': 'string', 'maxLength': 200, 'minLength': 3, 'description': 'Hugging Face repo id, e.g. "Qwen/Qwen3-8B", or its huggingface.co URL.'}, 'context': {'type': 'integer', 'default': 8192, 'maximum': 10000000, 'minimum': 1, 'description': 'Tokens held in context: prompt plus conversation.'}, 'kv_cache': {'enum': ['fp16', 'q8'], 'type': 'string', 'default': 'fp16', 'description': 'KV cache precision. fp16 is what most runtimes use; q8 halves the cache.'}, 'active_params_b': {'type': 'number', 'maximum': 10000, 'description': 'Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.', 'exclusiveMinimum': 0}}}
estimate_vram
Estimate VRAM
How much memory an LLM needs: weights + KV cache + overhead at each quantisation (or one), at a given context, and the smallest common card class that holds each. Model: a name or id from list_models, any Hugging Face repo id, or its architecture (params_b, layers, kv_heads, head_dim).
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', '$schema': 'https://json-schema.org/draft/2020-12/schema', 'properties': {'model': {'type': 'string', 'maxLength': 200, 'minLength': 1, 'description': 'A model from list_models (id or name, e.g. "llama-3.3-70b" or "Llama 3.3 70B"), or any Hugging Face repo id (e.g. "Qwen/Qwen3-8B"), read live from its config.json.'}, 'quant': {'enum': ['fp16', 'q8', 'q6', 'q5', 'q4', 'q3'], 'type': 'string', 'description': 'Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). Omit it for all 6.'}, 'context': {'type': 'integer', 'default': 8192, 'maximum': 10000000, 'minimum': 1, 'description': 'Tokens held in context: prompt plus conversation.'}, 'kv_cache': {'enum': ['fp16', 'q8'], 'type': 'string', 'default': 'fp16', 'description': 'KV cache precision. fp16 is what most runtimes use; q8 halves the cache.'}, 'architecture': {'type': 'object', 'required': ['params_b', 'layers', 'kv_heads', 'head_dim'], 'properties': {'layers': {'type': 'integer', 'maximum': 1000, 'description': 'num_hidden_layers', 'exclusiveMinimum': 0}, 'head_dim': {'type': 'integer', 'maximum': 4096, 'description': 'head_dim, or hidden_size ÷ num_attention_heads', 'exclusiveMinimum': 0}, 'kv_heads': {'type': 'integer', 'maximum': 1024, 'description': 'num_key_value_heads', 'exclusiveMinimum': 0}, 'params_b': {'type': 'number', 'maximum': 10000, 'description': 'Total parameters, billions; all experts for a mixture-of-experts model.', 'exclusiveMinimum': 0}, 'kv_groups': {'type': 'array', 'items': {'type': 'object', 'required': ['layers', 'values_per_token'], 'properties': {'layers': {'type': 'integer', 'maximum': 9007199254740991, 'exclusiveMinimum': 0}, 'window_tokens': {'type': 'integer', 'maximum': 9007199254740991, 'description': 'Sliding window: these layers keep only this many tokens.', 'exclusiveMinimum': 0}, 'values_per_token': {'type': 'number', 'description': 'Values each layer caches per token: 2 Ã\x97 KV heads Ã\x97 head dim, or the latent width for MLA.', 'exclusiveMinimum': 0}}}, 'maxItems': 8, 'description': 'Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers Ã\x97 kv_heads Ã\x97 head_dim.'}, 'fixed_state_gb': {'type': 'number', 'maximum': 100, 'minimum': 0, 'description': 'Fixed recurrent state of linear-attention or Mamba layers, GB.'}, 'native_context': {'type': 'integer', 'maximum': 9007199254740991, 'description': 'The context window the model supports, tokens.', 'exclusiveMinimum': 0}, 'active_params_b': {'type': 'number', 'maximum': 10000, 'description': 'Parameters read per token, billions (mixture-of-experts only).', 'exclusiveMinimum': 0}}, 'description': 'A model described by its config.json values instead of a name.'}, 'active_params_b': {'type': 'number', 'maximum': 10000, 'description': 'Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.', 'exclusiveMinimum': 0}}}
list_gpus
List GPUs
The GPUs and machines nodegrove.io covers: memory, the memory a runtime can use and bandwidth, from the makers' specs, with each one's page.
Solo lectura Idempotente
Esquema de entrada
{'type': 'object', '$schema': 'https://json-schema.org/draft/2020-12/schema', 'properties': {'search': {'type': 'string', 'maxLength': 100, 'description': 'Words to filter by, e.g. "qwen" or "24 GB".'}}}
list_models
List models
The open-weight LLMs nodegrove.io has verified against their config.json (data version 2026-09-25): id, size, attention design, native context, licence, memory at Q4 with 8k context and each model's page.
Solo lectura Idempotente
Esquema de entrada
{'type': 'object', '$schema': 'https://json-schema.org/draft/2020-12/schema', 'properties': {'search': {'type': 'string', 'maxLength': 100, 'description': 'Words to filter by, e.g. "qwen" or "24 GB".'}}}
what_fits
What fits my GPU?
Which open-weight LLMs fit this GPU: every model in list_models checked at one quantisation and context, with a recommended everyday model (the biggest class that fits with room for context at conversational speed), the largest that fits, the best at Q8 and the first out of reach. GPU: a name or id from list_gpus, or vram_gb for any other card.
Solo lectura Idempotente
Esquema de entrada
{'type': 'object', '$schema': 'https://json-schema.org/draft/2020-12/schema', 'properties': {'gpu': {'type': 'string', 'maxLength': 100, 'minLength': 1, 'description': 'A GPU from list_gpus (id or name, e.g. "rtx-4090", "4090" or "M4 Max").'}, 'quant': {'enum': ['fp16', 'q8', 'q6', 'q5', 'q4', 'q3'], 'type': 'string', 'default': 'q4', 'description': 'Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.'}, 'context': {'type': 'integer', 'default': 8192, 'maximum': 10000000, 'minimum': 1, 'description': 'Tokens held in context: prompt plus conversation.'}, 'vram_gb': {'type': 'number', 'maximum': 4096, 'description': 'Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.', 'exclusiveMinimum': 0}, 'apple_silicon': {'type': 'boolean', 'description': 'vram_gb is Apple unified memory; the GPU can use about 75% of it by default.'}, 'bandwidth_gb_s': {'type': 'number', 'maximum': 100000, 'description': "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.", 'exclusiveMinimum': 0}}}
Añadido
list_gpus
4 de October de 2026 a las 02:40
Añadido
list_models
4 de October de 2026 a las 02:40
Añadido
estimate_from_hf_repo
4 de October de 2026 a las 02:40
Añadido
estimate_vram
4 de October de 2026 a las 02:40
Añadido
what_fits
4 de October de 2026 a las 02:40
Añadido
can_i_run
4 de October de 2026 a las 02:40