MCP 服务器

FitLLM

run.fitllm/fitllm
开发者工具 公开且可连接 MCP 2025-11-25

此 MCP 可以做什么

Estimates whether local language models fit on specified GPUs or Apple Silicon systems based on memory and context requirements.

check_llm_fit
Check if an LLM fits on hardware
Check whether a specific local LLM fits in the memory of a specific GPU or Apple Silicon Mac. Returns fits/tight/won't-fit verdict with the memory breakdown (weights, KV cache, linear-attention state when present, runtime overhead, reserve), max context, and a concrete fix if it doesn't fit. Use this whenever a user asks anything like "can I run <model> on my <GPU/Mac>?", "will <model> fit in <N>GB?", or "what do I need to run <model>?". Estimates using curated, config-derived architecture fields (MLA, sliding-window, hybrid attention, MoE modeled).
只读 幂等
输入模式
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['model'], 'properties': {'ctx': {'type': 'integer', 'minimum': 1024, 'description': 'Alias of context_tokens â\x80\x94 accepted because the REST API uses this name. Do not pass both with different values.'}, 'gpu': {'type': 'string', 'description': 'GPU name, fuzzy â\x80\x94 e.g. "RTX 4090", "RX 7900 XTX", "A100 80GB". Multi-GPU rigs: join with + â\x80\x94 e.g. "RTX 5090 + RTX 3090" (VRAM pools across cards). Provide gpu OR mac_ram_gb.'}, 'model': {'type': 'string', 'description': 'LLM name, fuzzy â\x80\x94 e.g. "GLM-4.7-Flash", "gpt-oss-20b", "gemma 31b"'}, 'quant': {'type': 'string', 'description': 'Weight quantization. GPU: Q4_K_M(default)/Q5_K_M/Q6_K/Q8_0/FP16. Mac: 4/8(default)/16 (bits).'}, 'kv_bits': {'enum': [16, 8, 4], 'type': 'number', 'description': 'KV-cache quantization bits (default 16 = F16)'}, 'gpu_count': {'type': 'integer', 'maximum': 8, 'minimum': 1, 'description': 'Number of identical copies of the gpu (e.g. gpu="RTX 3090", gpu_count=2 for a 2Ã\x973090 rig). Default 1.'}, 'mac_ram_gb': {'type': 'integer', 'maximum': 2048, 'minimum': 8, 'description': 'Apple Silicon unified memory in GB â\x80\x94 e.g. 16, 64, 512. Provide gpu OR mac_ram_gb.'}, 'context_tokens': {'type': 'integer', 'minimum': 1024, 'description': 'Context length in tokens (default 8192). Alias: ctx (same field as the REST API).'}}, 'additionalProperties': False}
list_supported
List supported models & hardware
List the built-in model names and hardware names this fit-checker knows (for mapping user wording to exact names). Standard text-only HuggingFace transformer configs can also be checked via fitllm.run; unsupported architectures are rejected.
只读 幂等
输入模式
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {}}
what_fits_on_hardware
What LLMs fit on this hardware
Rank which popular local LLMs fit on a given GPU or Apple Silicon Mac (at ~4-bit quantization, 8K context) — models that fit come first, biggest first, with max context each. Use when a user asks "what can I run on my <GPU/Mac/N GB>?", "best local model for my machine?", or gives hardware without naming a model.
只读 幂等
输入模式
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'gpu': {'type': 'string', 'description': 'GPU name, fuzzy. Multi-GPU rigs: join with + (e.g. "RTX 5090 + RTX 3090"). Provide gpu OR mac_ram_gb.'}, 'gpu_count': {'type': 'integer', 'maximum': 8, 'minimum': 1, 'description': 'Number of identical copies of the gpu. Default 1.'}, 'mac_ram_gb': {'type': 'integer', 'maximum': 2048, 'minimum': 8, 'description': 'Apple Silicon unified memory GB. Provide gpu OR mac_ram_gb.'}}, 'additionalProperties': False}
已添加
list_supported
2026年9月17日 12:54
已添加
what_fits_on_hardware
2026年9月17日 12:54
已添加
check_llm_fit
2026年9月17日 12:54