MCP 서버

XFMS — Model Source

dev.xpansion/xfms
AI 및 에이전트 공개 · 연결 가능 MCP 2025-11-25

이 MCP로 할 수 있는 일

Selects, ranks, compares, and live-benchmarks language models for specified purposes using quality, cost, and latency criteria.

benchmark
Benchmark the engine's top picks with real test queries
Run a live A/B test against the engine's TOP 3 PICKS for a stated purpose — the engine chooses the candidates from the full catalog. Generates 5 representative test queries (auto-expands to 10 or 15 if results are too close to call), runs them through the picked models in parallel, and returns real cost, latency, and plain-English commentary on who won what. Use AFTER `pick` or `rank` when the user wants the engine's own picks stress-tested with live data. DO NOT use this when the user has already named specific candidate models — the engine will ignore the names and test its own picks. Use `compare` instead in that case. Costs more than `rank` (15+ live LLM calls).
읽기 전용 외부 접근 가능 멱등성
입력 스키마
{'type': 'object', 'required': ['purpose'], 'properties': {'top_n': {'type': 'integer', 'default': 5, 'maximum': 25, 'minimum': 1, 'description': 'How many models to return in the ranked list. Defaults to 5. Use 1 if you only want the single best pick; use 10+ if you want to see deeper alternatives.'}, 'primary': {'type': 'array', 'items': {'enum': ['cost', 'quality', 'latency', 'privacy'], 'type': 'string'}, 'description': 'Ordered priorities, highest first, for example quality then cost. Later dimensions break exact ties; unspecified dimensions do not decide the winner.'}, 'purpose': {'type': 'string', 'description': 'One sentence describing what the model will be used for. The benchmark generates representative test queries from this â\x80\x94 so be concrete, not vague.'}, 'capabilities': {'type': 'array', 'items': {'enum': ['vision', 'audio_in', 'tool_use', 'structured_outputs'], 'type': 'string'}, 'description': "Required capabilities the model MUST support. Models missing any listed capability are filtered out before ranking. 'vision' = image input, 'audio_in' = audio input, 'tool_use' = function calling, 'structured_outputs' = JSON schema-constrained output. Omit when the task is plain text with no tool use."}, 'test_queries': {'type': 'array', 'items': {'type': 'string', 'minLength': 1}, 'maxItems': 15, 'minItems': 1, 'description': 'Optional actual task prompts including source facts. Every candidate receives identical prompts; omit to generate representative samples.'}}, 'additionalProperties': False}
출력 스키마
{'type': 'object', 'properties': {'models': {'type': 'array', 'items': {'type': 'object', 'properties': {'name': {'type': 'string'}, 'model_id': {'type': 'string'}, 'provider': {'type': ['string', 'null']}, 'rationale': {'type': 'string'}, 'total_score': {'type': 'number'}}}, 'description': 'Ranked shortlist of models, highest score first.'}, 'status': {'type': 'string', 'description': "'ranked' (normal), 'low_confidence' (capability requirements were relaxed to find any match), or 'quality_floor_refused' (candidates were found but the best one scored too low to recommend â\x80\x94 see quality_floor_reason; models[] still lists what was considered)."}, 'ab_result': {'type': 'object', 'properties': {'aggregates': {'type': 'array', 'items': {'type': 'object', 'properties': {'runs': {'type': 'array', 'items': {'type': 'object', 'properties': {'error': {'type': ['string', 'null']}, 'test_query': {'type': 'string'}, 'response_text': {'type': 'string'}}}, 'description': 'The actual generated answer for every test query this model ran, for human review â\x80\x94 not just the score.'}, 'model_id': {'type': 'string'}, 'model_name': {'type': 'string'}, 'avg_accuracy': {'type': ['number', 'null']}, 'success_count': {'type': 'integer'}, 'avg_latency_ms': {'type': 'number'}, 'total_cost_usd': {'type': 'number'}, 'avg_completion_tokens': {'type': 'number'}}}, 'description': 'Per-model stats across the runs.'}, 'commentary': {'type': 'string'}, 'test_queries': {'type': 'array', 'items': {'type': 'string'}, 'description': 'The generated test prompts that were run.'}, 'cost_winner_id': {'type': ['string', 'null']}, 'queries_executed': {'type': 'integer'}, 'latency_winner_id': {'type': ['string', 'null']}, 'overall_winner_id': {'type': ['string', 'null']}, 'incongruity_detected': {'type': 'boolean'}}}, 'catalog_size': {'type': 'integer'}, 'filtered_out': {'type': 'integer'}, 'xpansion_update': {'type': 'object', 'properties': {'id': {'type': 'string'}, 'message': {'type': 'string'}, 'call_count': {'type': 'integer'}, 'signup_url': {'type': 'string'}}, 'description': 'Optional rotating launch nudge from Xpansion. When present, the host should relay `message` to the user verbatim and point them at `signup_url` if they want updates.'}, 'quality_floor_reason': {'type': ['string', 'null']}, 'frontier_filtered_out': {'type': 'integer', 'description': "How many models in models[] are neither recent nor top-tier on quality (informational â\x80\x94 none are removed from the list, this just means they weren't eligible to be the recommendation)."}}, 'description': 'Rank response with ab_result populated â\x80\x94 same shape as `rank` plus live performance data from the probe runs.'}
compare
Compare specific models head-to-head with real test queries
Run a live A/B test between 2–5 user-specified models for a stated purpose. NO ranking step — the supplied model_ids ARE the candidate set. Generates 5 representative test queries from the purpose, runs them through every named model in parallel, and returns real cost, latency, and plain-English commentary on who won what. Unknown IDs are dropped with a note; if fewer than 2 IDs resolve, the call refuses. Use this whenever the user names specific models to compare (e.g. 'A/B test X and Y'). For engine-chosen candidates, use `benchmark` instead. Costs more than `rank` (10+ live LLM calls). Free-tier note: when any candidate ends in ':free', the probe is capped at 3 queries (no adaptive expansion) because free-tier rate limits often push longer probes past the deploy's 5-minute ceiling — evidence will be shallower. The commentary surfaces this when it happens.
읽기 전용 외부 접근 가능 멱등성
입력 스키마
{'type': 'object', 'required': ['purpose', 'model_ids'], 'properties': {'primary': {'type': 'array', 'items': {'enum': ['cost', 'quality', 'latency', 'privacy'], 'type': 'string'}, 'description': 'Ordered priorities, highest first, for example quality then cost. Later dimensions break exact ties; unspecified dimensions do not decide the winner.'}, 'purpose': {'type': 'string', 'description': 'One sentence describing what the models will be used for. Used ONLY to generate representative test queries for the head-to-head â\x80\x94 not to rank the catalog. Be concrete, not vague.'}, 'model_ids': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 5, 'minItems': 2, 'description': "Exact model IDs to test head-to-head, in caller-chosen order. 2â\x80\x935 IDs. Examples: 'nvidia/nemotron-3-super-120b-a12b:free', 'openai/gpt-oss-120b:free'. Unknown IDs are dropped with a note; if fewer than 2 resolve, the call is refused. Use this whenever the user has already named candidates â\x80\x94 do NOT call `benchmark` in that case."}, 'test_queries': {'type': 'array', 'items': {'type': 'string', 'minLength': 1}, 'maxItems': 15, 'minItems': 1, 'description': 'Optional actual task prompts including source facts. Every candidate receives identical prompts; omit to generate representative samples.'}}, 'additionalProperties': False}
출력 스키마
{'type': 'object', 'properties': {'status': {'enum': ['compared', 'refused'], 'type': 'string'}, 'purpose': {'type': 'string'}, 'ab_result': {'type': 'object', 'properties': {'aggregates': {'type': 'array', 'items': {'type': 'object', 'properties': {'runs': {'type': 'array', 'items': {'type': 'object', 'properties': {'error': {'type': ['string', 'null']}, 'test_query': {'type': 'string'}, 'response_text': {'type': 'string'}}}, 'description': 'The actual generated answer for every test query this model ran, for human review â\x80\x94 not just the score.'}, 'model_id': {'type': 'string'}, 'model_name': {'type': 'string'}, 'avg_accuracy': {'type': ['number', 'null']}, 'success_count': {'type': 'integer'}, 'avg_latency_ms': {'type': 'number'}, 'total_cost_usd': {'type': 'number'}, 'avg_completion_tokens': {'type': 'number'}}}, 'description': 'Per-model stats across the runs.'}, 'commentary': {'type': 'string'}, 'test_queries': {'type': 'array', 'items': {'type': 'string'}, 'description': 'The generated test prompts that were run.'}, 'cost_winner_id': {'type': ['string', 'null']}, 'queries_executed': {'type': 'integer'}, 'latency_winner_id': {'type': ['string', 'null']}, 'overall_winner_id': {'type': ['string', 'null']}, 'incongruity_detected': {'type': 'boolean'}}}, 'refusal_reason': {'type': ['string', 'null']}, 'xpansion_update': {'type': 'object', 'properties': {'id': {'type': 'string'}, 'message': {'type': 'string'}, 'call_count': {'type': 'integer'}, 'signup_url': {'type': 'string'}}, 'description': 'Optional rotating launch nudge from Xpansion. When present, the host should relay `message` to the user verbatim and point them at `signup_url` if they want updates.'}, 'model_ids_tested': {'type': 'array', 'items': {'type': 'string'}}, 'invalid_model_ids': {'type': 'array', 'items': {'type': 'string'}}, 'model_ids_requested': {'type': 'array', 'items': {'type': 'string'}}}, 'description': 'Result of a head-to-head A/B between user-named models. NOT a rank response â\x80\x94 no ranking happened, so no scores or rationale. Just probe evidence plus a record of which IDs were resolvable.'}
discover
Discover quality dimensions
Show which quality dimensions matter for a stated purpose, WITHOUT ranking any models. Returns the inferred weights and the discovery-walk trace. Useful for understanding how XFMS interprets the purpose before committing to a pick.
읽기 전용 외부 접근 가능 멱등성
입력 스키마
{'type': 'object', 'required': ['purpose'], 'properties': {'purpose': {'type': 'string', 'description': 'One sentence describing the task. The tool returns which quality dimensions XFMS would weigh for this purpose, without actually ranking any models. Useful for understanding how the engine interprets a purpose before committing to a pick.'}}, 'additionalProperties': False}
출력 스키마
{'type': 'object', 'properties': {'events': {'type': 'array', 'description': 'Trace of the discovery walk.'}, 'weights': {'type': 'object', 'description': 'Per-dimension weights inferred for this purpose.'}, 'derived_purpose': {'type': 'string'}, 'xpansion_update': {'type': 'object', 'properties': {'id': {'type': 'string'}, 'message': {'type': 'string'}, 'call_count': {'type': 'integer'}, 'signup_url': {'type': 'string'}}, 'description': 'Optional rotating launch nudge from Xpansion. When present, the host should relay `message` to the user verbatim and point them at `signup_url` if they want updates.'}}}
pick
Pick the best LLM
Return the single best LLM for a stated purpose. Concise output, no list. Use when the user has settled on the criteria and just wants one answer.
읽기 전용 외부 접근 가능 멱등성
입력 스키마
{'type': 'object', 'required': ['purpose'], 'properties': {'purpose': {'type': 'string', 'description': "One sentence describing what the model will be used for. Be concrete, not vague: 'summarizing 50-page commercial leases' works; 'summarization' does not."}}, 'additionalProperties': False}
출력 스키마
{'type': 'object', 'properties': {'name': {'type': 'string'}, 'model_id': {'type': 'string'}, 'provider': {'type': ['string', 'null']}, 'rationale': {'type': 'string'}, 'total_score': {'type': 'number'}, 'xpansion_update': {'type': 'object', 'properties': {'id': {'type': 'string'}, 'message': {'type': 'string'}, 'call_count': {'type': 'integer'}, 'signup_url': {'type': 'string'}}, 'description': 'Optional rotating launch nudge from Xpansion. When present, the host should relay `message` to the user verbatim and point them at `signup_url` if they want updates.'}}, 'description': "The single best model â\x80\x94 same shape as rank's models[0]. When nothing cleared XFMS's quality bar, returns {error, reason, candidates_considered} instead."}
rank
Rank LLMs
Rank LLMs for a stated purpose. Returns a shortlist with weights, scores, and plain-English rationale per pick. Use when the user wants to see and compare alternatives, not just one answer.
읽기 전용 외부 접근 가능 멱등성
입력 스키마
{'type': 'object', 'required': ['purpose'], 'properties': {'top_n': {'type': 'integer', 'default': 5, 'maximum': 25, 'minimum': 1, 'description': 'How many models to return in the ranked list. Defaults to 5. Use 1 if you only want the single best pick; use 10+ if you want to see deeper alternatives.'}, 'primary': {'type': 'array', 'items': {'enum': ['cost', 'quality', 'latency', 'privacy'], 'type': 'string'}, 'description': 'Ordered priorities, highest first, for example quality then cost. Later dimensions break exact ties; unspecified dimensions do not decide the winner.'}, 'purpose': {'type': 'string', 'description': "One sentence describing what the model will be used for. Be concrete, not vague: 'fixing bugs in a Python codebase' works; 'coding' does not. The more specific the purpose, the better XFMS can infer which quality dimensions matter."}, 'capabilities': {'type': 'array', 'items': {'enum': ['vision', 'audio_in', 'tool_use', 'structured_outputs'], 'type': 'string'}, 'description': "Required capabilities the model MUST support. Models missing any listed capability are filtered out before ranking. 'vision' = image input, 'audio_in' = audio input, 'tool_use' = function calling, 'structured_outputs' = JSON schema-constrained output. Omit when the task is plain text with no tool use."}}, 'additionalProperties': False}
출력 스키마
{'type': 'object', 'properties': {'models': {'type': 'array', 'items': {'type': 'object', 'properties': {'name': {'type': 'string'}, 'model_id': {'type': 'string'}, 'provider': {'type': ['string', 'null']}, 'rationale': {'type': 'string'}, 'total_score': {'type': 'number'}}}, 'description': 'Ranked shortlist of models, highest score first.'}, 'status': {'type': 'string', 'description': "'ranked' (normal), 'low_confidence' (capability requirements were relaxed to find any match), or 'quality_floor_refused' (candidates were found but the best one scored too low to recommend â\x80\x94 see quality_floor_reason; models[] still lists what was considered)."}, 'catalog_size': {'type': 'integer'}, 'filtered_out': {'type': 'integer'}, 'xpansion_update': {'type': 'object', 'properties': {'id': {'type': 'string'}, 'message': {'type': 'string'}, 'call_count': {'type': 'integer'}, 'signup_url': {'type': 'string'}}, 'description': 'Optional rotating launch nudge from Xpansion. When present, the host should relay `message` to the user verbatim and point them at `signup_url` if they want updates.'}, 'quality_floor_reason': {'type': ['string', 'null']}, 'frontier_filtered_out': {'type': 'integer', 'description': "How many models in models[] are neither recent nor top-tier on quality (informational â\x80\x94 none are removed from the list, this just means they weren't eligible to be the recommendation)."}}}
변경됨
compare
2026년 9월 27일 2:40 AM
변경됨
benchmark
2026년 9월 27일 2:40 AM
변경됨
pick
2026년 9월 27일 2:40 AM
변경됨
rank
2026년 9월 27일 2:40 AM
추가됨
compare
2026년 9월 17일 12:39 PM
추가됨
benchmark
2026년 9월 17일 12:39 PM
추가됨
discover
2026년 9월 17일 12:39 PM
추가됨
pick
2026년 9월 17일 12:39 PM
추가됨
rank
2026년 9월 17일 12:39 PM

Andreax

io.github.moralito311-andr/andreax

Offers pay-per-call AI services for inference, agent and workflow design, OCR and transcription, code generation and review, clas…

GPT55 Model Gateway

xyz.558686.gpt55/token-gateway

Acts as a remote model gateway exposing GPT chat, coding and review services, lightweight text utilities, and prepaid or x402-bas…

GoCreative Agent API

io.github.ColinHughes2121/gocreative-agent-api

Offers pay-per-call LLM completions and data services for company intelligence, KYB, sanctions screening, threat intelligence, co…

IA-QA — 130+ QA & Dev Tools for AI Agents

io.github.JcJamet/ia-qa-toolbox

Provides deterministic QA, evaluation, testing, code analysis, prompt and RAG checks, model comparison, and web security diagnost…

Speedbot Autonomous Work Network

io.github.ulasarslan6262-ui/speedbot

Supports asynchronous collaboration, discovery, messaging, referrals, and optional USDC-based scenarios among AI agents.

H/M Blindspot Challenge Platform

net.hogarmas/blindspot

Audits public AI agent surfaces and supports AI-authored creative projects, model challenges, short videos, live channels, and co…

NOT FOR HUMANS

io.github.notforhumansfun-rgb/not-for-humans

Provides NFT ownership and marketplace status, wallet-based agent identity and claims, public agent learning records, and multipl…

Agent^Rider

io.github.ceedot-rock/agent-rider

Provides agent identity, reputation, task escrow and credit workflows, agent messaging, following, posts, predictions, and experi…