MCP Server

Cloud World Model

ai.cloudworldmodel/cloud-world-model
Cloud & Infrastructure Science & Engineering Public & reachable MCP 2025-11-25

What this MCP does

Creates and runs temporary cloud architecture simulations with traffic, failures, recovery, performance metrics, and cost estimates.

scenario.get
Get Scenario
Hydrate one built-in scenario from the live Cloud World Model scenario library. Prerequisite: a scenario id returned by scenario.list. Returns the complete selected scenario graph, including resources and connections plus optional seed, resilienceConfig, protectedResilienceConfig, traffic/failure presets, named traffic-phase summaries, activeFailurePhases and optionalFailurePhases (type, resource/zone target, severity, step range), retry-workload disclosure, and real-world incident metadata. The response includes both title and name for compatibility; pass resources and connections, and optionally seed/resilienceConfig, to simulation.create when you need to edit or inspect the graph. For the shorter handoff, pass the id as scenarioId instead. The likely next tool is simulation.create.
Read only Idempotent
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['scenarioId'], 'properties': {'scenarioId': {'type': 'string', 'minLength': 1, 'description': 'Scenario identifier returned by scenario.list'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'id': {'type': 'string', 'description': 'Stable scenario identifier'}, 'name': {'type': 'string', 'description': 'Scenario display name; equivalent to title'}, 'seed': {'type': 'integer', 'minimum': 0}, 'tags': {'type': 'array', 'items': {'type': 'string'}}, 'title': {'type': 'string', 'description': 'Scenario title'}, 'status': {'type': 'string', 'description': 'Result status; not_found when the requested scenario does not exist'}, 'message': {'type': 'string', 'description': 'Error or guidance message'}, 'category': {'type': 'string'}, 'duration': {'type': 'string'}, 'resources': {'type': 'array', 'items': {'type': 'object', 'additionalProperties': {}}, 'description': 'Full resource graph; pass to simulation.create'}, 'difficulty': {'type': 'string'}, 'connections': {'type': 'array', 'items': {'type': 'object', 'additionalProperties': {}}, 'description': 'Full connection graph; pass to simulation.create'}, 'description': {'type': 'string'}, 'primaryPurpose': {'enum': ['educational', 'chaos', 'predictive', 'optimization'], 'type': 'string', 'description': 'Primary purpose of the scenario; absent means legacy purpose not specified'}, 'resilienceConfig': {'type': 'object', 'additionalProperties': {}}, 'realWorldIncident': {'type': 'object', 'additionalProperties': {}}, 'activeFailurePhases': {'type': 'array', 'items': {'type': 'object', 'required': ['name', 'type', 'startStep'], 'properties': {'name': {'type': 'string'}, 'type': {'enum': ['instance_kill', 'instance_down', 'az_outage', 'region_outage', 'permanent_data_loss', 'database_overload', 'network_latency', 'spot_interruption'], 'type': 'string'}, 'endStep': {'type': 'number', 'minimum': 0}, 'isActive': {'type': 'boolean', 'default': True}, 'severity': {'enum': ['minor', 'moderate', 'severe'], 'type': 'string', 'default': 'moderate'}, 'startStep': {'type': 'number', 'minimum': 0}, 'targetZone': {'type': 'string'}, 'targetRegion': {'type': 'string'}, 'targetProvider': {'enum': ['aws', 'gcp', 'azure', 'oci', 'digitalocean'], 'type': 'string'}, 'targetResourceId': {'type': 'string'}}, 'additionalProperties': False}}, 'activeTrafficPhases': {'type': 'array', 'items': {'type': 'object', 'required': ['name', 'type', 'startStep', 'isActive', 'traffic'], 'properties': {'name': {'type': 'string'}, 'type': {'enum': ['ramp', 'burst', 'step', 'wave', 'custom'], 'type': 'string'}, 'endStep': {'type': 'number', 'minimum': 0}, 'traffic': {'type': 'object', 'properties': {'endRps': {'type': 'number', 'minimum': 0}, 'peakRps': {'type': 'number', 'minimum': 0}, 'startRps': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}, 'isActive': {'type': 'boolean'}, 'startStep': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}}, 'optionalFailurePhases': {'type': 'array', 'items': {'$ref': '#/properties/activeFailurePhases/items'}}, 'optionalTrafficPhases': {'type': 'array', 'items': {'$ref': '#/properties/activeTrafficPhases/items'}}, 'defaultTrafficPatterns': {'type': 'array', 'items': {'type': 'object', 'additionalProperties': {}}}, 'retryTrafficDisclosure': {'type': 'object', 'required': ['enabled', 'dependencyCount', 'maxConfiguredRetries', 'externalTraffic', 'internalRetryAttempts'], 'properties': {'enabled': {'type': 'boolean', 'const': True}, 'dependencyCount': {'type': 'integer', 'exclusiveMinimum': 0}, 'externalTraffic': {'type': 'string'}, 'maxConfiguredRetries': {'type': 'integer', 'exclusiveMinimum': 0}, 'internalRetryAttempts': {'type': 'string'}}, 'additionalProperties': False}, 'defaultFailureInjections': {'type': 'array', 'items': {'type': 'object', 'additionalProperties': {}}}, 'protectedResilienceConfig': {'type': 'object', 'additionalProperties': {}}}, 'additionalProperties': True}
scenario.list
List Scenarios
List the built-in demo scenarios as compact catalog cards — stable IDs, title/name, description, difficulty, tags, category, duration, provider summary, resource/connection counts, named active/optional traffic phases, and retry-workload disclosure. Use it as the first call when you want a ready-made architecture instead of designing one; the cards intentionally omit resource, connection, traffic-pattern, and failure-injection graphs. Anonymous discovery includes only scenarios with at most 10 resources so every listed card is demo-creatable. No prerequisites. Optionally narrow discovery with provider, category, and/or difficulty filters; omit them to receive the complete demo-creatable catalog. Pass a returned id as scenarioId to simulation.create for server-side expansion, or pass it to scenario.get when you need to inspect the full graph. Larger scenarios require an authenticated session. Returns named activeFailurePhases and optionalFailurePhases with type, resource/zone target, severity, and step range. No API key required. The likely next tool is scenario.get.
Read only Idempotent
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'category': {'type': 'string', 'minLength': 1, 'description': 'Only scenarios in this category, such as scaling, failure, reliability, networking, or cost'}, 'provider': {'enum': ['aws', 'gcp', 'azure', 'oci', 'digitalocean'], 'type': 'string', 'description': 'Only scenarios that include resources from this cloud provider'}, 'difficulty': {'enum': ['beginner', 'intermediate', 'advanced'], 'type': 'string', 'description': 'Only scenarios at this difficulty level'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['scenarios'], 'properties': {'scenarios': {'type': 'array', 'items': {'type': 'object', 'properties': {'id': {'type': 'string', 'description': 'Stable scenario identifier â\x80\x94 pass to scenario.get'}, 'name': {'type': 'string', 'description': 'Scenario display name; equivalent to title'}, 'tags': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Discovery tags'}, 'title': {'type': 'string', 'description': 'Scenario title'}, 'version': {'type': ['string', 'null'], 'description': 'Catalog version; null when unavailable'}, 'category': {'type': 'string', 'description': 'Scenario category'}, 'duration': {'type': 'string', 'description': 'Expected scenario duration'}, 'provider': {'type': 'string', 'description': 'Primary cloud provider'}, 'revision': {'type': ['string', 'null'], 'description': 'Catalog revision; null when unavailable'}, 'providers': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Cloud providers represented in the scenario'}, 'difficulty': {'type': 'string', 'description': 'Scenario difficulty'}, 'description': {'type': 'string', 'description': 'What this scenario demonstrates; capped at 500 characters with an ellipsis when truncated'}, 'resourceCount': {'type': 'number', 'description': 'Number of resources in the scenario graph'}, 'primaryPurpose': {'enum': ['educational', 'chaos', 'predictive', 'optimization'], 'type': 'string', 'description': 'Primary purpose: educational guided scenario, chaos injected failure, predictive capacity simulation, or optimization configuration comparison'}, 'connectionCount': {'type': 'number', 'description': 'Number of connections in the scenario graph'}, 'providerSummary': {'type': 'string', 'description': 'Compact provider summary'}, 'activeFailurePhases': {'type': 'array', 'items': {'type': 'object', 'required': ['name', 'type', 'startStep'], 'properties': {'name': {'type': 'string'}, 'type': {'enum': ['instance_kill', 'instance_down', 'az_outage', 'region_outage', 'permanent_data_loss', 'database_overload', 'network_latency', 'spot_interruption'], 'type': 'string'}, 'endStep': {'type': 'number', 'minimum': 0}, 'isActive': {'type': 'boolean', 'default': True}, 'severity': {'enum': ['minor', 'moderate', 'severe'], 'type': 'string', 'default': 'moderate'}, 'startStep': {'type': 'number', 'minimum': 0}, 'targetZone': {'type': 'string'}, 'targetRegion': {'type': 'string'}, 'targetProvider': {'enum': ['aws', 'gcp', 'azure', 'oci', 'digitalocean'], 'type': 'string'}, 'targetResourceId': {'type': 'string'}}, 'additionalProperties': False}, 'description': 'Scheduled active failures: name, type, target resource/zone, severity and step range; no parameters'}, 'activeTrafficPhases': {'type': 'array', 'items': {'type': 'object', 'required': ['name', 'type', 'startStep', 'isActive', 'traffic'], 'properties': {'name': {'type': 'string'}, 'type': {'enum': ['ramp', 'burst', 'step', 'wave', 'custom'], 'type': 'string'}, 'endStep': {'type': 'number', 'minimum': 0}, 'traffic': {'type': 'object', 'properties': {'endRps': {'type': 'number', 'minimum': 0}, 'peakRps': {'type': 'number', 'minimum': 0}, 'startRps': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}, 'isActive': {'type': 'boolean'}, 'startStep': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}, 'description': 'Named active catalog traffic phases with type, step range, and workload shape'}, 'optionalFailurePhases': {'type': 'array', 'items': {'$ref': '#/properties/scenarios/items/properties/activeFailurePhases/items'}, 'description': 'Disabled optional failure presets; not scheduled unless enabled'}, 'optionalTrafficPhases': {'type': 'array', 'items': {'$ref': '#/properties/scenarios/items/properties/activeTrafficPhases/items'}, 'description': 'Named optional catalog traffic phases; these are inactive until enabled in the workspace'}, 'retryTrafficDisclosure': {'type': 'object', 'required': ['enabled', 'dependencyCount', 'maxConfiguredRetries', 'externalTraffic', 'internalRetryAttempts'], 'properties': {'enabled': {'type': 'boolean', 'const': True}, 'dependencyCount': {'type': 'integer', 'exclusiveMinimum': 0}, 'externalTraffic': {'type': 'string'}, 'maxConfiguredRetries': {'type': 'integer', 'exclusiveMinimum': 0}, 'internalRetryAttempts': {'type': 'string'}}, 'description': 'When present, separates external catalog traffic from modeled internal retry attempts', 'additionalProperties': False}}, 'additionalProperties': True}, 'description': 'Available demo scenarios'}}, 'additionalProperties': True}
simulation.create
Create Simulation
Create a temporary anonymous demo cloud simulation from a list of resources and connections (max 2 active simulations per client, up to 10 resources; the returned simulationId is a short-lived unguessable capability that survives MCP transport teardown, but it is cleaned up when the demo lifetime expires or the simulation is deleted). No API key required for this temporary anonymous demo operation. Built-in scenario workflow: call `scenario.list` and pass a returned card's `id` as `scenarioId` to `simulation.create` for server-side graph expansion. For full control, call `scenario.get` and pass its hydrated `resources` and `connections` arrays instead. These are two alternatives — do not send `scenarioId` with `resources` or `connections`. For the catalog EKS Spot Interruption Migration scenario, you may set `scenarioOverrides: { eksSpotInterruption: { startupSeconds } }` with an integer startupSeconds from 0 through 3600 to test a different readiness deadline without copying the graph; this override requires scenarioId and is mutually exclusive with resources and connections. `scenario.list` returns graph-free cards with bounded active/optional traffic-phase and retry-workload summaries; it is not a source of resource, connection, traffic-pattern, or failure-injection graphs. Scenario traffic and failure presets are not applied automatically. Use it to start any simulation workflow — either with hydrated resources and connections from scenario.get or your own architecture. Do not use it to modify an existing simulation (use simulation.inject_traffic to change load). For the exact owned typical fit, set appWeight:'typical', location.regionKey:'us-east-2' on all four AWS nodes, one ALB with serviceFamily:'alb' and loadBalancerScheme:'internal', two m5.large compute apps with workload:'crud-typical', appRuntime:'node', appWorkerCount:2, appDbPoolSize:250, one db.r5.large MySQL with workloadDatabaseEngine:'mysql', workloadDatabaseVersion:'8.0', maxConnections:500, ALB→each app→DB connections, minInstances=maxInstances=2, autoscaling:false, traffic 20–300 RPS. Only typical-fit-20, typical-fit-100 and typical-fit-200 tuned the typical-v1-20260927c/6aa574d7ff9d3080b88b221bcd59f7d218ae37f0 fit; 300 is an independent holdout and 500 is diagnostic only. Other typical workloads are modeled, not owned. P99 has distinct per-percentile provenance: latencyP99Basis identifies the owned in-VPC internal-ALB fit, a scaled-from-fit estimate (not directly measured), or an uncalibrated generic model. On the exact healthy lean owned graph at 10–1,000 offered target RPS, P99=max(final P95, 7.021919127633514 + 0.00027013891327780484*T) ms; only 10/100/500 RPS were fit, 1,000 RPS was held out. Lean M5 scaling is not a new measurement; typical/heavy and active failures retain uncalibrated P99. predictionEvidence.latencyP99 has measured 0.65–1.35×, scaled 0.50–1.50× (beyond 1,000: 0.25–2×), or uncalibrated 0.50–2× (beyond: 0.25–3×) assumption bounds centered on final P99. These are not confidence intervals or provider measurements. Historical evidence may omit P99. latencyBasis describes the general modeled latency path; use latencyP99Basis specifically for P99. P99 is diagnostic, not scored. For compute, set characteristics.capacityRps for an explicit per-node RPS ceiling at which CPU reaches ~95%; do not use maxThroughput for that compute contract. Kubernetes rejects capacityRps: set maxThroughput for the total cluster RPS ceiling, or nodePools[].maxThroughput for per-node pool capacity. Omitted compute capacityRps uses the selected catalog tier and can intentionally produce a stressed baseline (for example, the AWS m5.large catalog denominator is 2,000 RPS); for a healthy, capacity-bounded compute experiment, declare an explicit per-node capacity such as 500 RPS. That value is an experiment control, not a universal hardware fact. For OCI flexible compute shapes, pass the documented positive integer characteristics.ocpus explicitly; VM.Standard.E4.Flex accepts 1–64 OCPUs and each OCPU maps to 2 vCPUs. OCPU count establishes capacity dimensions only, not provider-specific performance, throughput, or price. An uncounted flexible shape remains an unverified generic estimate. Check GET /api/prediction/generic-shapes for the catalog-derived generic fallback inventory. New prediction-only GCP standard capacity entries include e2-standard-2/4/8/16/32, n1-standard-1/2/4/8/16/32/64/96, and n2-standard-16/32/48/64/80/96/128; provider specifications establish vCPU/memory dimensions, not CWM performance or pricing. Version 1 predictionEvidence explicitly reports legacyGeneric at the top level and on every appCpuByResource item; its note identifies each generic resource's shape and fallback reason. For generic fixed compute, characteristics.instanceCount accepts integer 1–100 represented VMs; capacity aggregates and CPU is per VM. Do not combine it with autoscaling:true, minInstances, or maxInstances. Aurora Serverless v2 remains limited to 1 with multiAz:false or 2 with multiAz:true. For Aurora Serverless ACU limits, use characteristics.config.minCapacity/maxCapacity or flat characteristics.minCapacity/maxCapacity. The exact AWS database shape with serviceFamily: 'aurora-serverless' and size: 'db.serverless' also accepts flat characteristics.minAcu/maxAcu; those aliases are rejected elsewhere, including at the resource root or inside config. For that shape, multiAz:true with instanceCount:2 creates a separately billable reader (<writer-id>-reader) in another AZ. Inspect returned resources and metrics before using simulation.step to observe modeled failover; no AWS timing guarantee is implied. Capacity, node-bound, SKU, and autoscaling values supplied through this MCP tool are recorded as agent-supplied in the immutable normalizationReceipt; request responseMode: 'full' to inspect it. Generic GKE telemetry and recovery apply only to worker nodes; the control-plane management fee is cost-only, with no modeled control-plane CPU, API throttling, or cooldown. To bound the autoscaled fleet size, set the top-level maxInstances / minInstances parameters. If you do not set maxInstances, the engine uses the provider default — AWS 50, GCP 15, Azure/OCI/DigitalOcean 10 — which may be much larger than your intended fleet size. The response includes effectiveMaxInstances / effectiveMinInstances so you can confirm the bounds that will be enforced. For a targeted CPU HPA scale-out threshold, send the canonical autoscalingTargetCpu field in this create call (for example, autoscalingTargetCpu: 70 for GKE). The compatible aliases scaleOutCpuThreshold, scaleOutCpuPercent, and autoscaleTargetCpuPercent are also accepted; if more than one is sent, their values must agree. Every create response includes hpaAudit with the supplied field, persisted thresholds, and any provider default. For ECS Fargate CPU-only target tracking, set ecsCpuTargetTracking: true, autoscalingTargetCpu, minInstances/maxInstances, and optional scaleOutCooldownSeconds/scaleInCooldownSeconds with simulationSecondsPerStep (default 1). Inspect autoscalingConfig in the compact response or applicationAutoscalingPolicy in the full response. Latency and throughput do not trigger ECS scaling in this mode. These four TOP-LEVEL fields are simulation-wide — the engine applies one CPU threshold identically to every resource's scale decision by default. To make ONE resource scale at a different CPU target than the rest of the simulation (e.g. a GKE cluster scaling out at 60% while an EC2 fleet in the same simulation scales out at 80%), set characteristics.scaleOutCpuThreshold and/or characteristics.scaleInCpuThreshold on that specific resource instead — the per-resource value wins over the simulation-wide default for that resource only. A misnamed near-miss field nested under characteristics (e.g. targetCPUUtilizationPercentage) is rejected with a 400 explaining the correct field name — it is never silently dropped and defaulted. Responses are compact by default: id, name, status, traffic, and a per-resource summary (id, name, status, cpuPercent, routedRps, availabilityState, isRoutable, and recoveryBlockedReason when provided). Pass responseMode: 'full' to get the complete simulation object instead. During a failure workflow, lower traffic to serviceable levels before calling simulation.recover_resource, then use simulation.step until the recovered resource is healthy. Recovery progress is included when applicable: recoveryProgress.state is parked, cooling_down, or healthy, and its parkWindow/cooldown objects report totalSteps, completedSteps, remainingSteps, target, and requiredSteps. Poll simulation.step until state is healthy, then use simulation.metrics to inspect the resulting state and metrics. No prerequisites. Returns the created simulation's id, which every other simulation.* tool consumes; the new simulation also becomes this session's current simulation, so subsequent per-simulation tools may omit simulationId. The likely next tool is simulation.step to advance time. Do not call api.spec to learn the simulation workflow — the tool descriptions in this session contain everything needed. Authenticate with an API key for unlimited persistent simulations.
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'required': ['name'], 'properties': {'name': {'type': 'string', 'maxLength': 120, 'description': 'Human-readable name for the simulation'}, 'seed': {'type': 'integer', 'minimum': 0, 'description': 'Deterministic RNG seed for reproducible replays'}, 'traffic': {'type': 'number', 'default': 0, 'description': 'Initial traffic in requests per second (RPS)'}, 'appWeight': {'enum': ['lean', 'typical', 'heavy'], 'type': 'string', 'description': 'Immutable root-level workload weight; defaults to typical with appWeightDefaulted=true. Lean is measured only for the eligible owned graph; heavy is an unsupported assumption. predictionEvidence reports per-quantity provenance and assumption ranges, not confidence intervals.'}, 'resources': {'type': 'array', 'items': {'type': 'object', 'required': ['id', 'type', 'name'], 'properties': {'id': {'type': 'string', 'description': 'Unique identifier for this resource within the simulation'}, 'name': {'type': 'string', 'description': 'Display name for the resource'}, 'type': {'enum': ['compute', 'database', 'storage', 'network', 'cache', 'queue', 'kubernetes'], 'type': 'string', 'description': 'Resource category'}, 'maxAcu': {'description': 'Not accepted at resource level; use characteristics.maxAcu for the exact AWS Aurora Serverless v2 db.serverless shape.'}, 'minAcu': {'description': 'Not accepted at resource level; use characteristics.minAcu for the exact AWS Aurora Serverless v2 db.serverless shape.'}, 'location': {'type': 'object', 'required': ['regionKey'], 'properties': {'zoneKey': {'type': 'string'}, 'regionKey': {'type': 'string'}, 'localityType': {'enum': ['az', 'zone', 'availability_domain', 'fault_domain'], 'type': 'string'}, 'providerLabel': {'type': 'string'}, 'faultDomainKey': {'type': 'string'}}, 'description': "Resource region; set location.regionKey:'us-east-2' on each owned typical node (normalized to use2).", 'additionalProperties': False}, 'provider': {'enum': ['aws', 'gcp', 'azure', 'oci', 'digitalocean'], 'type': 'string', 'default': 'aws', 'description': 'Cloud provider hosting this resource'}, 'recoveryPolicy': {'type': 'object', 'properties': {'warningSteps': {'type': 'integer', 'minimum': 1}, 'criticalSteps': {'type': 'integer', 'minimum': 1}, 'failureParkSteps': {'type': 'integer', 'minimum': 1}, 'warningCpuThreshold': {'type': 'number', 'maximum': 100, 'minimum': 0}, 'criticalCpuThreshold': {'type': 'number', 'maximum': 100, 'minimum': 0}}, 'description': 'Optional recovery thresholds and cooldown lengths for this resource', 'additionalProperties': False}, 'characteristics': {'type': 'object', 'properties': {'size': {'type': 'string', 'description': "Instance or SKU size (e.g. 't3.medium', 'db.r5.large')"}, 'ocpus': {'type': 'integer', 'maximum': 126, 'description': 'OCI flexible compute shape capacity. Family-specific limits are validated against the provider shape; VM.Standard.E4.Flex accepts 1â\x80\x9364 OCPUs and 1 OCPU = 2 vCPUs. Not a performance or price claim.', 'exclusiveMinimum': 0}, 'config': {'type': 'object', 'properties': {'maxCapacity': {'type': 'number', 'exclusiveMinimum': 0}, 'minCapacity': {'type': 'number', 'exclusiveMinimum': 0}}, 'description': 'Aurora Serverless ACU bounds. Nested config uses minCapacity/maxCapacity; flat characteristics.minAcu/maxAcu aliases are accepted only on the exact AWS Aurora Serverless v2 db.serverless shape.', 'additionalProperties': True}, 'maxAcu': {'type': 'number', 'description': 'Flat maximum ACU alias only for AWS Aurora Serverless v2 size db.serverless; conflicts with maxCapacity are rejected.', 'exclusiveMinimum': 0}, 'minAcu': {'type': 'number', 'description': 'Flat minimum ACU alias only for AWS Aurora Serverless v2 size db.serverless; conflicts with minCapacity are rejected.', 'exclusiveMinimum': 0}, 'multiAz': {'type': 'boolean', 'description': 'AWS Aurora Serverless v2: true with instanceCount:2 creates a billable reader in a second AZ.'}, 'workload': {'enum': ['crud', 'crud-typical'], 'type': 'string', 'description': "Explicit 'crud-typical' on each owned typical app; 'crud' selects lean identity on the exact lean graph."}, 'appRuntime': {'type': 'string', 'description': 'Owned typical app runtime: node.'}, 'capacityGB': {'type': 'number', 'description': 'Storage capacity in GB â\x80\x94 drives per-GB snapshot/backup idle cost and stopped-instance storage cost'}, 'autoscaling': {'type': 'boolean', 'description': 'Mark this compute resource as the autoscaling primary target'}, 'capacityRps': {'type': 'number', 'description': 'Compute only: literal per-node RPS ceiling at which CPU reaches ~95%. Kubernetes rejects capacityRps; use maxThroughput for its total cluster ceiling.', 'exclusiveMinimum': 0}, 'maxCapacity': {'type': 'number', 'description': 'Aurora Serverless maximum ACU (flat form; nested config wins).', 'exclusiveMinimum': 0}, 'minCapacity': {'type': 'number', 'description': 'Aurora Serverless minimum ACU (flat form; nested config wins).', 'exclusiveMinimum': 0}, 'billingState': {'enum': ['active', 'idle', 'stopped', 'detached', 'deleted'], 'type': 'string', 'description': "Billing state: 'stopped' bills storage only, 'idle'/'detached' bill flat idle rates, 'deleted' bills nothing. Default 'active'."}, 'appDbPoolSize': {'type': 'integer', 'description': 'Owned typical: pool size 250 on each app.', 'exclusiveMinimum': 0}, 'instanceCount': {'type': 'integer', 'maximum': 100, 'minimum': 1, 'description': 'Generic fixed compute: represented VM count (integer 1â\x80\x93100); capacity aggregates and CPU is per VM. Cannot combine with autoscaling:true, minInstances, or maxInstances. AWS Aurora Serverless v2 remains 1 with multiAz:false or 2 with multiAz:true.'}, 'maxThroughput': {'type': 'number', 'description': 'Kubernetes: total cluster RPS ceiling; compute: legacy internal throughput scaling parameter (prefer capacityRps for compute).'}, 'serviceFamily': {'type': 'string', 'description': "Idle-billed service family (e.g. 'nat-gateway', 'vpc-endpoint', 'eip-detached', 'ebs-snapshot') â\x80\x94 determines the flat idle rate"}, 'appWorkerCount': {'type': 'integer', 'description': 'Owned typical: 2 workers on each app.', 'exclusiveMinimum': 0}, 'maxConnections': {'type': 'integer', 'maximum': 1000000, 'description': 'Fixed concurrent DB connection budget; Aurora Serverless defaults from maximum configured ACU.', 'exclusiveMinimum': 0}, 'loadBalancerScheme': {'enum': ['internal', 'internet-facing'], 'type': 'string', 'description': 'Owned typical ALB scheme: internal.'}, 'scaleInCpuThreshold': {'type': 'number', 'maximum': 100, 'minimum': 0, 'description': "Compute or Kubernetes only â\x80\x94 overrides the simulation-wide CPU HPA scale-in target for THIS resource's scale decisions only; every other resource keeps the simulation-wide default."}, 'scaleOutCpuThreshold': {'type': 'number', 'maximum': 100, 'minimum': 0, 'description': "Compute or Kubernetes only â\x80\x94 overrides the simulation-wide CPU HPA scale-out target (see the top-level autoscalingTargetCpu) for THIS resource's scale decisions only; every other resource keeps the simulation-wide default."}, 'workloadDatabaseEngine': {'enum': ['mysql'], 'type': 'string', 'description': 'MySQL workload identity on the db.r5.large database.'}, 'workloadDatabaseVersion': {'type': 'string', 'description': 'Owned typical MySQL workload version: 8.0.'}}, 'description': 'Capacity and sizing characteristics for the resource (extra keys such as costMultiplier pass through unchanged)', 'additionalProperties': True}}, 'additionalProperties': False}, 'description': 'List of cloud resources composing this simulation. For a built-in scenario, pass the hydrated resources from scenario.get; scenario.list cards are graph-free. (max 10 in demo mode; mutually exclusive with scenarioId)'}, 'scenarioId': {'type': 'string', 'minLength': 1, 'description': 'Live scenario identifier from scenario.list; mutually exclusive with resources and connections'}, 'connections': {'type': 'array', 'items': {'type': 'object', 'required': ['sourceId', 'targetId'], 'properties': {'label': {'type': 'string', 'description': 'Optional label describing the connection type'}, 'sourceId': {'type': 'string', 'description': 'ID of the source (upstream) resource'}, 'targetId': {'type': 'string', 'description': 'ID of the target (downstream) resource'}}, 'additionalProperties': False}, 'description': 'Directed connections between resources. For a built-in scenario, pass the hydrated connections from scenario.get; scenario.list cards are graph-free. Directed edges should describe traffic flow between resources; omit when using scenarioId.'}, 'description': {'type': 'string', 'maxLength': 500, 'description': "Optional description of the simulation's purpose"}, 'maxInstances': {'type': 'integer', 'minimum': 1, 'description': 'Hard ceiling on the autoscaled compute fleet size, stored as autoscalingConfig.maxInstances. If omitted, the provider default applies (AWS 50, GCP 15, Azure/OCI/DigitalOcean 10) â\x80\x94 which may be much larger than your intended fleet size.'}, 'minInstances': {'type': 'integer', 'minimum': 1, 'description': 'Floor on the autoscaled compute fleet size, stored as autoscalingConfig.minInstances.'}, 'responseMode': {'enum': ['compact', 'full'], 'type': 'string', 'default': 'compact', 'description': "Response detail level. 'compact' (default) returns id, name, status, traffic, and a per-resource summary (id, name, status, cpuPercent) â\x80\x94 keeps the response small for agent loops. 'full' returns the complete simulation object including all resource characteristics and connections."}, 'resilienceConfig': {'type': 'object', 'properties': {'enabled': {'type': 'boolean', 'description': 'Master switch for retry/cascade modeling'}, 'version': {'type': 'number', 'const': 1, 'description': 'Resilience model version'}, 'maxStepWork': {'type': 'integer', 'maximum': 2048, 'minimum': 1}, 'dependencies': {'type': 'array', 'items': {'type': 'object', 'required': ['id', 'sourceId', 'targetId'], 'properties': {'id': {'type': 'string', 'maxLength': 128, 'minLength': 1, 'description': 'Unique dependency edge ID'}, 'capacity': {'type': 'object', 'properties': {'maxRps': {'type': 'number', 'maximum': 500000, 'exclusiveMinimum': 0}, 'maxConcurrent': {'type': 'integer', 'maximum': 1000000, 'exclusiveMinimum': 0}, 'meanServiceTimeMs': {'type': 'number', 'maximum': 120000, 'exclusiveMinimum': 0}}, 'additionalProperties': False}, 'sourceId': {'type': 'string', 'minLength': 1, 'description': 'Upstream resource ID'}, 'targetId': {'type': 'string', 'minLength': 1, 'description': 'Downstream resource ID'}, 'protection': {'type': 'object', 'properties': {'loadShedding': {'type': 'boolean'}, 'rateLimitRps': {'type': 'number', 'maximum': 500000, 'exclusiveMinimum': 0}, 'circuitBreaker': {'type': 'object', 'properties': {'enabled': {'type': 'boolean'}, 'openSteps': {'type': 'integer', 'maximum': 120, 'minimum': 1}, 'minimumRequests': {'type': 'integer', 'maximum': 100000, 'minimum': 1}, 'halfOpenMaxRequests': {'type': 'integer', 'maximum': 10000, 'minimum': 1}, 'failureRateThreshold': {'type': 'number', 'maximum': 1, 'minimum': 0}}, 'additionalProperties': False}}, 'additionalProperties': False}, 'retryPolicy': {'type': 'object', 'properties': {'backoffMs': {'type': 'integer', 'maximum': 60000, 'minimum': 0, 'description': 'Initial retry backoff in milliseconds (default 100).'}, 'timeoutMs': {'type': 'integer', 'maximum': 120000, 'minimum': 1, 'description': 'Per-attempt client deadline in milliseconds (default 2000); includes queue wait, service, and latency.'}, 'maxRetries': {'type': 'integer', 'maximum': 8, 'minimum': 0, 'description': 'Maximum retries for this dependency edge (0 disables retries; default 2).'}, 'jitterRatio': {'type': 'number', 'maximum': 1, 'minimum': 0, 'description': 'Fractional backoff jitter from 0 to 1 (default 0.1).'}, 'retryActorId': {'type': 'string', 'maxLength': 128, 'minLength': 1, 'description': 'Actor producing retries, such as a gateway or client'}, 'retryBudgetRps': {'type': 'number', 'maximum': 500000, 'minimum': 0, 'description': 'Optional absolute retry RPS cap, also bounded by retryBudgetRatio.'}, 'retryBudgetRatio': {'type': 'number', 'maximum': 10, 'minimum': 0, 'description': 'Retry RPS budget as a multiple of original edge RPS (default 2).'}, 'backoffMultiplier': {'type': 'number', 'maximum': 10, 'minimum': 1, 'description': 'Exponential backoff multiplier (default 2).'}}, 'additionalProperties': False}, 'requestRatio': {'type': 'number', 'maximum': 20, 'minimum': 0, 'description': 'Requests to target per original request'}, 'authDependencyId': {'type': 'string', 'maxLength': 128, 'minLength': 1, 'description': 'Dependency receiving generated auth/token traffic'}, 'authRequestsPerAttempt': {'type': 'number', 'maximum': 10, 'minimum': 0, 'description': 'Auth/token requests generated per dependency attempt'}}, 'additionalProperties': False}, 'maxItems': 64, 'description': 'Dependency edges with retry and protection policies'}, 'maxCascadeDepth': {'type': 'integer', 'maximum': 8, 'minimum': 1}, 'maxGeneratedRps': {'type': 'number', 'maximum': 500000, 'exclusiveMinimum': 0}, 'scalingPolicies': {'type': 'array', 'items': {'type': 'object', 'required': ['id', 'dependencyId', 'constrainedMetric', 'observationMetric', 'scaleOutThresholdPercent'], 'properties': {'id': {'type': 'string', 'maxLength': 128, 'minLength': 1}, 'dependencyId': {'type': 'string', 'maxLength': 128, 'minLength': 1}, 'constrainedMetric': {'enum': ['rps', 'concurrency'], 'type': 'string'}, 'observationMetric': {'enum': ['source_cpu', 'capacity_utilization'], 'type': 'string'}, 'observedResourceId': {'type': 'string', 'minLength': 1}, 'scaleOutThresholdPercent': {'type': 'number', 'maximum': 100, 'minimum': 1}, 'scaleOutCapacityMultiplier': {'type': 'number', 'maximum': 20, 'minimum': 1}}, 'additionalProperties': False}, 'maxItems': 32, 'description': 'Capacity-observation policies, including intentional autoscaling blind spots'}, 'scheduledFaults': {'type': 'array', 'items': {'type': 'object', 'required': ['id', 'type', 'startStep'], 'properties': {'id': {'type': 'string', 'maxLength': 128, 'minLength': 1}, 'type': {'enum': ['capacity_limit', 'concurrency_limit', 'latency', 'error_rate', 'traffic_surge'], 'type': 'string', 'description': 'Fault type; traffic_surge increases root client demand'}, 'endStep': {'type': 'integer', 'minimum': 1}, 'errorRate': {'type': 'number', 'maximum': 1, 'minimum': 0}, 'startStep': {'type': 'integer', 'minimum': 0}, 'dependencyId': {'type': 'string'}, 'maxConcurrent': {'type': 'integer', 'maximum': 1000000, 'exclusiveMinimum': 0}, 'addedLatencyMs': {'type': 'number', 'maximum': 120000, 'minimum': 0}, 'capacityPercent': {'type': 'number', 'maximum': 100, 'minimum': 0}, 'targetResourceId': {'type': 'string'}, 'trafficMultiplier': {'type': 'number', 'maximum': 20, 'description': 'traffic_surge demand multiplier', 'exclusiveMinimum': 1}}, 'additionalProperties': False}, 'maxItems': 32, 'description': 'Scheduled capacity, error, latency, concurrency, or traffic-surge faults'}, 'retryGeneratedTrafficAffectsCost': {'type': 'boolean'}}, 'description': 'Optional retry/cascade resilience model returned by scenario.get', 'additionalProperties': False}, 'scaleOutCpuPercent': {'type': 'number', 'maximum': 100, 'minimum': 0, 'description': 'Grok-compatible alias for the CPU HPA scale-out target; if multiple target names are sent they must match.'}, 'autoscalingTargetCpu': {'type': 'number', 'maximum': 100, 'minimum': 0, 'description': 'Canonical CPU HPA scale-out target percent. CWM synthesizes unrelated autoscaling defaults.'}, 'ecsCpuTargetTracking': {'type': 'boolean', 'description': 'Opt into CPU-only ECS Fargate target tracking.'}, 'scaleOutCpuThreshold': {'type': 'number', 'maximum': 100, 'minimum': 0, 'description': 'Equivalent alias for autoscalingTargetCpu; if both are sent they must match.'}, 'scaleInCooldownSeconds': {'type': 'integer', 'maximum': 86400, 'minimum': 0, 'description': 'ECS CPU target-tracking scale-in cooldown in simulated seconds.'}, 'scaleOutCooldownSeconds': {'type': 'integer', 'maximum': 86400, 'minimum': 0, 'description': 'ECS CPU target-tracking scale-out cooldown in simulated seconds.'}, 'simulationSecondsPerStep': {'type': 'number', 'maximum': 60, 'description': 'Simulated seconds per step for ECS cooldowns (default 1).', 'exclusiveMinimum': 0}, 'autoscaleTargetCpuPercent': {'type': 'number', 'maximum': 100, 'minimum': 0, 'description': 'Grok-compatible alias for the CPU HPA scale-out target; if multiple target names are sent they must match.'}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'id': {'type': 'string', 'description': 'Unique simulation ID â\x80\x94 use with simulation.step, simulation.metrics, etc.'}, 'name': {'type': 'string', 'description': 'Simulation name'}, 'status': {'type': 'string', 'description': 'Current simulation status'}, 'traffic': {'type': 'number', 'description': 'Current traffic in RPS'}, 'hpaAudit': {'type': 'object', 'required': ['targetSupplied', 'suppliedField', 'requestedTargetCpu', 'effectiveScaleOutCpuPercent', 'effectiveScaleInCpuPercent', 'defaulted', 'provider', 'resourceCategory'], 'properties': {'provider': {'type': 'string'}, 'defaulted': {'type': 'boolean'}, 'suppliedField': {'type': ['string', 'null']}, 'targetSupplied': {'type': 'boolean'}, 'resourceCategory': {'type': 'string'}, 'defaultExplanation': {'type': 'string'}, 'requestedTargetCpu': {'type': ['number', 'null']}, 'effectiveScaleInCpuPercent': {'type': 'number'}, 'effectiveScaleOutCpuPercent': {'type': 'number'}}, 'description': 'CPU HPA create-time audit: whether a target arrived, its accepted field, the persisted thresholds, and the provider-default explanation when omitted.', 'additionalProperties': False}, 'appWeight': {'enum': ['lean', 'typical', 'heavy'], 'type': 'string'}, 'resources': {'type': 'array', 'items': {'type': 'object', 'properties': {'id': {'type': 'string', 'description': 'Resource ID'}, 'name': {'type': 'string', 'description': 'Resource display name'}, 'status': {'type': 'string', 'description': 'Health status (healthy/warning/critical/failed)'}, 'routedRps': {'type': 'number', 'description': 'Requests per second routed to this resource (compute/kubernetes only)'}, 'cpuPercent': {'type': 'number', 'description': 'CPU utilization (%)'}, 'isRoutable': {'type': 'boolean', 'description': 'Whether this compute/Kubernetes resource can receive traffic'}, 'availabilityState': {'enum': ['available', 'degraded', 'unavailable'], 'type': 'string', 'description': 'Availability derived from routed traffic; degraded can still serve, unavailable cannot'}, 'recoveryBlockedReason': {'type': 'string', 'description': 'Engine recovery guard currently blocking cooldown progress, when present'}}, 'additionalProperties': True}, 'description': 'Per-resource summary (compact mode) or full resource states (full mode)'}, 'scenarioHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Canonical SHA-256 of the persisted scenario graph and attached traffic-pattern order'}, 'engineVersion': {'type': 'string', 'description': 'Simulation engine version used for this prediction.'}, 'replayIdentity': {'type': 'object', 'required': ['scenarioHash', 'effectiveConfigHash'], 'properties': {'scenarioHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Canonical SHA-256 of the persisted scenario graph and attached traffic-pattern order'}, 'effectiveConfigHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Original replay-only SHA-256 of the effective six-control startup configuration; top-level effectiveConfigHash is the versioned prediction identity.'}}, 'additionalProperties': False}, 'normalizedConfig': {'type': 'object', 'properties': {'resources': {'type': 'array', 'items': {'type': 'object', 'properties': {'id': {'type': 'string', 'description': 'Resource ID'}, 'name': {'type': 'string', 'description': 'Resource display name'}, 'type': {'type': 'string', 'description': 'Resource type (compute, database, kubernetes, â\x80¦)'}, 'maxNodes': {'type': 'number', 'description': 'Autoscale ceiling (scale-out stops here)'}, 'minNodes': {'type': 'number', 'description': 'Autoscale floor (scale-in stops here)'}, 'provider': {'type': 'string', 'description': 'Cloud provider'}, 'nodeCount': {'type': 'number', 'description': 'Current node count'}, 'nodePools': {'type': 'array', 'items': {'type': 'object', 'required': ['nodeCount', 'minNodes', 'maxNodes', 'perNodeRatePerHour'], 'properties': {'name': {'type': 'string'}, 'maxNodes': {'type': 'number'}, 'minNodes': {'type': 'number'}, 'nodeCount': {'type': 'number'}, 'perNodeRatePerHour': {'type': 'number'}, 'perNodeMaxThroughputRps': {'type': 'number'}}, 'additionalProperties': True}, 'description': 'Per-pool billing details for multi-pool clusters (absent on single-pool clusters)'}, 'spotFactor': {'type': 'number', 'description': 'Spot-instance discount factor applied to the node cost (absent = 1, i.e. no discount)'}, 'behaviorModel': {'type': 'object', 'properties': {'description': {'type': 'string', 'description': 'Human-readable description of the latency model applied to this resource'}, 'latencyPath': {'type': 'string', 'description': "Canonical latency simulation path identifier, e.g. 'gpu-inference/ttft-decode' or 'generic-kubernetes/load-saturation'. Machine-match against LATENCY_PATHS constants â\x80\x94 do not string-parse."}, 'modelingNote': {'type': 'string', 'description': 'Provenance disclaimer for latency figures. On inferenceMode clusters states that TTFT/P95/P99 are CWM simulation-model estimates from accelerator catalog parameters, not externally measured benchmarks. When topology is omitted also notes the legacy-baseline assumption (topologyThroughputFactor=1.0).'}, 'behaviourFidelity': {'enum': ['modeled', 'estimated'], 'type': 'string', 'description': "'modeled' = parameters from the performance catalog; 'estimated' = defaults substituted"}, 'topologyInterNode': {'type': ['string', 'null'], 'description': 'Raw characteristics.topology.interNode value or null when absent'}, 'topologyIntraNode': {'type': ['string', 'null'], 'description': "Raw characteristics.topology.intraNode value ('pcie', 'nvlink', 'nvlink-nvswitch', 'infiniband') or null when absent"}, 'topologyNodeCount': {'type': 'number', 'description': 'Node count used when resolving topology latency factors (â\x89¥ 1)'}, 'topologyTtftFactor': {'type': 'number', 'description': "TTFT latency scaling factor applied by the engine's GPU inference model (resolveTopologyLatencyFactors). Describes the TTFT+decode latency dimension â\x80\x94 distinct from the throughput scaling factor."}, 'topologyDecodeFactor': {'type': 'number', 'description': "Decode-rate latency scaling factor applied by the engine's GPU inference model (resolveTopologyLatencyFactors). Describes the TTFT+decode latency dimension â\x80\x94 distinct from the throughput scaling factor."}, 'acceleratorResolution': {'enum': ['catalog', 'fallback'], 'type': 'string', 'description': "'catalog' = accelerator matched ACCELERATOR_TOKENS_PER_SEC; 'fallback' = unrecognised, defaults substituted (TTFT 600 ms, decode 700 tok/s, perNodeTokens 500)"}, 'topologyIsLegacyBaseline': {'type': 'boolean', 'description': 'true when topology.intraNode was absent or unrecognised, meaning topologyThroughputFactor=1.0 reflects the pre-topology legacy baseline (not a fabric-specific measurement). false when an explicit, recognised intraNode value was supplied and the throughput factor is topology-modelled.'}, 'topologyThroughputFactor': {'type': 'number', 'description': "Throughput scaling coefficient applied by resolveTopologyScalingFactor to this cluster's token capacity (nodes Ã\x97 perNodeTokensPerSec Ã\x97 factor Ã\x97 gpuUtil/100). 1.0 when topology.intraNode is absent or unrecognised (legacy baseline, topologyIsLegacyBaseline=true). Known fabric penalties: pcieâ\x89\x880.70, nvlinkâ\x89\x880.85, nvlink-nvswitchâ\x89\x880.92, infinibandâ\x89\x881.0. Entirely separate from topologyTtftFactor/topologyDecodeFactor, which describe the TTFT+decode latency model."}, 'topologyCalibrationStatus': {'enum': ['calibrated', 'estimated', 'not_applicable'], 'type': 'string', 'description': 'Calibration status of the TTFT+decode latency factors (resolveTopologyLatencyFactors). Entirely separate from topologyThroughputCalibrationStatus, which covers throughput scaling.'}, 'topologyThroughputCalibrationStatus': {'enum': ['calibrated', 'estimated', 'not_applicable'], 'type': 'string', 'description': "Calibration status of the throughput scaling factor (resolveTopologyScalingFactor). Distinct from topologyCalibrationStatus, which covers TTFT+decode latency factors. 'estimated' for all current intraNode coefficients (engineering assumptions from bandwidth arithmetic). 'not_applicable' on non-inferenceMode resources. Reserve 'calibrated' for when a directly measured per-request throughput ratio is added."}}, 'description': 'Latency simulation model and GPU interconnect topology parameters for this resource. latencyPath identifies which engine path runs at step time. Two distinct model components: (1) throughput scaling â\x80\x94 topologyThroughputFactor (resolveTopologyScalingFactor) bounds effective token capacity; topologyThroughputCalibrationStatus and topologyIsLegacyBaseline qualify that factor. (2) latency shape â\x80\x94 topologyTtftFactor + topologyDecodeFactor (resolveTopologyLatencyFactors) scale TTFT and decode curves; topologyCalibrationStatus qualifies those factors. Both components are present on inferenceMode Kubernetes clusters; latencyPath is present on all resource types.', 'additionalProperties': True}, 'rateProvenance': {'type': 'object', 'required': ['basis', 'resolution', 'lastVerified', 'region'], 'properties': {'sku': {'type': 'string', 'description': 'Exact catalog SKU used for this rate, when available'}, 'basis': {'enum': ['on-demand', 'spot', 'estimated'], 'type': 'string', 'description': 'Pricing basis: published on-demand list price, spot, or a flat estimate'}, 'region': {'type': ['string', 'null'], 'description': 'Region the verified rate applies to; null for region-uniform pricing'}, 'source': {'type': 'string', 'description': 'Official pricing source URL, when verified'}, 'resolution': {'enum': ['large-sku', 'entry-level-fallback', 'base-rate-multiplier', 'fargate-task-vcpu-memory', 'aurora-acu-hour', 'cloud-run-request-based', 'request-serving-usage-based'], 'type': 'string', 'description': 'Which lookup path resolved the rate'}, 'architecture': {'type': 'string', 'description': 'Fargate task architecture when architecture-specific billing applies'}, 'lastVerified': {'type': ['string', 'null'], 'description': "ISO-8601 date the constant was last cross-checked against the provider's pricing page; null when unknown"}}, 'description': "Trust metadata for the resolved rate: pricing basis (on-demand/spot/estimated), lookup path, and the date the constant was last verified against the provider's public pricing page. Present on compute, database, and GPU inference resources; absent on resource types where no rate is resolved.", 'additionalProperties': False}, 'requestServing': {'type': 'object', 'properties': {'maxTaskCount': {'type': 'integer', 'exclusiveMinimum': 0}, 'minTaskCount': {'type': 'integer', 'minimum': 0}, 'serviceFamily': {'type': 'string'}, 'desiredTaskCount': {'type': 'integer', 'minimum': 0}, 'perTaskCapacityRps': {'type': 'number', 'exclusiveMinimum': 0}, 'aggregateCapacityRps': {'type': 'number', 'minimum': 0}, 'maxAggregateCapacityRps': {'type': 'number', 'minimum': 0}, 'perTaskCapacityEvidence': {'type': 'object', 'required': ['basis', 'origin'], 'properties': {'basis': {'enum': ['assumption', 'documented', 'measured', 'catalog'], 'type': 'string'}, 'origin': {'enum': ['caller-provided', 'fargate-size-heuristic', 'policy-default', 'provider-catalog'], 'type': 'string'}, 'source': {'type': 'string'}}, 'additionalProperties': False}}, 'description': "Resolved request-serving settings. For Fargate, aggregateCapacityRps is desiredTaskCount Ã\x97 this service's own rate, maxAggregateCapacityRps uses maxTaskCount, and perTaskCapacityEvidence distinguishes assumptions from caller-asserted documentation/measurement or provider catalog data.", 'additionalProperties': True}, 'resolvedSkuLabel': {'type': 'string', 'description': "The GPU SKU that was matched (e.g. 'a100-80gb') or a fallback label naming the provider default â\x80\x94 present only on inference-mode kubernetes resources"}, 'billingFloorNodes': {'type': 'number', 'description': 'Minimum node count the engine will ever bill â\x80\x94 scale-in cannot go below this'}, 'autoscaleThreshold': {'type': 'object', 'required': ['scaleOutCpuPercent', 'scaleInCpuPercent'], 'properties': {'scaleInCpuPercent': {'type': 'number', 'description': 'CPU % below which a scale-in event fires'}, 'scaleOutCpuPercent': {'type': 'number', 'description': 'CPU % that triggers a scale-out event'}}, 'description': 'Effective autoscale thresholds (provider profile merged with any explicit config)', 'additionalProperties': False}, 'capacityProvenance': {'enum': ['absent', 'capacityRps', 'explicit_maxThroughput', 'explicit_requestsPerSecond', 'provider_catalog', 'generic_fallback'], 'type': 'string', 'description': 'Source of the effective compute capacity: explicit caller field, provider catalog, or generic fallback'}, 'perNodeRatePerHour': {'type': 'number', 'description': 'Per-node billing rate in USD/hr'}, 'resolvedHourlyRate': {'type': 'number', 'description': 'Resolved hourly billing rate in USD/hr (base rate Ã\x97 multiplier)'}, 'perNodeTokensPerSec': {'type': 'number', 'description': 'Modelled per-node token throughput at full utilisation (tokens/sec) â\x80\x94 present only on inference-mode kubernetes resources'}, 'resolvedStorageTier': {'type': 'string', 'description': 'Storage tier / size string used to resolve pricing (database resources)'}, 'controlPlaneFeePerHour': {'type': 'number', 'description': 'Cost-only cluster management fee in USD/hr; it does not represent control-plane CPU, throttling, or recovery telemetry'}, 'resolvedCostMultiplier': {'type': 'number', 'description': 'Effective cost multiplier the engine applies to the base provider rate for this resource'}, 'resolvedGpuRatePerNode': {'type': 'number', 'description': 'Resolved GPU node billing rate in USD/hr â\x80\x94 present only on inference-mode kubernetes resources'}, 'billingFloorCostPerHour': {'type': 'number', 'description': 'Minimum cluster cost in USD/hr (cost at billingFloorNodes)'}, 'currentNodesCostPerHour': {'type': 'number', 'description': 'Total cluster cost at the current node count (control plane + nodes Ã\x97 rate) Ã\x97 spotFactor in USD/hr'}, 'effectiveCpuCapacityRps': {'type': 'number', 'description': 'CPU-curve denominator after applying the provider threshold to explicit capacityRps'}, 'resolvedConnectionLimit': {'type': 'number', 'description': 'Effective max-connection limit the engine uses for connection-pressure modelling (database resources)'}, 'resolvedMaxThroughputRps': {'type': 'number', 'description': 'Per-node RPS ceiling the engine uses for load and autoscale calculations'}}, 'additionalProperties': True}, 'description': 'Per-resource billing parameters resolved by the engine at create time'}}, 'description': 'Engine-resolved billing parameters for every resource: cost multipliers, hourly rates, autoscale thresholds (scaleOut/scaleIn CPU %), GPU SKU, per-node token throughput, billing floor, connection limits. Use this immediately after create to verify the simulation was set up as intended â\x80\x94 e.g. confirm which GPU SKU was resolved, the effective billing floor, or the autoscale CPU threshold that will drive scale-out.', 'additionalProperties': True}, 'resilienceConfig': {'type': 'object', 'properties': {'enabled': {'type': 'boolean', 'description': 'Master switch for retry/cascade modeling'}, 'version': {'type': 'number', 'const': 1, 'description': 'Resilience model version'}, 'maxStepWork': {'type': 'integer', 'maximum': 2048, 'minimum': 1}, 'dependencies': {'type': 'array', 'items': {'type': 'object', 'required': ['id', 'sourceId', 'targetId'], 'properties': {'id': {'type': 'string', 'maxLength': 128, 'minLength': 1, 'description': 'Unique dependency edge ID'}, 'capacity': {'type': 'object', 'properties': {'maxRps': {'type': 'number', 'maximum': 500000, 'exclusiveMinimum': 0}, 'maxConcurrent': {'type': 'integer', 'maximum': 1000000, 'exclusiveMinimum': 0}, 'meanServiceTimeMs': {'type': 'number', 'maximum': 120000, 'exclusiveMinimum': 0}}, 'additionalProperties': False}, 'sourceId': {'type': 'string', 'minLength': 1, 'description': 'Upstream resource ID'}, 'targetId': {'type': 'string', 'minLength': 1, 'description': 'Downstream resource ID'}, 'protection': {'type': 'object', 'properties': {'loadShedding': {'type': 'boolean'}, 'rateLimitRps': {'type': 'number', 'maximum': 500000, 'exclusiveMinimum': 0}, 'circuitBreaker': {'type': 'object', 'properties': {'enabled': {'type': 'boolean'}, 'openSteps': {'type': 'integer', 'maximum': 120, 'minimum': 1}, 'minimumRequests': {'type': 'integer', 'maximum': 100000, 'minimum': 1}, 'halfOpenMaxRequests': {'type': 'integer', 'maximum': 10000, 'minimum': 1}, 'failureRateThreshold': {'type': 'number', 'maximum': 1, 'minimum': 0}}, 'additionalProperties': False}}, 'additionalProperties': False}, 'retryPolicy': {'type': 'object', 'properties': {'backoffMs': {'type': 'integer', 'maximum': 60000, 'minimum': 0, 'description': 'Initial retry backoff in milliseconds (default 100).'}, 'timeoutMs': {'type': 'integer', 'maximum': 120000, 'minimum': 1, 'description': 'Per-attempt client deadline in milliseconds (default 2000); includes queue wait, service, and latency.'}, 'maxRetries': {'type': 'integer', 'maximum': 8, 'minimum': 0, 'description': 'Maximum retries for this dependency edge (0 disables retries; default 2).'}, 'jitterRatio': {'type': 'number', 'maximum': 1, 'minimum': 0, 'description': 'Fractional backoff jitter from 0 to 1 (default 0.1).'}, 'retryActorId': {'type': 'string', 'maxLength': 128, 'minLength': 1, 'description': 'Actor producing retries, such as a gateway or client'}, 'retryBudgetRps': {'type': 'number', 'maximum': 500000, 'minimum': 0, 'description': 'Optional absolute retry RPS cap, also bounded by retryBudgetRatio.'}, 'retryBudgetRatio': {'type': 'number', 'maximum': 10, 'minimum': 0, 'description': 'Retry RPS budget as a multiple of original edge RPS (default 2).'}, 'backoffMultiplier': {'type': 'number', 'maximum': 10, 'minimum': 1, 'description': 'Exponential backoff multiplier (default 2).'}}, 'additionalProperties': False}, 'requestRatio': {'type': 'number', 'maximum': 20, 'minimum': 0, 'description': 'Requests to target per original request'}, 'authDependencyId': {'type': 'string', 'maxLength': 128, 'minLength': 1, 'description': 'Dependency receiving generated auth/token traffic'}, 'authRequestsPerAttempt': {'type': 'number', 'maximum': 10, 'minimum': 0, 'description': 'Auth/token requests generated per dependency attempt'}}, 'additionalProperties': False}, 'maxItems': 64, 'description': 'Dependency edges with retry and protection policies'}, 'maxCascadeDepth': {'type': 'integer', 'maximum': 8, 'minimum': 1}, 'maxGeneratedRps': {'type': 'number', 'maximum': 500000, 'exclusiveMinimum': 0}, 'scalingPolicies': {'type': 'array', 'items': {'type': 'object', 'required': ['id', 'dependencyId', 'constrainedMetric', 'observationMetric', 'scaleOutThresholdPercent'], 'properties': {'id': {'type': 'string', 'maxLength': 128, 'minLength': 1}, 'dependencyId': {'type': 'string', 'maxLength': 128, 'minLength': 1}, 'constrainedMetric': {'enum': ['rps', 'concurrency'], 'type': 'string'}, 'observationMetric': {'enum': ['source_cpu', 'capacity_utilization'], 'type': 'string'}, 'observedResourceId': {'type': 'string', 'minLength': 1}, 'scaleOutThresholdPercent': {'type': 'number', 'maximum': 100, 'minimum': 1}, 'scaleOutCapacityMultiplier': {'type': 'number', 'maximum': 20, 'minimum': 1}}, 'additionalProperties': False}, 'maxItems': 32, 'description': 'Capacity-observation policies, including intentional autoscaling blind spots'}, 'scheduledFaults': {'type': 'array', 'items': {'type': 'object', 'required': ['id', 'type', 'startStep'], 'properties': {'id': {'type': 'string', 'maxLength': 128, 'minLength': 1}, 'type': {'enum': ['capacity_limit', 'concurrency_limit', 'latency', 'error_rate', 'traffic_surge'], 'type': 'string', 'description': 'Fault type; traffic_surge increases root client demand'}, 'endStep': {'type': 'integer', 'minimum': 1}, 'errorRate': {'type': 'number', 'maximum': 1, 'minimum': 0}, 'startStep': {'type': 'integer', 'minimum': 0}, 'dependencyId': {'type': 'string'}, 'maxConcurrent': {'type': 'integer', 'maximum': 1000000, 'exclusiveMinimum': 0}, 'addedLatencyMs': {'type': 'number', 'maximum': 120000, 'minimum': 0}, 'capacityPercent': {'type': 'number', 'maximum': 100, 'minimum': 0}, 'targetResourceId': {'type': 'string'}, 'trafficMultiplier': {'type': 'number', 'maximum': 20, 'description': 'traffic_surge demand multiplier', 'exclusiveMinimum': 1}}, 'additionalProperties': False}, 'maxItems': 32, 'description': 'Scheduled capacity, error, latency, concurrency, or traffic-surge faults'}, 'retryGeneratedTrafficAffectsCost': {'type': 'boolean'}}, 'description': 'Effective resilience model returned after create; per-edge retryPolicy values include defaults for omitted fields.', 'additionalProperties': False}, 'autoscalingConfig': {'type': 'object', 'properties': {}, 'description': "Effective simulation scaling config; full response also contains the ECS resource's linked CPU-only target-tracking policy.", 'additionalProperties': True}, 'appWeightDefaulted': {'type': 'boolean'}, 'predictionEvidence': {'anyOf': [{'type': 'object', 'required': ['version', 'evidenceLevel', 'legacyGeneric', 'appWeight', 'appWeightDefaulted', 'loadScope', 'sourceIds', 'formulaIds', 'appCpu', 'appCpuByResource', 'latencyP50', 'latencyP95', 'note'], 'properties': {'note': {'type': 'string'}, 'appCpu': {'anyOf': [{'type': 'object', 'required': ['low', 'central', 'high'], 'properties': {'low': {'type': 'number', 'minimum': 0}, 'high': {'type': 'number', 'minimum': 0}, 'central': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}, {'type': 'null'}]}, 'version': {'type': 'number', 'const': 1}, 'appWeight': {'enum': ['lean', 'typical', 'heavy'], 'type': 'string'}, 'loadScope': {'enum': ['within measured load', 'beyond measured load', 'below measured load', 'not calibrated'], 'type': 'string'}, 'sourceIds': {'type': 'array', 'items': {'type': 'string'}}, 'formulaIds': {'type': 'array', 'items': {'type': 'string'}}, 'latencyP50': {'anyOf': [{'type': 'object', 'required': ['low', 'central', 'high', 'evidenceLevel'], 'properties': {'low': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/low'}, 'high': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/high'}, 'central': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/central'}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}}, 'additionalProperties': False}, {'type': 'null'}]}, 'latencyP95': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'latencyP99': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'evidenceLevel': {'enum': ['measured', 'scaled from measured', 'reference estimate'], 'type': 'string'}, 'legacyGeneric': {'type': 'boolean'}, 'appCpuByResource': {'type': 'array', 'items': {'type': 'object', 'required': ['low', 'central', 'high', 'evidenceLevel', 'resourceId', 'legacyGeneric'], 'properties': {'low': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/low'}, 'high': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/high'}, 'central': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/central'}, 'resourceId': {'type': 'string'}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}, 'legacyGeneric': {'type': 'boolean'}}, 'additionalProperties': False}}, 'appWeightDefaulted': {'type': 'boolean'}, 'latencyAvailability': {'type': 'object', 'required': ['status', 'reason', 'resourceIds'], 'properties': {'reason': {'type': 'string', 'minLength': 1}, 'status': {'type': 'string', 'const': 'unavailable'}, 'resourceIds': {'type': 'array', 'items': {'type': 'string', 'minLength': 1}}}, 'additionalProperties': False}, 'requestServingCapacity': {'type': 'array', 'items': {'type': 'object', 'required': ['resourceId', 'resourceName', 'taskCount', 'perTaskCapacityRps', 'aggregateCapacityRps', 'maxTaskCount', 'maxAggregateCapacityRps', 'capacityEvidence'], 'properties': {'taskCount': {'type': 'integer', 'minimum': 0}, 'resourceId': {'type': 'string'}, 'maxTaskCount': {'type': 'integer', 'exclusiveMinimum': 0}, 'resourceName': {'type': 'string'}, 'capacityEvidence': {'type': 'object', 'required': ['basis', 'origin'], 'properties': {'basis': {'enum': ['assumption', 'documented', 'measured', 'catalog'], 'type': 'string'}, 'origin': {'enum': ['caller-provided', 'fargate-size-heuristic', 'policy-default', 'provider-catalog'], 'type': 'string'}, 'source': {'type': 'string'}}, 'additionalProperties': False}, 'perTaskCapacityRps': {'type': 'number', 'exclusiveMinimum': 0}, 'aggregateCapacityRps': {'type': 'number', 'minimum': 0}, 'maxAggregateCapacityRps': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}}, 'latencyPercentileStatus': {'enum': ['modeled', 'unavailable'], 'type': 'string'}}, 'additionalProperties': False}, {'type': 'object', 'required': ['version', 'evidenceLevel', 'appWeight', 'appWeightDefaulted', 'loadScope', 'sourceIds', 'formulaIds', 'appCpu', 'appCpuByResource', 'latencyP50', 'latencyP95', 'note'], 'properties': {'note': {'type': 'string'}, 'appCpu': {'type': 'null'}, 'version': {'type': 'number', 'const': 1}, 'appWeight': {'type': 'null'}, 'loadScope': {'type': 'null'}, 'sourceIds': {'type': 'array', 'items': {'type': 'string'}}, 'formulaIds': {'type': 'array', 'items': {'type': 'string'}}, 'latencyP50': {'type': 'null'}, 'latencyP95': {'type': 'null'}, 'latencyP99': {'type': 'null'}, 'evidenceLevel': {'type': 'string', 'const': 'unavailable/legacy'}, 'legacyGeneric': {'type': 'null'}, 'appCpuByResource': {'type': 'array', 'items': {'type': 'object', 'required': ['low', 'central', 'high', 'evidenceLevel', 'resourceId'], 'properties': {'low': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/low'}, 'high': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/high'}, 'central': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/central'}, 'resourceId': {'type': 'string'}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}}, 'additionalProperties': False}, 'maxItems': 0}, 'appWeightDefaulted': {'type': 'null'}}, 'additionalProperties': False}, {'type': 'object', 'required': ['version', 'evidenceLevel', 'appWeight', 'appWeightDefaulted', 'loadScope', 'sourceIds', 'formulaIds', 'appCpu', 'appCpuByResource', 'latencyP50', 'latencyP95', 'note'], 'properties': {'note': {'type': 'string'}, 'appCpu': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0'}, {'type': 'null'}]}, 'version': {'type': 'number', 'const': 1}, 'appWeight': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appWeight'}, 'loadScope': {'enum': ['within measured load', 'beyond measured load', 'below measured load', 'not calibrated'], 'type': 'string'}, 'sourceIds': {'type': 'array', 'items': {'type': 'string'}}, 'formulaIds': {'type': 'array', 'items': {'type': 'string'}}, 'latencyP50': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'latencyP95': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'latencyP99': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}, 'legacyGeneric': {'type': 'boolean'}, 'appCpuByResource': {'type': 'array', 'items': {'type': 'object', 'required': ['low', 'central', 'high', 'evidenceLevel', 'resourceId'], 'properties': {'low': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/low'}, 'high': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/high'}, 'central': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/central'}, 'resourceId': {'type': 'string'}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}, 'legacyGeneric': {'type': 'boolean'}}, 'additionalProperties': False}}, 'appWeightDefaulted': {'type': 'boolean'}, 'latencyAvailability': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyAvailability'}, 'latencyPercentileStatus': {'enum': ['modeled', 'unavailable'], 'type': 'string'}}, 'additionalProperties': False}]}, 'calibrationEvidence': {'type': 'object', 'required': ['kind', 'latencyBasis', 'note'], 'properties': {'kind': {'enum': ['owned', 'modeled'], 'type': 'string', 'description': 'owned for the exact AWS CRUD fit; modeled when the gate fails or another generic model applies.'}, 'note': {'type': 'string', 'description': "Names the fit scope and limitations. For canonical workload inference, states it is a modeling assumption. For modeled fallback, names all actual failed checks in stable order with resource IDs and safe expected/actual values; never only a generic 'not applied'."}, 'latencyBasis': {'type': 'string', 'description': 'General modeled latency path, such as in-VPC ALB rather than end-to-end; see latencyP99Basis for percentile-specific P99 provenance.'}, 'calibrationId': {'type': 'string', 'description': 'Versioned owned calibration identifier; present only when that calibration applies.'}, 'latencyP99Basis': {'type': 'string', 'description': 'Percentile-specific P99 basis: owned in-VPC internal-ALB fit, scaled from that fit (not directly measured), or uncalibrated generic model.'}}, 'description': 'Owned-versus-modeled evidence and latency boundary.', 'additionalProperties': False}, 'effectiveConfigHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Versioned prediction hash over replay startup inputs, engine version, and calibration identity; see replayIdentity.effectiveConfigHash for the original replay-only hash.'}, 'scenarioAttribution': {'type': 'object', 'required': ['id', 'version', 'revision', 'source'], 'properties': {'id': {'type': 'string', 'description': 'Live scenario catalog ID that supplied the simulation graph'}, 'source': {'type': 'string', 'const': 'scenario-catalog', 'description': 'Trusted server-side attribution source'}, 'version': {'type': ['string', 'null'], 'description': 'Catalog version, when provided'}, 'revision': {'type': ['string', 'null'], 'description': 'Catalog revision, when provided'}}, 'description': 'Trusted server-side attribution copied from the live scenario catalog; absent for explicit resource-graph creates', 'additionalProperties': False}, 'effectiveMaxInstances': {'type': 'number', 'description': 'The fleet-size ceiling the engine will enforce (autoscalingConfig.maxInstances, or the provider default when unset)'}, 'effectiveMinInstances': {'type': 'number', 'description': 'The fleet-size floor the engine will enforce (autoscalingConfig.minInstances, or the provider default when unset)'}, 'predictionEffectiveConfigHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Versioned prediction hash over replay startup inputs, engine version, and calibration identity.'}}, 'description': "Compact created-simulation summary by default (id, name, status, traffic, per-resource summary), the complete simulation object with responseMode 'full', or a structured limit response (status: limit_reached) when a demo cap is reached", 'additionalProperties': True}
simulation.delete
Delete Simulation
Permanently delete an owned temporary anonymous demo simulation and its metrics, events, failures, and capability. This is the explicit way to free a simulation slot; deletion is irreversible, while the existing demo TTL remains the safety net for abandoned simulations. Prerequisite: a simulationId from simulation.create, or an active simulation in the preserved MCP session. The likely next tool is simulation.create to use the freed slot. A successful response is { deleted: true, id }; failed ownership checks do not delete or revoke anything. Authenticate with an API key to unlock all 62 tools and persistent simulation management.
Destructive Idempotent
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'simulationId': {'type': 'string', 'description': "Simulation ID returned by simulation.create. Preserve Mcp-Session-Id to omit this field and use the session's current simulation; if your connector starts a fresh MCP session for each call (for example Grok Bot or Cursor), pass this explicit ID after every fresh initialization. A fresh session has no current-simulation pointer and returns NO_ACTIVE_SIMULATION when the ID is omitted. Anonymous capabilities are short-lived (30 minutes by default), unguessable, and revoked when the demo expires or is deleted; proxy IP changes do not invalidate them. Do not treat the ID as a durable share link."}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'id': {'type': 'string', 'description': 'ID of the deleted simulation'}, 'deleted': {'type': 'boolean', 'description': 'True when the simulation was deleted'}, 'simulationId': {'type': 'string', 'description': 'ID used for the deletion'}, 'simulationIdSource': {'enum': ['explicit', 'session_default'], 'type': 'string'}}, 'additionalProperties': True}
simulation.inject_failure
Inject Failure
Fail one compute/Kubernetes node or database in a temporary anonymous demo simulation (marked critical, not removed). Targeted database quick failures are bounded and reversible; with no serving database at positive load, simulation.step reports 100% errors, zero goodput and errorBreakdown.dbFailure: 100. simulation.recover_resource can restore the database early. An AWS Aurora writer with characteristics.auroraStandbyResourceId pointing to a healthy related replicaOf database has an opt-in modeled failover: the first step is unavailable, the second shows the standby serving without a residual writer-outage penalty. MultiAz or an unrelated second database alone does not establish a standby. Read metrics.databases[].auroraFailover and the promotion event for the failed and standby IDs and success, plus errorRate and throughput. Each quick step is one modeled simulation second, so second-step promotion is one second after injection. Chaos database_crash samples every 10 seconds and defaults to a 30-second promotion and a 1800-second sole-writer restart after its injection duration; compare matching phases, not equal step indices or wall-clock times. These are deterministic assumptions, not observed AWS behavior. No API key required for this temporary anonymous demo operation. Exact targeting: pass resourceName (human-readable name, e.g. 'app-server-01'; exact match preferred, an unambiguous prefix is accepted) or resourceId to fail a specific resource — including an individual named instance, not only a group. If resourceName matches multiple resources the call fails with a 400 listing every matching candidate by name — retry with one exact name (or its resourceId) from that list. The database must be healthy; an already failed database returns a 400 describing its current status. When neither parameter is supplied, a RANDOM healthy compute/Kubernetes node is selected (not a database) — this path is non-deterministic and NOT suitable for controlled scenarios or replay; always target by name/id when reproducing a precise fault sequence. The response always echoes the applied outcome via resolvedResourceId, resolvedResourceName, and previousHealth (populated from the selected resource on the random path too). Network, storage, cache, queue, and security resource types are not supported by quick injection. For typed database_overload use authenticated failure.create (unavailable anonymously); instance_kill PERMANENTLY removes the instance — failure.delete does not restore it; use instance_down instead for a reversible single-node outage. Returns the updated resource list and the failure event that was logged. The likely next tool is simulation.step to observe how the architecture degrades under failure, then simulation.metrics to review the health impact. Do not use it to advance simulation time — that is simulation.step. Pass simulationId from simulation.create when this call is made from a fresh MCP session; otherwise you may omit it to target the current simulation in the preserved MCP session. Authenticate with an API key to unlock all 62 tools including typed durational failures and chaos engineering.
Destructive
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'resourceId': {'type': 'string', 'description': 'Optional: ID of the resource to fail. Takes precedence over resourceName.'}, 'resourceName': {'type': 'string', 'description': 'Optional: name of the resource to fail (exact match preferred; unambiguous prefix accepted). Ambiguous names return a 400 with a candidate list.'}, 'simulationId': {'type': 'string', 'description': "Simulation ID returned by simulation.create. Preserve Mcp-Session-Id to omit this field and use the session's current simulation; if your connector starts a fresh MCP session for each call (for example Grok Bot or Cursor), pass this explicit ID after every fresh initialization. A fresh session has no current-simulation pointer and returns NO_ACTIVE_SIMULATION when the ID is omitted. Anonymous capabilities are short-lived (30 minutes by default), unguessable, and revoked when the demo expires or is deleted; proxy IP changes do not invalidate them. Do not treat the ID as a durable share link."}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'event': {'type': 'object', 'description': 'Failure event that was logged', 'additionalProperties': {}}, 'resources': {'type': 'array', 'items': {'type': 'object', 'additionalProperties': {}}, 'description': 'Updated resource list after failure injection'}, 'previousHealth': {'type': 'string', 'description': "The resource's health status immediately before the failure was applied"}, 'resolvedResourceId': {'type': 'string', 'description': 'ID of the resource that was failed (targeted or randomly selected)'}, 'resolvedResourceName': {'type': 'string', 'description': 'Name of the resource that was failed'}}, 'additionalProperties': True}
simulation.inject_traffic
Inject Traffic
Change the traffic load on a demo simulation. Omit traffic to trigger a random 2×–5× spike (sends random: true internally); provide traffic to set an absolute RPS level (capped at 10000 RPS in demo mode). Use it to stress-test the architecture before stepping; the change only affects metrics after the next simulation.step. Do not use it to read metrics (simulation.metrics) or advance time (simulation.step). Pass the simulationId returned by simulation.create when your connector opens a fresh MCP session; preserve Mcp-Session-Id to use the omitted-ID current-simulation default. Returns the updated simulation with its new traffic level; the likely next tool is simulation.step.
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'traffic': {'type': 'number', 'description': 'Absolute traffic level in RPS to set. Omit to trigger a random spike instead. Server-capped at 10000 RPS in demo mode.'}, 'simulationId': {'type': 'string', 'description': "Simulation ID returned by simulation.create. Preserve Mcp-Session-Id to omit this field and use the session's current simulation; if your connector starts a fresh MCP session for each call (for example Grok Bot or Cursor), pass this explicit ID after every fresh initialization. A fresh session has no current-simulation pointer and returns NO_ACTIVE_SIMULATION when the ID is omitted. Anonymous capabilities are short-lived (30 minutes by default), unguessable, and revoked when the demo expires or is deleted; proxy IP changes do not invalidate them. Do not treat the ID as a durable share link."}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'id': {'type': 'string', 'description': 'Simulation ID'}, 'status': {'type': 'string', 'description': 'Updated simulation status'}, 'traffic': {'type': 'number', 'description': 'New traffic level in RPS after injection'}}, 'additionalProperties': True}
simulation.metrics
Get Simulation Metrics
Read the latest metrics and resource states for a temporary anonymous demo simulation: latency, CPU, throughput, error rate, cost per hour, and per-resource health. Use it to inspect current state and metrics history without advancing time; do not use it to move the simulation forward — that is simulation.step. Responses are compact by default: principal current metrics plus explicit modeled goodputRps (a post-step point rate sourced from throughput, with provenance), goodputWindow (recorded only from persisted simulation-clock bounds, otherwise unavailable with provenance; never derive it from retrieval time or currentStep), errorBreakdown when available, per-resource status (id, name, status, cpuPercent, routedRps, availabilityState, isRoutable, and recoveryBlockedReason when provided), seeded EKS Spot checkpoint history and migrationEvaluationComplete/provenance when present, and the last 10 metrics-history entries. Pass responseMode: 'full' to get the complete simulation state, normalizedConfig, and full metrics history instead. During recovery, each resource may include recoveryProgress.state (parked, cooling_down, or healthy) with parkWindow and cooldown counters; poll simulation.step until healthy, then use simulation.metrics to inspect the resulting state and metrics. GPU / inference workflow: when the simulation includes a kubernetes resource with characteristics.inferenceMode: true, the response also includes top-level gpuUtilization (%), tokensPerSecond, costPerMillionTokens (USD/M tokens), idleGpuCostPerHour (USD/hr of standby GPU spend), and idleGpuFraction (0-1 idle HA overhead share) from the latest step, and each history entry carries the same inference fields. Pass the simulationId returned by simulation.create when your connector opens a fresh MCP session; preserve Mcp-Session-Id to use the omitted-ID current-simulation default. A fresh session has no current-simulation pointer. At least one simulation.step is needed for meaningful metrics. Read-only; repeated calls may be subject to demo usage limits. The likely next tool is simulation.step or simulation.inject_traffic.
Read only Idempotent
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'responseMode': {'enum': ['compact', 'full'], 'type': 'string', 'default': 'compact', 'description': "Response detail level. 'compact' (default) returns principal current metrics, errorBreakdown when available, per-resource status (id, name, status, cpuPercent, routedRps, availabilityState, isRoutable, recoveryBlockedReason, and failureLifecycle/routingState when provided), and only the last 10 metrics-history entries â\x80\x94 keeps polling cheap for agent loops. 'full' returns the complete simulation object (all resource characteristics and connections) plus the entire metrics history."}, 'simulationId': {'type': 'string', 'description': "Simulation ID returned by simulation.create. Preserve Mcp-Session-Id to omit this field and use the session's current simulation; if your connector starts a fresh MCP session for each call (for example Grok Bot or Cursor), pass this explicit ID after every fresh initialization. A fresh session has no current-simulation pointer and returns NO_ACTIVE_SIMULATION when the ID is omitted. Anonymous capabilities are short-lived (30 minutes by default), unguessable, and revoked when the demo expires or is deleted; proxy IP changes do not invalidate them. Do not treat the ID as a durable share link."}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'metrics': {'type': 'array', 'items': {'type': 'object', 'properties': {'cpuUsage': {'type': 'number', 'description': 'CPU utilization (%)'}, 'metricId': {'type': 'string', 'description': 'Storage-assigned persisted metric ID, when this history entry was stored'}, 'errorRate': {'type': 'number', 'description': 'Error rate (%)'}, 'latencyP50': {'type': ['number', 'null'], 'description': 'Median latency in ms; null when workload-specific latency evidence is unavailable.'}, 'latencyP95': {'type': ['number', 'null'], 'description': '95th-percentile latency in ms; null when workload-specific latency evidence is unavailable.'}, 'latencyP99': {'type': ['number', 'null'], 'description': 'Modeled 99th-percentile latency in ms; null when workload-specific latency evidence is unavailable.'}, 'throughput': {'type': 'number', 'description': 'Effective RPS'}, 'costPerHour': {'type': 'number', 'description': 'Estimated cost in USD/hr'}, 'latencyBasis': {'type': 'string', 'description': 'General modeled latency path/boundary; see latencyP99Basis for P99-specific provenance.'}, 'errorBreakdown': {'type': 'object', 'required': ['poolSaturation', 'dbFailure', 'computeFailure', 'capacityOverload', 'cpuOverload', 'ociStorage', 'queueAbsorption'], 'properties': {'dbFailure': {'type': 'number'}, 'ociStorage': {'type': 'number'}, 'cpuOverload': {'type': 'number'}, 'runtimeMemory': {'type': 'number'}, 'computeFailure': {'type': 'number'}, 'poolSaturation': {'type': 'number'}, 'queueAbsorption': {'type': 'number'}, 'capacityOverload': {'type': 'number'}, 'dependencyFailure': {'type': 'number'}}, 'description': 'Validated additive error contributors for this history entry', 'additionalProperties': False}, 'gpuUtilization': {'type': 'number', 'description': 'GPU utilization (%) for this step â\x80\x94 present only on GPU inference simulations'}, 'latencyP99Basis': {'type': 'string', 'description': 'Percentile-specific P99 basis for this history record.'}, 'tokensPerSecond': {'type': 'number', 'description': 'Inference throughput in tokens/second for this step â\x80\x94 present only on GPU inference simulations'}, 'latencyAvailability': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyAvailability'}, 'costPerMillionTokens': {'type': ['number', 'null'], 'description': 'Self-hosted inference cost in USD per million tokens for this step â\x80\x94 present only on GPU inference simulations'}, 'eksSpotInterruptions': {'type': 'array', 'items': {'$ref': '#/properties/eksSpotInterruptions/items'}}, 'retryAmplificationFactor': {'type': ['number', 'null'], 'description': 'Per-step retry amplification factor â\x80\x94 present only on resilience-enabled simulations'}}, 'additionalProperties': True}, 'description': 'Metrics history â\x80\x94 bounded to the last 10 entries in compact mode, full history in full mode'}, 'traffic': {'type': 'number', 'description': 'Current traffic level in RPS'}, 'metricId': {'type': 'string', 'description': 'Storage-assigned ID of the latest persisted metric'}, 'appWeight': {'enum': ['lean', 'typical', 'heavy'], 'type': 'string'}, 'errorRate': {'type': 'number', 'description': 'Latest error rate (%)'}, 'resources': {'type': 'array', 'items': {'type': 'object', 'properties': {'id': {'type': 'string', 'description': 'Resource ID'}, 'name': {'type': 'string', 'description': 'Resource display name'}, 'status': {'type': 'string', 'description': 'Health status (healthy/warning/critical/failed)'}, 'routedRps': {'type': 'number', 'description': 'Requests per second routed to this resource (compute/kubernetes only)'}, 'cpuPercent': {'type': 'number', 'description': 'CPU utilization (%)'}, 'isRoutable': {'type': 'boolean', 'description': 'Whether this compute/Kubernetes resource can receive traffic in this step; false distinguishes a failed/parked node from a critical node still serving'}, 'routingState': {'enum': ['unavailable', 'serving'], 'type': 'string', 'description': 'Read-only current routing state for a resource with failure lifecycle telemetry'}, 'failureLifecycle': {'enum': ['quick_injection_parked', 'quick_injection_rejoined', 'instance_down', 'instance_kill'], 'type': 'string', 'description': 'Read-only failure lifecycle marker when the backend identifies a quick injection or typed instance failure'}, 'recoveryProgress': {'type': 'object', 'required': ['state', 'parkWindow', 'cooldown'], 'properties': {'state': {'enum': ['parked', 'blocked', 'cooling_down', 'healthy'], 'type': 'string'}, 'cooldown': {'type': 'object', 'required': ['target', 'completedSteps', 'requiredSteps', 'remainingSteps'], 'properties': {'target': {'anyOf': [{'enum': ['warning', 'healthy'], 'type': 'string'}, {'type': 'null'}]}, 'requiredSteps': {'type': 'number'}, 'completedSteps': {'type': 'number'}, 'remainingSteps': {'type': 'number'}}, 'additionalProperties': False}, 'parkWindow': {'type': 'object', 'required': ['totalSteps', 'completedSteps', 'remainingSteps'], 'properties': {'totalSteps': {'type': 'number'}, 'completedSteps': {'type': 'number'}, 'remainingSteps': {'type': 'number'}}, 'additionalProperties': False}}, 'description': 'Read-only recovery progress; poll simulation.step until state is healthy, then use simulation.metrics to inspect the result', 'additionalProperties': False}, 'availabilityState': {'enum': ['available', 'degraded', 'unavailable'], 'type': 'string', 'description': 'Availability derived from routed traffic: available = serving normally, degraded = critical/warning but still serving, unavailable = no traffic and not routable'}, 'recoveryBlockedReason': {'type': 'string', 'description': 'Engine recovery guard currently preventing cooldown progress, when present (for example failure_park_window or idle_cpu_floor)'}}, 'additionalProperties': True}, 'description': 'Per-resource status summary (compact mode)'}, 'goodputRps': {'type': 'number', 'description': 'Modeled successful requests per second; a post-step point rate sourced from metrics.throughput'}, 'latencyP50': {'type': ['number', 'null'], 'description': 'Latest 50th-percentile latency in ms; null when workload-specific latency evidence is unavailable.'}, 'latencyP95': {'type': ['number', 'null'], 'description': 'Latest 95th-percentile latency in ms; null when workload-specific latency evidence is unavailable.'}, 'latencyP99': {'type': ['number', 'null'], 'description': 'Latest modeled 99th-percentile latency in ms; null when workload-specific latency evidence is unavailable. Interpret with latencyP99Basis and predictionEvidence.latencyP99.'}, 'offeredRps': {'type': 'number', 'description': 'Latest aggregate offered requests per second'}, 'simulation': {'type': 'object', 'properties': {'id': {'type': 'string', 'description': 'Simulation ID'}, 'name': {'type': 'string', 'description': 'Simulation name'}, 'traffic': {'type': 'number', 'description': 'Current RPS'}, 'resources': {'type': 'array', 'items': {'type': 'object', 'additionalProperties': {}}, 'description': 'Resource states with health and status'}, 'currentTime': {'type': 'number', 'description': 'Current time step'}}, 'description': 'Complete simulation state (full mode only; absent in compact mode and when status is not_found/access_denied)', 'additionalProperties': True}, 'throughput': {'type': 'number', 'description': 'Latest effective requests per second'}, 'costPerHour': {'type': 'number', 'description': 'Latest estimated cost in USD/hr'}, 'currentStep': {'type': 'number', 'description': 'Current simulation time step'}, 'latencyBasis': {'type': 'string', 'description': 'Latest general modeled latency path/boundary; see latencyP99Basis for P99-specific provenance.'}, 'scenarioHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Canonical replay scenario graph hash.'}, 'simulationId': {'type': 'string', 'description': 'ID of the queried simulation'}, 'engineVersion': {'type': 'string', 'description': 'Simulation engine version used for this prediction.'}, 'goodputWindow': {'anyOf': [{'type': 'object', 'required': ['status', 'provenance'], 'properties': {'status': {'type': 'string', 'const': 'unavailable'}, 'provenance': {'type': 'object', 'required': ['kind', 'reason'], 'properties': {'kind': {'type': 'string', 'const': 'unavailable'}, 'reason': {'type': 'string', 'maxLength': 256, 'minLength': 1}}, 'additionalProperties': False}}, 'additionalProperties': False}, {'type': 'object', 'required': ['status', 'window', 'aggregate', 'provenance'], 'properties': {'status': {'type': 'string', 'const': 'recorded'}, 'window': {'type': 'object', 'required': ['windowStartSeconds', 'windowEndSeconds', 'inclusion'], 'properties': {'inclusion': {'type': 'string', 'const': '[start,end)'}, 'windowEndSeconds': {'type': 'number', 'minimum': 0}, 'windowStartSeconds': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}, 'aggregate': {'type': 'object', 'required': ['window', 'coveredDurationSeconds', 'goodputRps', 'goodputRequestTotal', 'sampleIds', 'provenance'], 'properties': {'window': {'$ref': '#/properties/goodputWindow/anyOf/1/properties/window'}, 'sampleIds': {'type': 'array', 'items': {'type': 'string', 'minLength': 1}}, 'goodputRps': {'type': 'number', 'minimum': 0}, 'provenance': {'anyOf': [{'type': 'object', 'required': ['kind', 'source'], 'properties': {'kind': {'type': 'string', 'const': 'recorded'}, 'source': {'type': 'string', 'maxLength': 256, 'minLength': 1}}, 'additionalProperties': False}, {'type': 'object', 'required': ['kind', 'sourceFields'], 'properties': {'kind': {'type': 'string', 'const': 'derived'}, 'sourceFields': {'type': 'array', 'items': {'type': 'string', 'maxLength': 256, 'minLength': 1}, 'maxItems': 256, 'minItems': 1}}, 'additionalProperties': False}, {'type': 'object', 'required': ['kind', 'reason'], 'properties': {'kind': {'type': 'string', 'const': 'unavailable'}, 'reason': {'type': 'string', 'maxLength': 256, 'minLength': 1}}, 'additionalProperties': False}]}, 'goodputRequestTotal': {'type': 'number', 'minimum': 0}, 'coveredDurationSeconds': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}, 'provenance': {'type': 'object', 'required': ['kind', 'source'], 'properties': {'kind': {'type': 'string', 'const': 'recorded'}, 'source': {'type': 'string', 'maxLength': 256, 'minLength': 1}}, 'additionalProperties': False}}, 'additionalProperties': False}], 'description': 'Interval goodput aggregate when every point has persisted simulation-clock bounds; otherwise status=unavailable. Never derive this from retrieval time or currentStep.'}, 'errorBreakdown': {'type': 'object', 'required': ['poolSaturation', 'dbFailure', 'computeFailure', 'capacityOverload', 'cpuOverload', 'ociStorage', 'queueAbsorption'], 'properties': {'dbFailure': {'type': 'number'}, 'ociStorage': {'type': 'number'}, 'cpuOverload': {'type': 'number'}, 'runtimeMemory': {'type': 'number'}, 'computeFailure': {'type': 'number'}, 'poolSaturation': {'type': 'number'}, 'queueAbsorption': {'type': 'number'}, 'capacityOverload': {'type': 'number'}, 'dependencyFailure': {'type': 'number'}}, 'description': 'Validated additive error contributors from the latest metrics entry, in percentage-point units', 'additionalProperties': False}, 'gpuUtilization': {'type': 'number', 'description': 'Latest GPU utilization (%) â\x80\x94 present only on GPU inference simulations'}, 'modeledShedRps': {'type': 'number', 'description': 'Latest aggregate modeled shed requests per second'}, 'replayIdentity': {'type': 'object', 'required': ['scenarioHash', 'effectiveConfigHash'], 'properties': {'scenarioHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Canonical SHA-256 of the persisted scenario graph and attached traffic-pattern order'}, 'effectiveConfigHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Original replay-only SHA-256 of the effective six-control startup configuration; top-level effectiveConfigHash is the versioned prediction identity.'}}, 'additionalProperties': False}, 'auroraFailovers': {'type': 'array', 'items': {'type': 'object', 'required': ['failedResourceId', 'standbyResourceId', 'succeeded', 'phase'], 'properties': {'phase': {'enum': ['promoting', 'serving', 'unavailable'], 'type': 'string'}, 'succeeded': {'type': 'boolean'}, 'failedResourceId': {'type': 'string'}, 'standbyResourceId': {'type': 'string'}}, 'additionalProperties': False}, 'description': 'Latest compact per-writer Aurora failover state, with failed and standby IDs and phase.'}, 'idleGpuFraction': {'type': 'number', 'description': 'Latest share (0-1) of the GPU bill that is idle/standby capacity â\x80\x94 present only on GPU inference simulations'}, 'latencyP99Basis': {'type': 'string', 'description': 'Latest percentile-specific P99 basis: owned fit, scaled-from-fit (not directly measured), or uncalibrated generic model.'}, 'tokensPerSecond': {'type': 'number', 'description': 'Latest inference throughput in tokens/second â\x80\x94 present only on GPU inference simulations'}, 'goodputSemantics': {'type': 'string', 'const': 'post_step_point_rate', 'description': 'Goodput is a point rate, not an interval total'}, 'goodputProvenance': {'anyOf': [{'type': 'object', 'required': ['kind'], 'properties': {'kind': {'type': 'string', 'const': 'recorded'}}, 'additionalProperties': False}, {'type': 'object', 'required': ['kind', 'sourceFields'], 'properties': {'kind': {'type': 'string', 'const': 'derived'}, 'sourceFields': {'type': 'array', 'items': {'type': 'string', 'minLength': 1}, 'minItems': 1}}, 'additionalProperties': False}, {'type': 'object', 'required': ['kind', 'reason'], 'properties': {'kind': {'type': 'string', 'const': 'unavailable'}, 'reason': {'type': 'string', 'minLength': 1}}, 'additionalProperties': False}], 'description': 'Provenance for the modeled goodput field'}, 'appWeightDefaulted': {'type': 'boolean'}, 'idleGpuCostPerHour': {'type': 'number', 'description': 'Latest USD/hr of GPU spend funding idle/standby capacity (HA overhead) â\x80\x94 present only on GPU inference simulations'}, 'predictionEvidence': {'anyOf': [{'type': 'object', 'required': ['version', 'evidenceLevel', 'legacyGeneric', 'appWeight', 'appWeightDefaulted', 'loadScope', 'sourceIds', 'formulaIds', 'appCpu', 'appCpuByResource', 'latencyP50', 'latencyP95', 'note'], 'properties': {'note': {'type': 'string'}, 'appCpu': {'anyOf': [{'type': 'object', 'required': ['low', 'central', 'high'], 'properties': {'low': {'type': 'number', 'minimum': 0}, 'high': {'type': 'number', 'minimum': 0}, 'central': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}, {'type': 'null'}]}, 'version': {'type': 'number', 'const': 1}, 'appWeight': {'enum': ['lean', 'typical', 'heavy'], 'type': 'string'}, 'loadScope': {'enum': ['within measured load', 'beyond measured load', 'below measured load', 'not calibrated'], 'type': 'string'}, 'sourceIds': {'type': 'array', 'items': {'type': 'string'}}, 'formulaIds': {'type': 'array', 'items': {'type': 'string'}}, 'latencyP50': {'anyOf': [{'type': 'object', 'required': ['low', 'central', 'high', 'evidenceLevel'], 'properties': {'low': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/low'}, 'high': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/high'}, 'central': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/central'}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}}, 'additionalProperties': False}, {'type': 'null'}]}, 'latencyP95': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'latencyP99': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'evidenceLevel': {'enum': ['measured', 'scaled from measured', 'reference estimate'], 'type': 'string'}, 'legacyGeneric': {'type': 'boolean'}, 'appCpuByResource': {'type': 'array', 'items': {'type': 'object', 'required': ['low', 'central', 'high', 'evidenceLevel', 'resourceId', 'legacyGeneric'], 'properties': {'low': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/low'}, 'high': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/high'}, 'central': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/central'}, 'resourceId': {'type': 'string'}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}, 'legacyGeneric': {'type': 'boolean'}}, 'additionalProperties': False}}, 'appWeightDefaulted': {'type': 'boolean'}, 'latencyAvailability': {'type': 'object', 'required': ['status', 'reason', 'resourceIds'], 'properties': {'reason': {'type': 'string', 'minLength': 1}, 'status': {'type': 'string', 'const': 'unavailable'}, 'resourceIds': {'type': 'array', 'items': {'type': 'string', 'minLength': 1}}}, 'additionalProperties': False}, 'requestServingCapacity': {'type': 'array', 'items': {'type': 'object', 'required': ['resourceId', 'resourceName', 'taskCount', 'perTaskCapacityRps', 'aggregateCapacityRps', 'maxTaskCount', 'maxAggregateCapacityRps', 'capacityEvidence'], 'properties': {'taskCount': {'type': 'integer', 'minimum': 0}, 'resourceId': {'type': 'string'}, 'maxTaskCount': {'type': 'integer', 'exclusiveMinimum': 0}, 'resourceName': {'type': 'string'}, 'capacityEvidence': {'type': 'object', 'required': ['basis', 'origin'], 'properties': {'basis': {'enum': ['assumption', 'documented', 'measured', 'catalog'], 'type': 'string'}, 'origin': {'enum': ['caller-provided', 'fargate-size-heuristic', 'policy-default', 'provider-catalog'], 'type': 'string'}, 'source': {'type': 'string'}}, 'additionalProperties': False}, 'perTaskCapacityRps': {'type': 'number', 'exclusiveMinimum': 0}, 'aggregateCapacityRps': {'type': 'number', 'minimum': 0}, 'maxAggregateCapacityRps': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}}, 'latencyPercentileStatus': {'enum': ['modeled', 'unavailable'], 'type': 'string'}}, 'additionalProperties': False}, {'type': 'object', 'required': ['version', 'evidenceLevel', 'appWeight', 'appWeightDefaulted', 'loadScope', 'sourceIds', 'formulaIds', 'appCpu', 'appCpuByResource', 'latencyP50', 'latencyP95', 'note'], 'properties': {'note': {'type': 'string'}, 'appCpu': {'type': 'null'}, 'version': {'type': 'number', 'const': 1}, 'appWeight': {'type': 'null'}, 'loadScope': {'type': 'null'}, 'sourceIds': {'type': 'array', 'items': {'type': 'string'}}, 'formulaIds': {'type': 'array', 'items': {'type': 'string'}}, 'latencyP50': {'type': 'null'}, 'latencyP95': {'type': 'null'}, 'latencyP99': {'type': 'null'}, 'evidenceLevel': {'type': 'string', 'const': 'unavailable/legacy'}, 'legacyGeneric': {'type': 'null'}, 'appCpuByResource': {'type': 'array', 'items': {'type': 'object', 'required': ['low', 'central', 'high', 'evidenceLevel', 'resourceId'], 'properties': {'low': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/low'}, 'high': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/high'}, 'central': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/central'}, 'resourceId': {'type': 'string'}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}}, 'additionalProperties': False}, 'maxItems': 0}, 'appWeightDefaulted': {'type': 'null'}}, 'additionalProperties': False}, {'type': 'object', 'required': ['version', 'evidenceLevel', 'appWeight', 'appWeightDefaulted', 'loadScope', 'sourceIds', 'formulaIds', 'appCpu', 'appCpuByResource', 'latencyP50', 'latencyP95', 'note'], 'properties': {'note': {'type': 'string'}, 'appCpu': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0'}, {'type': 'null'}]}, 'version': {'type': 'number', 'const': 1}, 'appWeight': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appWeight'}, 'loadScope': {'enum': ['within measured load', 'beyond measured load', 'below measured load', 'not calibrated'], 'type': 'string'}, 'sourceIds': {'type': 'array', 'items': {'type': 'string'}}, 'formulaIds': {'type': 'array', 'items': {'type': 'string'}}, 'latencyP50': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'latencyP95': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'latencyP99': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}, 'legacyGeneric': {'type': 'boolean'}, 'appCpuByResource': {'type': 'array', 'items': {'type': 'object', 'required': ['low', 'central', 'high', 'evidenceLevel', 'resourceId'], 'properties': {'low': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/low'}, 'high': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/high'}, 'central': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/central'}, 'resourceId': {'type': 'string'}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}, 'legacyGeneric': {'type': 'boolean'}}, 'additionalProperties': False}}, 'appWeightDefaulted': {'type': 'boolean'}, 'latencyAvailability': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyAvailability'}, 'latencyPercentileStatus': {'enum': ['modeled', 'unavailable'], 'type': 'string'}}, 'additionalProperties': False}]}, 'calibrationEvidence': {'type': 'object', 'required': ['kind', 'latencyBasis', 'note'], 'properties': {'kind': {'enum': ['owned', 'modeled'], 'type': 'string', 'description': 'owned for the exact AWS CRUD fit; modeled when the gate fails or another generic model applies.'}, 'note': {'type': 'string', 'description': "Names the fit scope and limitations. For canonical workload inference, states it is a modeling assumption. For modeled fallback, names all actual failed checks in stable order with resource IDs and safe expected/actual values; never only a generic 'not applied'."}, 'latencyBasis': {'type': 'string', 'description': 'General modeled latency path, such as in-VPC ALB rather than end-to-end; see latencyP99Basis for percentile-specific P99 provenance.'}, 'calibrationId': {'type': 'string', 'description': 'Versioned owned calibration identifier; present only when that calibration applies.'}, 'latencyP99Basis': {'type': 'string', 'description': 'Percentile-specific P99 basis: owned in-VPC internal-ALB fit, scaled from that fit (not directly measured), or uncalibrated generic model.'}}, 'description': 'Owned-versus-modeled evidence and latency boundary.', 'additionalProperties': False}, 'effectiveConfigHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Versioned prediction hash over replay startup inputs, engine version, and calibration identity.'}, 'latencyAvailability': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyAvailability'}, 'costPerMillionTokens': {'type': ['number', 'null'], 'description': 'Latest self-hosted inference cost in USD per million tokens (null when no tokens are being processed) â\x80\x94 present only on GPU inference simulations'}, 'eksSpotInterruptions': {'type': 'array', 'items': {'type': 'object', 'properties': {'name': {'type': 'string'}, 'resourceId': {'type': 'string'}, 'checkpointEvidence': {'type': 'object', 'required': ['engineInputStepIndex', 'simulationSeconds', 'tickDurationSeconds', 'pointRateSemantics', 'intervalAttribution', 'fieldProvenance', 'trafficFieldProvenance'], 'properties': {'fieldProvenance': {'type': 'object', 'required': ['engineInputStepIndex', 'simulationSeconds', 'tickDurationSeconds', 'pointRateSemantics', 'intervalAttribution'], 'properties': {'simulationSeconds': {'$ref': '#/properties/goodputProvenance'}, 'pointRateSemantics': {'$ref': '#/properties/goodputProvenance'}, 'intervalAttribution': {'$ref': '#/properties/goodputProvenance'}, 'tickDurationSeconds': {'$ref': '#/properties/goodputProvenance'}, 'engineInputStepIndex': {'$ref': '#/properties/goodputProvenance'}}, 'additionalProperties': False}, 'simulationSeconds': {'type': 'number', 'minimum': 0}, 'pointRateSemantics': {'type': 'string', 'const': 'post_step_point_rate'}, 'intervalAttribution': {'anyOf': [{'anyOf': [{'type': 'object', 'required': ['status', 'window', 'provenance'], 'properties': {'status': {'type': 'string', 'const': 'recorded'}, 'window': {'$ref': '#/properties/goodputWindow/anyOf/1/properties/window'}, 'provenance': {'type': 'object', 'required': ['kind', 'source'], 'properties': {'kind': {'type': 'string', 'const': 'recorded'}, 'source': {'type': 'string', 'maxLength': 256, 'minLength': 1}}, 'additionalProperties': False}}, 'additionalProperties': False}, {'type': 'object', 'required': ['status', 'provenance'], 'properties': {'status': {'type': 'string', 'const': 'unavailable'}, 'provenance': {'type': 'object', 'required': ['kind', 'reason'], 'properties': {'kind': {'type': 'string', 'const': 'unavailable'}, 'reason': {'type': 'string', 'maxLength': 256, 'minLength': 1}}, 'additionalProperties': False}}, 'additionalProperties': False}]}, {'type': 'null'}]}, 'tickDurationSeconds': {'type': 'number', 'exclusiveMinimum': 0}, 'engineInputStepIndex': {'type': 'integer', 'minimum': 0}, 'trafficFieldProvenance': {'type': 'object', 'required': ['throughputRps', 'offeredRps', 'errorRatePercent', 'latencyP50Ms'], 'properties': {'offeredRps': {'$ref': '#/properties/goodputProvenance'}, 'latencyP50Ms': {'$ref': '#/properties/goodputProvenance'}, 'throughputRps': {'$ref': '#/properties/goodputProvenance'}, 'errorRatePercent': {'$ref': '#/properties/goodputProvenance'}}, 'additionalProperties': False}}, 'description': 'Persisted seeded-EKS checkpoint captured with its interruption JSONB record. engineInputStepIndex is an engine input-step index, simulationSeconds is derived from that index and the recorded tick, and traffic provenance points to outer metric fields rather than duplicating float values.', 'additionalProperties': False}, 'migrationEvaluation': {'type': 'object', 'required': ['status', 'interruptionNoticeAtSimulationSeconds', 'deadlineAtSimulationSeconds', 'migrationStartedAtSimulationSeconds', 'allAffectedWorkloadsReadyAtSimulationSeconds', 'migrationDurationSeconds', 'deadlineSeconds', 'deadlineMet', 'affectedWorkloadCount', 'readyWorkloadCountAtDeadline', 'missedDeadlineWorkloadCount', 'evaluatedAtSimulationSeconds', 'limitingReasons', 'fieldProvenance'], 'properties': {'status': {'enum': ['not_started', 'in_progress', 'completed'], 'type': 'string'}, 'deadlineMet': {'type': ['boolean', 'null']}, 'deadlineSeconds': {'type': 'number', 'const': 120}, 'fieldProvenance': {'type': 'object', 'required': ['status', 'interruptionNoticeAtSimulationSeconds', 'deadlineAtSimulationSeconds', 'migrationStartedAtSimulationSeconds', 'allAffectedWorkloadsReadyAtSimulationSeconds', 'migrationDurationSeconds', 'deadlineSeconds', 'deadlineMet', 'affectedWorkloadCount', 'readyWorkloadCountAtDeadline', 'missedDeadlineWorkloadCount', 'evaluatedAtSimulationSeconds', 'limitingReasons'], 'properties': {'status': {'$ref': '#/properties/goodputProvenance'}, 'deadlineMet': {'$ref': '#/properties/goodputProvenance'}, 'deadlineSeconds': {'$ref': '#/properties/goodputProvenance'}, 'limitingReasons': {'$ref': '#/properties/goodputProvenance'}, 'affectedWorkloadCount': {'$ref': '#/properties/goodputProvenance'}, 'migrationDurationSeconds': {'$ref': '#/properties/goodputProvenance'}, 'deadlineAtSimulationSeconds': {'$ref': '#/properties/goodputProvenance'}, 'missedDeadlineWorkloadCount': {'$ref': '#/properties/goodputProvenance'}, 'evaluatedAtSimulationSeconds': {'$ref': '#/properties/goodputProvenance'}, 'readyWorkloadCountAtDeadline': {'$ref': '#/properties/goodputProvenance'}, 'migrationStartedAtSimulationSeconds': {'$ref': '#/properties/goodputProvenance'}, 'interruptionNoticeAtSimulationSeconds': {'$ref': '#/properties/goodputProvenance'}, 'allAffectedWorkloadsReadyAtSimulationSeconds': {'$ref': '#/properties/goodputProvenance'}}, 'additionalProperties': False}, 'limitingReasons': {'type': 'array', 'items': {'type': 'object', 'required': ['code', 'workloadCount', 'pendingWorkloadCount', 'pullingWorkloadCount', 'startingWorkloadCount', 'residualPullSeconds', 'residualStartupSeconds', 'schedulingCapacity', 'interruptionHandling', 'provenance'], 'properties': {'code': {'enum': ['no_reschedule', 'scheduling_capacity_exhausted', 'image_pull_incomplete', 'startup_incomplete', 'unclassified_runtime_work_remaining'], 'type': 'string'}, 'provenance': {'type': 'string', 'const': 'recorded'}, 'workloadCount': {'type': 'integer', 'minimum': 0}, 'schedulingCapacity': {'type': 'integer', 'minimum': 0}, 'residualPullSeconds': {'type': 'number', 'minimum': 0}, 'interruptionHandling': {'enum': ['reschedule', 'drain-only', 'fail-fast'], 'type': 'string'}, 'pendingWorkloadCount': {'type': 'integer', 'minimum': 0}, 'pullingWorkloadCount': {'type': 'integer', 'minimum': 0}, 'startingWorkloadCount': {'type': 'integer', 'minimum': 0}, 'residualStartupSeconds': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}, 'maxItems': 5}, 'affectedWorkloadCount': {'anyOf': [{'type': 'integer', 'minimum': 0}, {'type': 'null'}]}, 'migrationDurationSeconds': {'anyOf': [{'type': 'number', 'minimum': 0}, {'type': 'null'}]}, 'deadlineAtSimulationSeconds': {'anyOf': [{'type': 'number', 'minimum': 0}, {'type': 'null'}]}, 'missedDeadlineWorkloadCount': {'anyOf': [{'type': 'integer', 'minimum': 0}, {'type': 'null'}]}, 'evaluatedAtSimulationSeconds': {'anyOf': [{'type': 'number', 'minimum': 0}, {'type': 'null'}]}, 'readyWorkloadCountAtDeadline': {'anyOf': [{'type': 'integer', 'minimum': 0}, {'type': 'null'}]}, 'migrationStartedAtSimulationSeconds': {'anyOf': [{'type': 'number', 'minimum': 0}, {'type': 'null'}]}, 'interruptionNoticeAtSimulationSeconds': {'anyOf': [{'type': 'number', 'minimum': 0}, {'type': 'null'}]}, 'allAffectedWorkloadsReadyAtSimulationSeconds': {'anyOf': [{'type': 'number', 'minimum': 0}, {'type': 'null'}]}}, 'description': 'Authoritative additive seeded EKS interruption evaluation. Field values carry recorded/derived/unavailable provenance; deadline verdict/counts/reasons are frozen once status is completed and do not describe later service health.', 'additionalProperties': False}}, 'additionalProperties': True}, 'description': 'Latest seeded EKS interruption telemetry, including additive migrationEvaluation/provenance when recorded.'}, 'metricsHistoryLength': {'type': 'number', 'description': 'Total number of metrics-history entries (compact mode returns only the last 10)'}, 'resilienceDiagnostics': {'type': 'object', 'properties': {'bounded': {'type': 'boolean', 'description': 'True when a generated-traffic, traversal-work, or cascade-depth bound truncated model work'}, 'pathCount': {'type': 'number', 'description': 'Number of dependency paths evaluated in the latest step'}, 'incidentOutcome': {'type': 'string', 'description': 'Latest incident outcome: stable | degraded | cascading | protected | recovered'}}, 'description': 'Bounded resilience diagnostics from the latest step (compact mode). Absent when the resilience model did not run.', 'additionalProperties': True}, 'retryAmplificationFactor': {'type': ['number', 'null'], 'description': 'Latest retry amplification factor (attemptedRps / originalRps). Values > 1.0 = amplification risk. null = model ran but no traffic. Absent = resilience model disabled.'}, 'predictionEffectiveConfigHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Versioned prediction hash over replay startup inputs, engine version, and calibration identity.'}}, 'description': "Compact metrics summary by default (principal current metrics, per-resource status, bounded history tail). With responseMode 'full', the complete simulation object and full metrics history are passed through instead.", 'additionalProperties': True}
simulation.recover_resource
Recover Failed Resource
Recover one reversible failed resource in a temporary anonymous demo simulation. No API key required for this temporary anonymous demo operation. Lower traffic to a serviceable level first, then provide resourceId or resourceName from simulation.create, simulation.step, or simulation.metrics. This deactivates applicable instance_down/database_overload failures for only the selected resource and returns recoveryProgress with parked, cooling_down, or healthy state plus cooldown counters. It cannot restore an instance_kill because that failure permanently removes the resource. The likely next tool is simulation.step; keep stepping and inspect the targeted resource until recoveryProgress.state is healthy. Pass simulationId from simulation.create when using a fresh MCP session; a preserved session may omit it. Authenticate with an API key to unlock all 62 tools and unlimited simulations.
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'resourceId': {'type': 'string', 'description': 'ID of the failed resource to recover'}, 'resourceName': {'type': 'string', 'description': 'Exact case-insensitive name of the failed resource to recover'}, 'simulationId': {'type': 'string', 'description': "Simulation ID returned by simulation.create. Preserve Mcp-Session-Id to omit this field and use the session's current simulation; if your connector starts a fresh MCP session for each call (for example Grok Bot or Cursor), pass this explicit ID after every fresh initialization. A fresh session has no current-simulation pointer and returns NO_ACTIVE_SIMULATION when the ID is omitted. Anonymous capabilities are short-lived (30 minutes by default), unguessable, and revoked when the demo expires or is deleted; proxy IP changes do not invalidate them. Do not treat the ID as a durable share link."}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'simulationId': {'type': 'string'}, 'recoveryState': {'type': 'string'}, 'previousHealth': {'type': 'string'}, 'stepsToHealthy': {'type': 'number'}, 'recoveryProgress': {'type': 'object', 'required': ['state', 'parkWindow', 'cooldown'], 'properties': {'state': {'enum': ['parked', 'blocked', 'cooling_down', 'healthy'], 'type': 'string'}, 'cooldown': {'type': 'object', 'required': ['target', 'completedSteps', 'requiredSteps', 'remainingSteps'], 'properties': {'target': {'anyOf': [{'enum': ['warning', 'healthy'], 'type': 'string'}, {'type': 'null'}]}, 'requiredSteps': {'type': 'number'}, 'completedSteps': {'type': 'number'}, 'remainingSteps': {'type': 'number'}}, 'additionalProperties': False}, 'parkWindow': {'type': 'object', 'required': ['totalSteps', 'completedSteps', 'remainingSteps'], 'properties': {'totalSteps': {'type': 'number'}, 'completedSteps': {'type': 'number'}, 'remainingSteps': {'type': 'number'}}, 'additionalProperties': False}}, 'additionalProperties': False}, 'resolvedResourceId': {'type': 'string'}, 'simulationIdSource': {'enum': ['explicit', 'session_default'], 'type': 'string'}, 'resolvedResourceName': {'type': 'string'}, 'deactivatedFailureIds': {'type': 'array', 'items': {'type': 'string'}}, 'stepsToHealthyIsLowerBound': {'type': 'boolean'}}, 'additionalProperties': True}
simulation.step
Simulate Step
Advance a temporary anonymous demo simulation by one time step and return updated metrics — CPU, latency, throughput, error rate, cost (max 20 persisted steps per demo). Concurrent simulation.step calls on one simulation either serialize as distinct consecutive steps (in the browser Workspace queue) or receive HTTP 409 simulation_step_in_progress without advancing or consuming a demo credit. Wait for the running call to finish, then retry only rejected calls; distinct simulations can run concurrently. After a timeout, inspect simulation.metrics before retrying an uncertain result. Use it to observe how the architecture behaves over time, typically right after simulation.create or simulation.inject_traffic. Do not use it to read current state without advancing time — that is simulation.metrics. Pass the simulationId returned by simulation.create when your connector opens a fresh MCP session; preserve Mcp-Session-Id to use the omitted-ID current-simulation default. The likely next tool is simulation.step again (to keep observing) or simulation.inject_traffic (to change load first). Responses are compact by default: principal metrics plus per-resource status (id, name, status, cpuPercent, routedRps, availabilityState, isRoutable, and recoveryBlockedReason when provided) and this step's events. Seeded characteristics.eksSpotInterruption telemetry retains its additive migrationEvaluation beside the interruption lifecycle: use its recorded/derived/unavailable field provenance, frozen deadline verdict/counts/reasons, and simulation-clock milestones rather than final service health. The distinct eksSpotMigration contract remains separately reported when configured. Compact responses also include errorBreakdown when the engine provides it. A critical resource with isRoutable: true is degraded but still serving; availabilityState: unavailable and isRoutable: false identify a failed or parked node. Pass responseMode: 'full' to get the complete simulation state instead. During recovery, each resource may include recoveryProgress with state parked, cooling_down, or healthy, plus parkWindow and cooldown counters. Poll simulation.step until the targeted resource's recoveryProgress.state is healthy, then use simulation.metrics to inspect the resulting state and metrics. GPU / inference workflow: when the simulation includes a kubernetes resource with characteristics.inferenceMode: true, each step response also includes gpuUtilization (%), tokensPerSecond, costPerMillionTokens (USD/M tokens), idleGpuCostPerHour (USD/hr of standby GPU spend), and idleGpuFraction (0-1 idle HA overhead share) so you can track inference economics step by step. Authenticate with an API key for unlimited steps and GPU right-sizing hints.
Input schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'responseMode': {'enum': ['compact', 'full'], 'type': 'string', 'default': 'compact', 'description': "Response detail level. 'compact' (default) returns only principal metrics, errorBreakdown when available, per-resource status (id, name, status, cpuPercent, routedRps, availabilityState, isRoutable, recoveryBlockedReason, and failureLifecycle/routingState when provided), and this step's events â\x80\x94 keeps observations small for agent loops. 'full' returns the complete backend step response including the entire simulation object with all resource characteristics and connections."}, 'simulationId': {'type': 'string', 'description': "Simulation ID returned by simulation.create. Preserve Mcp-Session-Id to omit this field and use the session's current simulation; if your connector starts a fresh MCP session for each call (for example Grok Bot or Cursor), pass this explicit ID after every fresh initialization. A fresh session has no current-simulation pointer and returns NO_ACTIVE_SIMULATION when the ID is omitted. Anonymous capabilities are short-lived (30 minutes by default), unguessable, and revoked when the demo expires or is deleted; proxy IP changes do not invalidate them. Do not treat the ID as a durable share link."}}, 'additionalProperties': False}
Output schema
{'type': 'object', '$schema': 'http://json-schema.org/draft-07/schema#', 'properties': {'events': {'type': 'array', 'items': {'type': 'object', 'additionalProperties': {}}, 'description': 'Events generated during this step'}, 'metrics': {'type': 'object', 'properties': {'latencyP50': {'type': ['number', 'null'], 'description': '50th-percentile latency in ms; null when workload-specific latency evidence is unavailable.'}, 'latencyP95': {'type': ['number', 'null'], 'description': '95th-percentile latency in ms; null when workload-specific latency evidence is unavailable.'}, 'latencyP99': {'type': ['number', 'null'], 'description': "Modeled 99th-percentile latency in ms; null when workload-specific latency evidence is unavailable. Interpret with this metric's latencyP99Basis and predictionEvidence.latencyP99."}, 'latencyBasis': {'type': 'string', 'description': 'General modeled latency path/boundary; see latencyP99Basis for P99-specific provenance.'}, 'latencyP99Basis': {'type': 'string', 'description': 'Percentile-specific P99 basis for this metric checkpoint.'}, 'eksSpotInterruptions': {'type': 'array', 'items': {'$ref': '#/properties/eksSpotInterruptions/items'}}}, 'description': 'Full-mode backend metric record; migrationEvaluation is exact-typed when a seeded interruption is present.', 'additionalProperties': True}, 'traffic': {'type': 'number', 'description': 'Current traffic level in RPS'}, 'metricId': {'type': 'string', 'description': 'Storage-assigned persisted metric ID when available'}, 'appWeight': {'enum': ['lean', 'typical', 'heavy'], 'type': 'string'}, 'errorRate': {'type': 'number', 'description': 'Error rate (%)'}, 'resources': {'type': 'array', 'items': {'type': 'object', 'properties': {'id': {'type': 'string', 'description': 'Resource ID'}, 'name': {'type': 'string', 'description': 'Resource display name'}, 'status': {'type': 'string', 'description': 'Health status (healthy/warning/critical/failed)'}, 'routedRps': {'type': 'number', 'description': 'Requests per second routed to this resource this step (compute/kubernetes only)'}, 'cpuPercent': {'type': 'number', 'description': 'CPU utilization (%)'}, 'isRoutable': {'type': 'boolean', 'description': 'Whether this compute/Kubernetes resource can receive traffic in this step; false distinguishes a failed/parked node from a critical node still serving'}, 'routingState': {'enum': ['unavailable', 'serving'], 'type': 'string', 'description': 'Read-only current routing state for a resource with failure lifecycle telemetry'}, 'failureLifecycle': {'enum': ['quick_injection_parked', 'quick_injection_rejoined', 'instance_down', 'instance_kill'], 'type': 'string', 'description': 'Read-only failure lifecycle marker when the backend identifies a quick injection or typed instance failure'}, 'recoveryProgress': {'type': 'object', 'required': ['state', 'parkWindow', 'cooldown'], 'properties': {'state': {'enum': ['parked', 'blocked', 'cooling_down', 'healthy'], 'type': 'string'}, 'cooldown': {'type': 'object', 'required': ['target', 'completedSteps', 'requiredSteps', 'remainingSteps'], 'properties': {'target': {'anyOf': [{'enum': ['warning', 'healthy'], 'type': 'string'}, {'type': 'null'}]}, 'requiredSteps': {'type': 'number'}, 'completedSteps': {'type': 'number'}, 'remainingSteps': {'type': 'number'}}, 'additionalProperties': False}, 'parkWindow': {'type': 'object', 'required': ['totalSteps', 'completedSteps', 'remainingSteps'], 'properties': {'totalSteps': {'type': 'number'}, 'completedSteps': {'type': 'number'}, 'remainingSteps': {'type': 'number'}}, 'additionalProperties': False}}, 'description': 'Read-only recovery progress; poll simulation.step until state is healthy, then use simulation.metrics to inspect the result', 'additionalProperties': False}, 'availabilityState': {'enum': ['available', 'degraded', 'unavailable'], 'type': 'string', 'description': 'Availability derived from routed traffic: available = serving normally, degraded = critical/warning but still serving, unavailable = no traffic and not routable'}, 'recoveryBlockedReason': {'type': 'string', 'description': 'Engine recovery guard currently preventing cooldown progress, when present (for example failure_park_window or idle_cpu_floor)'}}, 'additionalProperties': True}, 'description': 'Per-resource status summary (compact mode)'}, 'goodputRps': {'type': 'number', 'description': 'Modeled successful requests per second; a post-step point rate sourced from metrics.throughput'}, 'latencyP50': {'type': ['number', 'null'], 'description': '50th-percentile latency in ms; null when workload-specific latency evidence is unavailable.'}, 'latencyP95': {'type': ['number', 'null'], 'description': '95th-percentile latency in ms; null when workload-specific latency evidence is unavailable.'}, 'latencyP99': {'type': ['number', 'null'], 'description': 'Modeled 99th-percentile latency in ms; null when workload-specific latency evidence is unavailable. Interpret with latencyP99Basis and predictionEvidence.latencyP99.'}, 'offeredRps': {'type': 'number', 'description': 'Aggregate offered requests per second represented by metrics.offeredRps provenance'}, 'throughput': {'type': 'number', 'description': 'Effective requests per second'}, 'costPerHour': {'type': 'number', 'description': 'Estimated cost in USD/hr'}, 'currentStep': {'type': 'number', 'description': 'New simulation time step index'}, 'latencyBasis': {'type': 'string', 'description': 'General modeled latency path/boundary; see latencyP99Basis for P99-specific provenance.'}, 'scenarioHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Canonical SHA-256 of the persisted scenario graph and attached traffic-pattern order'}, 'simulationId': {'type': 'string', 'description': 'ID of the stepped simulation'}, 'engineVersion': {'type': 'string', 'description': 'Simulation engine version used for this prediction.'}, 'goodputWindow': {'anyOf': [{'type': 'object', 'required': ['status', 'provenance'], 'properties': {'status': {'type': 'string', 'const': 'unavailable'}, 'provenance': {'type': 'object', 'required': ['kind', 'reason'], 'properties': {'kind': {'type': 'string', 'const': 'unavailable'}, 'reason': {'type': 'string', 'maxLength': 256, 'minLength': 1}}, 'additionalProperties': False}}, 'additionalProperties': False}, {'type': 'object', 'required': ['status', 'window', 'aggregate', 'provenance'], 'properties': {'status': {'type': 'string', 'const': 'recorded'}, 'window': {'type': 'object', 'required': ['windowStartSeconds', 'windowEndSeconds', 'inclusion'], 'properties': {'inclusion': {'type': 'string', 'const': '[start,end)'}, 'windowEndSeconds': {'type': 'number', 'minimum': 0}, 'windowStartSeconds': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}, 'aggregate': {'type': 'object', 'required': ['window', 'coveredDurationSeconds', 'goodputRps', 'goodputRequestTotal', 'sampleIds', 'provenance'], 'properties': {'window': {'$ref': '#/properties/goodputWindow/anyOf/1/properties/window'}, 'sampleIds': {'type': 'array', 'items': {'type': 'string', 'minLength': 1}}, 'goodputRps': {'type': 'number', 'minimum': 0}, 'provenance': {'anyOf': [{'type': 'object', 'required': ['kind', 'source'], 'properties': {'kind': {'type': 'string', 'const': 'recorded'}, 'source': {'type': 'string', 'maxLength': 256, 'minLength': 1}}, 'additionalProperties': False}, {'type': 'object', 'required': ['kind', 'sourceFields'], 'properties': {'kind': {'type': 'string', 'const': 'derived'}, 'sourceFields': {'type': 'array', 'items': {'type': 'string', 'maxLength': 256, 'minLength': 1}, 'maxItems': 256, 'minItems': 1}}, 'additionalProperties': False}, {'type': 'object', 'required': ['kind', 'reason'], 'properties': {'kind': {'type': 'string', 'const': 'unavailable'}, 'reason': {'type': 'string', 'maxLength': 256, 'minLength': 1}}, 'additionalProperties': False}]}, 'goodputRequestTotal': {'type': 'number', 'minimum': 0}, 'coveredDurationSeconds': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}, 'provenance': {'type': 'object', 'required': ['kind', 'source'], 'properties': {'kind': {'type': 'string', 'const': 'recorded'}, 'source': {'type': 'string', 'maxLength': 256, 'minLength': 1}}, 'additionalProperties': False}}, 'additionalProperties': False}], 'description': 'Interval goodput aggregate when every point has persisted simulation-clock bounds; otherwise status=unavailable. Never derive this from retrieval time or currentStep.'}, 'errorBreakdown': {'type': 'object', 'required': ['poolSaturation', 'dbFailure', 'computeFailure', 'capacityOverload', 'cpuOverload', 'ociStorage', 'queueAbsorption'], 'properties': {'dbFailure': {'type': 'number'}, 'ociStorage': {'type': 'number'}, 'cpuOverload': {'type': 'number'}, 'runtimeMemory': {'type': 'number'}, 'computeFailure': {'type': 'number'}, 'poolSaturation': {'type': 'number'}, 'queueAbsorption': {'type': 'number'}, 'capacityOverload': {'type': 'number'}, 'dependencyFailure': {'type': 'number'}}, 'description': 'Validated additive error contributors in percentage-point units; separates pool/DB, compute, capacity, CPU, storage, runtime-memory, and queue absorption effects', 'additionalProperties': False}, 'gpuUtilization': {'type': 'number', 'description': 'GPU utilization (%) â\x80\x94 present only on simulations with a GPU inference kubernetes resource'}, 'modeledShedRps': {'type': 'number', 'description': 'Aggregate modeled requests per second shed by bounded capacity'}, 'replayIdentity': {'type': 'object', 'required': ['scenarioHash', 'effectiveConfigHash'], 'properties': {'scenarioHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Canonical SHA-256 of the persisted scenario graph and attached traffic-pattern order'}, 'effectiveConfigHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Original replay-only SHA-256 of the effective six-control startup configuration; top-level effectiveConfigHash is the versioned prediction identity.'}}, 'additionalProperties': False}, 'auroraFailovers': {'type': 'array', 'items': {'type': 'object', 'required': ['failedResourceId', 'standbyResourceId', 'succeeded', 'phase'], 'properties': {'phase': {'enum': ['promoting', 'serving', 'unavailable'], 'type': 'string'}, 'succeeded': {'type': 'boolean'}, 'failedResourceId': {'type': 'string'}, 'standbyResourceId': {'type': 'string'}}, 'additionalProperties': False}, 'description': 'Compact per-writer Aurora failover state; promoting has no serving writer, serving means the explicitly related standby took over.'}, 'idleGpuFraction': {'type': 'number', 'description': 'Share (0-1) of the GPU bill that is idle/standby capacity â\x80\x94 present only on GPU inference simulations; values above 0.5 mean over half the GPU spend is HA overhead'}, 'latencyP99Basis': {'type': 'string', 'description': 'Percentile-specific P99 basis: owned fit, scaled-from-fit (not directly measured), or uncalibrated generic model.'}, 'tokensPerSecond': {'type': 'number', 'description': 'Inference throughput in tokens/second â\x80\x94 present only on GPU inference simulations'}, 'goodputSemantics': {'type': 'string', 'const': 'post_step_point_rate', 'description': 'Goodput is a point rate, not an interval total'}, 'goodputProvenance': {'anyOf': [{'type': 'object', 'required': ['kind'], 'properties': {'kind': {'type': 'string', 'const': 'recorded'}}, 'additionalProperties': False}, {'type': 'object', 'required': ['kind', 'sourceFields'], 'properties': {'kind': {'type': 'string', 'const': 'derived'}, 'sourceFields': {'type': 'array', 'items': {'type': 'string', 'minLength': 1}, 'minItems': 1}}, 'additionalProperties': False}, {'type': 'object', 'required': ['kind', 'reason'], 'properties': {'kind': {'type': 'string', 'const': 'unavailable'}, 'reason': {'type': 'string', 'minLength': 1}}, 'additionalProperties': False}], 'description': 'Provenance for the modeled goodput field'}, 'appWeightDefaulted': {'type': 'boolean'}, 'idleGpuCostPerHour': {'type': 'number', 'description': 'USD/hr of GPU spend funding idle/standby capacity (HA overhead) â\x80\x94 present only on GPU inference simulations'}, 'predictionEvidence': {'anyOf': [{'type': 'object', 'required': ['version', 'evidenceLevel', 'legacyGeneric', 'appWeight', 'appWeightDefaulted', 'loadScope', 'sourceIds', 'formulaIds', 'appCpu', 'appCpuByResource', 'latencyP50', 'latencyP95', 'note'], 'properties': {'note': {'type': 'string'}, 'appCpu': {'anyOf': [{'type': 'object', 'required': ['low', 'central', 'high'], 'properties': {'low': {'type': 'number', 'minimum': 0}, 'high': {'type': 'number', 'minimum': 0}, 'central': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}, {'type': 'null'}]}, 'version': {'type': 'number', 'const': 1}, 'appWeight': {'enum': ['lean', 'typical', 'heavy'], 'type': 'string'}, 'loadScope': {'enum': ['within measured load', 'beyond measured load', 'below measured load', 'not calibrated'], 'type': 'string'}, 'sourceIds': {'type': 'array', 'items': {'type': 'string'}}, 'formulaIds': {'type': 'array', 'items': {'type': 'string'}}, 'latencyP50': {'anyOf': [{'type': 'object', 'required': ['low', 'central', 'high', 'evidenceLevel'], 'properties': {'low': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/low'}, 'high': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/high'}, 'central': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/central'}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}}, 'additionalProperties': False}, {'type': 'null'}]}, 'latencyP95': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'latencyP99': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'evidenceLevel': {'enum': ['measured', 'scaled from measured', 'reference estimate'], 'type': 'string'}, 'legacyGeneric': {'type': 'boolean'}, 'appCpuByResource': {'type': 'array', 'items': {'type': 'object', 'required': ['low', 'central', 'high', 'evidenceLevel', 'resourceId', 'legacyGeneric'], 'properties': {'low': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/low'}, 'high': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/high'}, 'central': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/central'}, 'resourceId': {'type': 'string'}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}, 'legacyGeneric': {'type': 'boolean'}}, 'additionalProperties': False}}, 'appWeightDefaulted': {'type': 'boolean'}, 'latencyAvailability': {'type': 'object', 'required': ['status', 'reason', 'resourceIds'], 'properties': {'reason': {'type': 'string', 'minLength': 1}, 'status': {'type': 'string', 'const': 'unavailable'}, 'resourceIds': {'type': 'array', 'items': {'type': 'string', 'minLength': 1}}}, 'additionalProperties': False}, 'requestServingCapacity': {'type': 'array', 'items': {'type': 'object', 'required': ['resourceId', 'resourceName', 'taskCount', 'perTaskCapacityRps', 'aggregateCapacityRps', 'maxTaskCount', 'maxAggregateCapacityRps', 'capacityEvidence'], 'properties': {'taskCount': {'type': 'integer', 'minimum': 0}, 'resourceId': {'type': 'string'}, 'maxTaskCount': {'type': 'integer', 'exclusiveMinimum': 0}, 'resourceName': {'type': 'string'}, 'capacityEvidence': {'type': 'object', 'required': ['basis', 'origin'], 'properties': {'basis': {'enum': ['assumption', 'documented', 'measured', 'catalog'], 'type': 'string'}, 'origin': {'enum': ['caller-provided', 'fargate-size-heuristic', 'policy-default', 'provider-catalog'], 'type': 'string'}, 'source': {'type': 'string'}}, 'additionalProperties': False}, 'perTaskCapacityRps': {'type': 'number', 'exclusiveMinimum': 0}, 'aggregateCapacityRps': {'type': 'number', 'minimum': 0}, 'maxAggregateCapacityRps': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}}, 'latencyPercentileStatus': {'enum': ['modeled', 'unavailable'], 'type': 'string'}}, 'additionalProperties': False}, {'type': 'object', 'required': ['version', 'evidenceLevel', 'appWeight', 'appWeightDefaulted', 'loadScope', 'sourceIds', 'formulaIds', 'appCpu', 'appCpuByResource', 'latencyP50', 'latencyP95', 'note'], 'properties': {'note': {'type': 'string'}, 'appCpu': {'type': 'null'}, 'version': {'type': 'number', 'const': 1}, 'appWeight': {'type': 'null'}, 'loadScope': {'type': 'null'}, 'sourceIds': {'type': 'array', 'items': {'type': 'string'}}, 'formulaIds': {'type': 'array', 'items': {'type': 'string'}}, 'latencyP50': {'type': 'null'}, 'latencyP95': {'type': 'null'}, 'latencyP99': {'type': 'null'}, 'evidenceLevel': {'type': 'string', 'const': 'unavailable/legacy'}, 'legacyGeneric': {'type': 'null'}, 'appCpuByResource': {'type': 'array', 'items': {'type': 'object', 'required': ['low', 'central', 'high', 'evidenceLevel', 'resourceId'], 'properties': {'low': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/low'}, 'high': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/high'}, 'central': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/central'}, 'resourceId': {'type': 'string'}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}}, 'additionalProperties': False}, 'maxItems': 0}, 'appWeightDefaulted': {'type': 'null'}}, 'additionalProperties': False}, {'type': 'object', 'required': ['version', 'evidenceLevel', 'appWeight', 'appWeightDefaulted', 'loadScope', 'sourceIds', 'formulaIds', 'appCpu', 'appCpuByResource', 'latencyP50', 'latencyP95', 'note'], 'properties': {'note': {'type': 'string'}, 'appCpu': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0'}, {'type': 'null'}]}, 'version': {'type': 'number', 'const': 1}, 'appWeight': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appWeight'}, 'loadScope': {'enum': ['within measured load', 'beyond measured load', 'below measured load', 'not calibrated'], 'type': 'string'}, 'sourceIds': {'type': 'array', 'items': {'type': 'string'}}, 'formulaIds': {'type': 'array', 'items': {'type': 'string'}}, 'latencyP50': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'latencyP95': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'latencyP99': {'anyOf': [{'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyP50/anyOf/0'}, {'type': 'null'}]}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}, 'legacyGeneric': {'type': 'boolean'}, 'appCpuByResource': {'type': 'array', 'items': {'type': 'object', 'required': ['low', 'central', 'high', 'evidenceLevel', 'resourceId'], 'properties': {'low': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/low'}, 'high': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/high'}, 'central': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/appCpu/anyOf/0/properties/central'}, 'resourceId': {'type': 'string'}, 'evidenceLevel': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/evidenceLevel'}, 'legacyGeneric': {'type': 'boolean'}}, 'additionalProperties': False}}, 'appWeightDefaulted': {'type': 'boolean'}, 'latencyAvailability': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyAvailability'}, 'latencyPercentileStatus': {'enum': ['modeled', 'unavailable'], 'type': 'string'}}, 'additionalProperties': False}]}, 'calibrationEvidence': {'type': 'object', 'required': ['kind', 'latencyBasis', 'note'], 'properties': {'kind': {'enum': ['owned', 'modeled'], 'type': 'string', 'description': 'owned for the exact AWS CRUD fit; modeled when the gate fails or another generic model applies.'}, 'note': {'type': 'string', 'description': "Names the fit scope and limitations. For canonical workload inference, states it is a modeling assumption. For modeled fallback, names all actual failed checks in stable order with resource IDs and safe expected/actual values; never only a generic 'not applied'."}, 'latencyBasis': {'type': 'string', 'description': 'General modeled latency path, such as in-VPC ALB rather than end-to-end; see latencyP99Basis for percentile-specific P99 provenance.'}, 'calibrationId': {'type': 'string', 'description': 'Versioned owned calibration identifier; present only when that calibration applies.'}, 'latencyP99Basis': {'type': 'string', 'description': 'Percentile-specific P99 basis: owned in-VPC internal-ALB fit, scaled from that fit (not directly measured), or uncalibrated generic model.'}}, 'description': 'Owned-versus-modeled evidence and latency boundary.', 'additionalProperties': False}, 'effectiveConfigHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Versioned prediction hash over replay startup inputs, engine version, and calibration identity; see replayIdentity.effectiveConfigHash for the original replay-only hash.'}, 'latencyAvailability': {'$ref': '#/properties/predictionEvidence/anyOf/0/properties/latencyAvailability'}, 'costPerMillionTokens': {'type': ['number', 'null'], 'description': 'Self-hosted inference cost in USD per million tokens (null when no tokens are being processed) â\x80\x94 present only on GPU inference simulations'}, 'eksSpotInterruptions': {'type': 'array', 'items': {'type': 'object', 'properties': {'name': {'type': 'string'}, 'resourceId': {'type': 'string'}, 'checkpointEvidence': {'type': 'object', 'required': ['engineInputStepIndex', 'simulationSeconds', 'tickDurationSeconds', 'pointRateSemantics', 'intervalAttribution', 'fieldProvenance', 'trafficFieldProvenance'], 'properties': {'fieldProvenance': {'type': 'object', 'required': ['engineInputStepIndex', 'simulationSeconds', 'tickDurationSeconds', 'pointRateSemantics', 'intervalAttribution'], 'properties': {'simulationSeconds': {'$ref': '#/properties/goodputProvenance'}, 'pointRateSemantics': {'$ref': '#/properties/goodputProvenance'}, 'intervalAttribution': {'$ref': '#/properties/goodputProvenance'}, 'tickDurationSeconds': {'$ref': '#/properties/goodputProvenance'}, 'engineInputStepIndex': {'$ref': '#/properties/goodputProvenance'}}, 'additionalProperties': False}, 'simulationSeconds': {'type': 'number', 'minimum': 0}, 'pointRateSemantics': {'type': 'string', 'const': 'post_step_point_rate'}, 'intervalAttribution': {'anyOf': [{'anyOf': [{'type': 'object', 'required': ['status', 'window', 'provenance'], 'properties': {'status': {'type': 'string', 'const': 'recorded'}, 'window': {'$ref': '#/properties/goodputWindow/anyOf/1/properties/window'}, 'provenance': {'type': 'object', 'required': ['kind', 'source'], 'properties': {'kind': {'type': 'string', 'const': 'recorded'}, 'source': {'type': 'string', 'maxLength': 256, 'minLength': 1}}, 'additionalProperties': False}}, 'additionalProperties': False}, {'type': 'object', 'required': ['status', 'provenance'], 'properties': {'status': {'type': 'string', 'const': 'unavailable'}, 'provenance': {'type': 'object', 'required': ['kind', 'reason'], 'properties': {'kind': {'type': 'string', 'const': 'unavailable'}, 'reason': {'type': 'string', 'maxLength': 256, 'minLength': 1}}, 'additionalProperties': False}}, 'additionalProperties': False}]}, {'type': 'null'}]}, 'tickDurationSeconds': {'type': 'number', 'exclusiveMinimum': 0}, 'engineInputStepIndex': {'type': 'integer', 'minimum': 0}, 'trafficFieldProvenance': {'type': 'object', 'required': ['throughputRps', 'offeredRps', 'errorRatePercent', 'latencyP50Ms'], 'properties': {'offeredRps': {'$ref': '#/properties/goodputProvenance'}, 'latencyP50Ms': {'$ref': '#/properties/goodputProvenance'}, 'throughputRps': {'$ref': '#/properties/goodputProvenance'}, 'errorRatePercent': {'$ref': '#/properties/goodputProvenance'}}, 'additionalProperties': False}}, 'description': 'Persisted seeded-EKS checkpoint captured with its interruption JSONB record. engineInputStepIndex is an engine input-step index, simulationSeconds is derived from that index and the recorded tick, and traffic provenance points to outer metric fields rather than duplicating float values.', 'additionalProperties': False}, 'migrationEvaluation': {'type': 'object', 'required': ['status', 'interruptionNoticeAtSimulationSeconds', 'deadlineAtSimulationSeconds', 'migrationStartedAtSimulationSeconds', 'allAffectedWorkloadsReadyAtSimulationSeconds', 'migrationDurationSeconds', 'deadlineSeconds', 'deadlineMet', 'affectedWorkloadCount', 'readyWorkloadCountAtDeadline', 'missedDeadlineWorkloadCount', 'evaluatedAtSimulationSeconds', 'limitingReasons', 'fieldProvenance'], 'properties': {'status': {'enum': ['not_started', 'in_progress', 'completed'], 'type': 'string'}, 'deadlineMet': {'type': ['boolean', 'null']}, 'deadlineSeconds': {'type': 'number', 'const': 120}, 'fieldProvenance': {'type': 'object', 'required': ['status', 'interruptionNoticeAtSimulationSeconds', 'deadlineAtSimulationSeconds', 'migrationStartedAtSimulationSeconds', 'allAffectedWorkloadsReadyAtSimulationSeconds', 'migrationDurationSeconds', 'deadlineSeconds', 'deadlineMet', 'affectedWorkloadCount', 'readyWorkloadCountAtDeadline', 'missedDeadlineWorkloadCount', 'evaluatedAtSimulationSeconds', 'limitingReasons'], 'properties': {'status': {'$ref': '#/properties/goodputProvenance'}, 'deadlineMet': {'$ref': '#/properties/goodputProvenance'}, 'deadlineSeconds': {'$ref': '#/properties/goodputProvenance'}, 'limitingReasons': {'$ref': '#/properties/goodputProvenance'}, 'affectedWorkloadCount': {'$ref': '#/properties/goodputProvenance'}, 'migrationDurationSeconds': {'$ref': '#/properties/goodputProvenance'}, 'deadlineAtSimulationSeconds': {'$ref': '#/properties/goodputProvenance'}, 'missedDeadlineWorkloadCount': {'$ref': '#/properties/goodputProvenance'}, 'evaluatedAtSimulationSeconds': {'$ref': '#/properties/goodputProvenance'}, 'readyWorkloadCountAtDeadline': {'$ref': '#/properties/goodputProvenance'}, 'migrationStartedAtSimulationSeconds': {'$ref': '#/properties/goodputProvenance'}, 'interruptionNoticeAtSimulationSeconds': {'$ref': '#/properties/goodputProvenance'}, 'allAffectedWorkloadsReadyAtSimulationSeconds': {'$ref': '#/properties/goodputProvenance'}}, 'additionalProperties': False}, 'limitingReasons': {'type': 'array', 'items': {'type': 'object', 'required': ['code', 'workloadCount', 'pendingWorkloadCount', 'pullingWorkloadCount', 'startingWorkloadCount', 'residualPullSeconds', 'residualStartupSeconds', 'schedulingCapacity', 'interruptionHandling', 'provenance'], 'properties': {'code': {'enum': ['no_reschedule', 'scheduling_capacity_exhausted', 'image_pull_incomplete', 'startup_incomplete', 'unclassified_runtime_work_remaining'], 'type': 'string'}, 'provenance': {'type': 'string', 'const': 'recorded'}, 'workloadCount': {'type': 'integer', 'minimum': 0}, 'schedulingCapacity': {'type': 'integer', 'minimum': 0}, 'residualPullSeconds': {'type': 'number', 'minimum': 0}, 'interruptionHandling': {'enum': ['reschedule', 'drain-only', 'fail-fast'], 'type': 'string'}, 'pendingWorkloadCount': {'type': 'integer', 'minimum': 0}, 'pullingWorkloadCount': {'type': 'integer', 'minimum': 0}, 'startingWorkloadCount': {'type': 'integer', 'minimum': 0}, 'residualStartupSeconds': {'type': 'number', 'minimum': 0}}, 'additionalProperties': False}, 'maxItems': 5}, 'affectedWorkloadCount': {'anyOf': [{'type': 'integer', 'minimum': 0}, {'type': 'null'}]}, 'migrationDurationSeconds': {'anyOf': [{'type': 'number', 'minimum': 0}, {'type': 'null'}]}, 'deadlineAtSimulationSeconds': {'anyOf': [{'type': 'number', 'minimum': 0}, {'type': 'null'}]}, 'missedDeadlineWorkloadCount': {'anyOf': [{'type': 'integer', 'minimum': 0}, {'type': 'null'}]}, 'evaluatedAtSimulationSeconds': {'anyOf': [{'type': 'number', 'minimum': 0}, {'type': 'null'}]}, 'readyWorkloadCountAtDeadline': {'anyOf': [{'type': 'integer', 'minimum': 0}, {'type': 'null'}]}, 'migrationStartedAtSimulationSeconds': {'anyOf': [{'type': 'number', 'minimum': 0}, {'type': 'null'}]}, 'interruptionNoticeAtSimulationSeconds': {'anyOf': [{'type': 'number', 'minimum': 0}, {'type': 'null'}]}, 'allAffectedWorkloadsReadyAtSimulationSeconds': {'anyOf': [{'type': 'number', 'minimum': 0}, {'type': 'null'}]}}, 'description': 'Authoritative additive seeded EKS interruption evaluation. Field values carry recorded/derived/unavailable provenance; deadline verdict/counts/reasons are frozen once status is completed and do not describe later service health.', 'additionalProperties': False}}, 'additionalProperties': True}, 'description': 'Seeded EKS interruption telemetry, including the authoritative additive migrationEvaluation when recorded.'}, 'resilienceDiagnostics': {'type': 'object', 'properties': {'bounded': {'type': 'boolean', 'description': 'True when a generated-traffic, traversal-work, or cascade-depth bound truncated model work'}, 'pathCount': {'type': 'number', 'description': 'Number of dependency paths evaluated this step'}, 'incidentOutcome': {'type': 'string', 'description': 'Incident outcome: stable | degraded | cascading | protected | recovered'}}, 'description': 'Bounded resilience diagnostics summary (compact mode). Absent when the resilience model did not run. Use simulation.compare_resilience for full per-path detail.', 'additionalProperties': True}, 'retryAmplificationFactor': {'type': ['number', 'null'], 'description': 'Retry amplification factor for this step (attemptedRps / originalRps). Values > 1.0 = amplification risk. null = model ran but no traffic. Absent = resilience model disabled.'}, 'predictionEffectiveConfigHash': {'type': 'string', 'maxLength': 64, 'minLength': 64, 'description': 'Versioned prediction hash over replay startup inputs, engine version, and calibration identity.'}}, 'description': "Compact step summary by default (principal metrics + per-resource status), the complete backend step response with responseMode 'full', or a structured limit response (status: limit_reached) when the step cap is reached", 'additionalProperties': True}
Changed
simulation.metrics
Oct. 2, 2026, 2:40 a.m.
Changed
simulation.step
Oct. 2, 2026, 2:40 a.m.
Changed
simulation.create
Oct. 2, 2026, 2:40 a.m.
Added
simulation.inject_traffic
Sept. 30, 2026, 2:40 a.m.
Added
simulation.metrics
Sept. 30, 2026, 2:40 a.m.
Added
simulation.step
Sept. 30, 2026, 2:40 a.m.
Added
simulation.create
Sept. 30, 2026, 2:40 a.m.
Added
scenario.get
Sept. 30, 2026, 2:40 a.m.
Added
scenario.list
Sept. 30, 2026, 2:40 a.m.
Added
simulation.delete
Sept. 30, 2026, 2:40 a.m.
Added
simulation.recover_resource
Sept. 30, 2026, 2:40 a.m.
Added
simulation.inject_failure
Sept. 30, 2026, 2:40 a.m.