LastPing
Ce que fait ce MCP
Monitors scheduled jobs, CI/CD pipelines, HTTP endpoints, and agent runs, with incident tracking, alert routing, status pages, and Terraform export.
Outils
Schéma d’entrée
{'type': 'object', 'required': ['incident_id', 'body'], 'properties': {'body': {'type': 'string', 'description': 'The diagnosis, in plain words and one or two sentences: what actually failed, whether it is the same failure as before (compare failure_signature.occurrences from list_open_incidents), and what you did about it. Must not be empty or whitespace-only, and must be at most 8192 bytes. An oversized body is REJECTED, never truncated — a truncated diagnosis reads as a complete one that trails off, and the reader cannot tell that the sentence naming the cause was the one cut — so shorten it and call again. A pasted stack trace is a note nobody reads: the full failure output already lives on the run that produced it.'}, 'incident_id': {'type': 'number', 'description': "The incident's numeric id, taken straight from an entry's incident_id in list_open_incidents. An integer, not a UUID."}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Discovered source UUID (from list_discovered_agents).'}, 'agent_id': {'type': 'string', 'description': 'Optional agent UUID (from list_agents) to merge the source into. Omit to create a new agent.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['name'], 'properties': {'name': {'type': 'string', 'description': 'Label for the key, e.g. "github-actions".'}, 'scope': {'enum': ['read', 'write', 'admin', 'ingest'], 'type': 'string', 'description': 'What the new key may do. "read" is every GET; "write" is everything except managing API keys; "admin" is everything, key management included. Omit for "write", which is the right tier for a credential handed to a job or an agent: it can do the work and cannot mint itself a replacement. A key can never be given a HIGHER scope than the key that creates it; asking for one is refused and the refusal names the ceiling. "ingest" can only send pings and telemetry (traces, metrics, logs) and cannot call the REST API at all: use it for a key that lives in a dotfile or an exporter\'s config. This tool needs an admin key; with a write key, use create_ingest_key, which mints a tracing key bound to one monitor.'}, 'check_id': {'type': 'string', 'description': 'Only with scope "ingest": binds the key to this one monitor (UUID from list_monitors), so it can send telemetry for that monitor and nothing else. Required for an exporter that cannot name its monitor, such as Codex.'}, 'expires_at': {'type': 'string', 'description': 'Optional RFC 3339 expiry, e.g. "2026-12-31T00:00:00Z". Omit for a 90-day key, capped at the creating key\'s own expiry. A key can never be given a longer life than the key that creates it.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['kind', 'name'], 'properties': {'url': {'type': 'string', 'description': 'webhook: the POST target URL.'}, 'kind': {'type': 'string', 'description': 'One of: webhook, telegram, discord, slack, ntfy, pushover, msteams, googlechat, email. Every destination URL must be https. A BRANDED kind must point at its vendor\'s host: discord at discord.com or discordapp.com, slack at hooks.slack.com, msteams at webhook.office.com or outlook.office.com or logic.azure.com or logic.azure.us or environment.api.powerplatform.com, googlechat at chat.googleapis.com. For any other endpoint use kind "webhook", which accepts any https host; ntfy is unpinned too, so a self-hosted ntfy server is fine. A pin narrows the destination to the vendor\'s own platform; it does NOT prove the endpoint belongs to the person or project that created it, because every pinned domain is multi-tenant and open to anyone who signs up. Do not report a pinned destination as verified or as owned by anyone on the strength of its host.'}, 'name': {'type': 'string', 'description': "Human-readable destination name, e.g. 'On-call Slack'."}, 'token': {'type': 'string', 'description': 'pushover: the application API token.'}, 'secret': {'type': 'string', 'description': 'webhook: shared secret used to sign the HMAC-SHA256 payload.'}, 'address': {'type': 'string', 'description': 'email: the destination email address (a confirmation link is sent).'}, 'chat_id': {'type': 'string', 'description': 'telegram: the target chat id.'}, 'user_key': {'type': 'string', 'description': 'pushover: the user or group key.'}, 'bot_token': {'type': 'string', 'description': 'telegram: the bot token from @BotFather.'}, 'topic_url': {'type': 'string', 'description': "ntfy: the full topic URL, e.g. 'https://ntfy.sh/my-topic'."}, 'webhook_url': {'type': 'string', 'description': 'slack / discord / msteams / googlechat: the incoming-webhook URL.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['monitor_id'], 'properties': {'name': {'type': 'string', 'description': 'Optional label for the key. Defaults to "Tracing:" followed by the monitor\'s name.'}, 'expires_at': {'type': 'string', 'description': 'Optional RFC 3339 expiry, e.g. "2026-12-31T00:00:00Z". Omit for a 90-day key, capped at the creating key\'s own expiry.'}, 'monitor_id': {'type': 'string', 'description': 'Monitor UUID the key is bound to (from create_monitor or list_monitors).'}}}
Schéma d’entrée
{'type': 'object', 'required': ['name'], 'properties': {'tz': {'type': 'string', 'description': 'IANA timezone for cron evaluation. Defaults to UTC.'}, 'name': {'type': 'string', 'description': "Human-readable monitor name, e.g. 'Daily backup job'."}, 'slug': {'type': 'string', 'description': 'Optional stable ID. If a monitor with this slug exists, it will be updated (upsert). Trimmed and lowercased automatically. Must match ^[a-z0-9][a-z0-9-]{1,48}[a-z0-9]$ (3-50 chars, lowercase alphanumeric and hyphens, starting and ending alphanumeric) after normalisation. UUID-shaped slugs are rejected — they would be ambiguous with a monitor id when importing into Terraform. Omit entirely for no slug.'}, 'tags': {'type': 'string', 'description': "Comma-separated labels for namespace scoping, e.g. 'agent:claude,env:prod'. Max 20 tags, each max 50 chars."}, 'grace_s': {'type': 'number', 'description': 'Grace period in seconds after a ping is due before alerting. Omit on an on_demand monitor and LastPing uses 300 seconds; on_demand has no cadence, so grace only sets the first-run deadline and the overrun fallback. On an upsert (existing slug), omitting it on an on_demand monitor sets 300: pass the current value to keep it.'}, 'agent_id': {'type': 'string', 'description': "Attach this monitor to an agent from the registry, by the agent's id OR its slug (both are returned by register_agent). Omit for a monitor with no owning agent. Naming an agent that does not exist is an error — 400 UNKNOWN_AGENT — it is NEVER created implicitly; call register_agent first to get a valid agent_id. On an upsert (existing slug), omitting this leaves the monitor's current attachment (or lack of one) unchanged; supplying it re-applies the attachment, so an agent re-running its own registration converges to 'attached' every time rather than silently no-opping after the first call."}, 'period_s': {'type': 'number', 'description': "Ping interval in seconds. Required when schedule_kind='simple'."}, 'ci_branch': {'type': 'string', 'description': "CI filter: only count runs on this branch, e.g. 'main'. REQUIRES ci_provider, and the API enforces it: without a CI binding the request is refused with 400 FIELD_NOT_IN_SHAPE rather than accepted and discarded. Subject to the SAME upsert exception as ci_workflow — create_monitor on an existing slug never writes this filter; use update_monitor. WITHOUT IT a run on ANY branch — a feature branch, a fork's pull request — reports to this monitor, so somebody else's broken branch marks your monitor down. Set it to the branch whose health you actually care about, which is almost always the default branch."}, 'cron_expr': {'type': 'string', 'description': "5-field cron expression, e.g. '0 3 * * *'. Required when schedule_kind='cron'."}, 'probe_url': {'type': 'string', 'description': "http monitors only: the absolute http/https URL to probe. Required when monitor_type='http'. The host is resolved at write time and rejected if it resolves only to private/link-local addresses."}, 'ci_provider': {'type': 'string', 'description': "Bind this monitor to a CI system, so the CI system itself reports every run by webhook and the job needs NO ping code at all. One of: 'github', 'gitlab', 'jenkins'. SET-ONCE: ci_provider can only be chosen when the monitor is created — update_monitor cannot change or remove it, so a monitor bound to the wrong provider must be deleted and recreated. Setting it generates a webhook secret that is returned exactly ONCE, in THIS call's response, together with the webhook URL. It is never retrievable afterwards — no MCP tool and no API read returns it again — so copy both out of the response and configure the CI webhook before doing anything else. Omit for a monitor that pings for itself. Also set ci_workflow and ci_branch unless the repository really has exactly one workflow on one branch. NOT ACCEPTED on monitor_type='http': an http probe is never bound to CI, and the API returns 400 FIELD_NOT_IN_SHAPE. It used to accept the provider, create no binding, and report success."}, 'ci_workflow': {'type': 'string', 'description': "CI filter: only count runs of the workflow / pipeline / job with this exact name. REQUIRES ci_provider, and the API enforces it: without a CI binding this filter has nowhere to be stored, so the request is refused with 400 FIELD_NOT_IN_SHAPE rather than accepted and discarded. Note that monitor_type='ci' does NOT bind anything on its own — ci_provider does. ONE EXCEPTION, and it is on the path agents use most, so do not rely on the enforcement here: create_monitor on a slug that ALREADY EXISTS is an upsert, and the upsert never writes this filter. With ci_provider in the same call the request is accepted and the filter is silently discarded; without it the request is refused, and doing what the error advises — adding ci_provider — reaches the discarding case instead. Set this filter with update_monitor, which does persist it. WITHOUT IT, EVERY workflow in the repository reports to this monitor — so one unrelated failing workflow opens an incident against a job that is perfectly healthy, and a green run of a different workflow clears an incident the real job never recovered from. Set it whenever the repository has more than one workflow."}, 'monitor_from': {'type': 'string', 'description': "DORMANT UNTIL: an RFC 3339 timestamp before which no deadline is computed and no incident can open — the monitor is fully configured but not yet armed. Use it when you provision ahead of the work: a monitor for a job that does not start running until next Monday is otherwise 'late' from the moment you create it, which is a false alert on day one. The first-run deadline is seeded as monitor_from + grace_s. Default: unset, meaning deadlines start immediately. Example: '2026-01-01T00:00:00Z'. On an upsert (existing slug), omitting this clears the monitor's monitor_from and arms it immediately — pass the current value to keep it."}, 'monitor_type': {'type': 'string', 'description': "'heartbeat' (default), 'ci', or 'http'. Any other value is refused with 400 UNKNOWN_MONITOR_TYPE. 'ci' is a label, not a binding: a CI monitor is a heartbeat monitor with ci_provider set, so passing monitor_type='ci' WITHOUT ci_provider creates an ordinary heartbeat and its ci_workflow/ci_branch filters are refused."}, 'probe_method': {'type': 'string', 'description': "http monitors only: the HTTP method the probe sends. One of 'GET', 'HEAD', 'POST'. Default 'GET'. Use 'HEAD' for a cheap liveness check when the body does not matter — but note it returns no body, so probe_expected_body cannot match anything."}, 'max_runtime_s': {'type': 'number', 'description': "Maximum seconds a single run may take before it is reported overdue (the 'overrun' rule), measured from the run's start ping. Omit to fall back to grace_s. This is how a long job avoids being flagged overdue while still being detected quickly if it goes silent: e.g. grace_s=600 with max_runtime_s=14400 alerts 10 minutes after a missed ping but tolerates a 4-hour run. It replaces grace_s for the overrun deadline ONLY — the silence rule and the first-run deadline still use grace_s. Range 60-31536000. Not supported on http monitors: a probe has no start/success pair, so the overrun rule can never fire and the API returns 400 MAX_RUNTIME_NOT_SUPPORTED (use probe_timeout_s to bound a single probe). On an upsert (existing slug), omitting this clears the monitor's max_runtime_s — pass the current value to keep it."}, 'schedule_kind': {'type': 'string', 'description': "'simple' (requires period_s), 'cron' (requires cron_expr), or 'on_demand' (requires neither). Required for heartbeat/ci monitors. NOT ACCEPTED on monitor_type='http', together with period_s, cron_expr and tz: an http monitor's schedule is derived from probe_interval_s, so the API refuses all four with 400 FIELD_NOT_IN_SHAPE instead of accepting and ignoring them. 'on_demand' means no cadence at all: no period_s, no cron_expr — the API returns 400 if either is supplied — and, by default, NO ABSENCE DEADLINES ARE ARMED BETWEEN RUNS. What this trades away: nothing tells you if the agent is never invoked again; silence between runs is invisible unless you opt in to expect_every_s. What it buys: a healthy agent that nobody happens to invoke for a week never generates a false 'late' or 'down' for simply not having been asked to run. Only run-scoped detection still applies once a run starts — max_runtime_s (overrun), step_timeout_s (stall), blocked_timeout_s (stuck on a human) — because those are anchored to a run's own start ping, not to a cadence. IMPORTANT: if you would be alarmed to find this agent silent for hours, set expect_every_s as well — it is the silence floor, and it is the only thing that makes an on_demand monitor detect absence at all. Choose 'simple'/'cron' when the agent is supposed to run on a cadence; choose 'on_demand' when invocation is inherently irregular and a quiet stretch between runs is expected, not a symptom."}, 'trace_content': {'enum': ['dropped', 'redacted'], 'type': 'string', 'description': "What this monitor's traces keep of prompt, command and tool content. 'dropped' (the default) removes it; 'redacted' keeps it, with every secret-shaped value redacted when it arrives. Only a person should choose 'redacted': never set it on your own initiative, only when the person you work for has asked for content to be stored. Omit on a create for dropped; on an upsert (existing slug), omitting it leaves the stored value unchanged."}, 'expect_every_s': {'type': 'number', 'description': "SILENCE FLOOR in seconds: open a 'silence' incident if NO ping of any kind — success, start, fail, step — has arrived within this window, regardless of the schedule. It is anchored on the monitor's last activity, not on a cadence, which is what makes it the ONLY absence rule an 'on_demand' monitor can have: that schedule_kind arms nothing between runs, so without this field an on_demand monitor reads 'up' forever no matter how long the agent stays dark. Set it on any on_demand agent monitor you would be alarmed to find silent — that is what it is for. It does NOT fire mid-run: while a run is in flight (a start ping is outstanding) the floor stands down entirely and the run clock owns detection (max_runtime_s, step_timeout_s), so a legitimate 4-hour run that reports nothing is still not an incident. A 'blocked' ping also pauses it, bounded by blocked_timeout_s. On 'simple'/'cron' monitors it is a backstop rather than the main rule: it joins the existing deadline as whichever is SOONER, so it can tighten detection under a long cadence (a daily cron has a ~25-hour blind window) but can never loosen it. Default: unset, which means no floor and is exactly how every monitor behaved before this field existed. Range 60-31536000. Accepted on every monitor_type and every schedule_kind. On an upsert (existing slug), omitting this clears the monitor's expect_every_s and turns the silence floor back off — pass the current value to keep it."}, 'step_timeout_s': {'type': 'number', 'description': "Progress budget in seconds: how long an armed run may go without reporting a step before a 'stalled' incident opens (the stall rule). The clock is anchored on the LATER of the run's start ping and its most recent step, so a run that wedges before its first step is caught too. Reach for this when 'still running' and 'still making progress' are different things — a long agent loop, a multi-stage pipeline, a migration. max_runtime_s alone tells you nothing until the whole budget expires; step_timeout_s=300 on a 4-hour budget tells you within five minutes, and names the last step that reported. To use it the run must report steps: call get_ping_instructions and use curl_step (POST <ping_url>/step?rid=<run-id>&step=<name>). A monitor with step_timeout_s set whose job never reports a step will open a stalled incident on EVERY run — set the field and instrument the job in the same change. Stall detection needs your job to call /step: the Claude Code hook and the reporting prompt do not send steps, so leave step_timeout_s unset for them. Default: unset, which disables stall detection entirely; a monitor that sets nothing behaves exactly as it did before this field existed. Range 10-86400. Two constraints. (1) It must be strictly LESS than the effective run budget, COALESCE(max_runtime_s, grace_s), or the API returns 400 STEP_TIMEOUT_EXCEEDS_BUDGET — at or above the budget the run overruns first, so the stall rule could never fire. (2) Not supported on http monitors: a probe never arms a run and has no /step endpoint to call, so the API returns 400 STEP_TIMEOUT_NOT_SUPPORTED. A step resets the stall clock ONLY — it never extends max_runtime_s, so an agent that reports progress forever still overruns. On an upsert (existing slug), omitting this clears the monitor's step_timeout_s and turns stall detection back off — pass the current value to keep it."}, 'probe_timeout_s': {'type': 'number', 'description': "http monitors only: how many seconds a single probe may take before it counts as a failure. Range 1-30, default 10. This is the http equivalent of max_runtime_s, which http monitors reject: it is the only way to say 'answering, but far too slowly to be healthy'."}, 'runaway_ceiling': {'type': 'number', 'description': "PING-RATE CEILING: the maximum number of pings this monitor may receive in a rolling one-hour window. Exceeding it opens a 'runaway' incident. This is the rule that catches a job or agent stuck in a LOOP — the failure every other rule misses, because a looping agent is pinging enthusiastically and therefore reads 'up' the whole time it is burning tokens or money. Set it a little above the monitor's real cadence: a job that runs every 15 minutes sends about 4 pings/hour, so 20 absorbs retries and still catches a loop. It is RATE-based, so failure_threshold does not gate it and neither does any run budget. Default: unset, which disables the runaway rule entirely. On an upsert (existing slug), omitting this clears the monitor's ceiling and turns the runaway rule back off — pass the current value to keep it."}, 'notify_min_run_s': {'type': 'number', 'description': "NOTIFICATION DURATION FLOOR in seconds: a run SHORTER than this does not produce an INFO-CLASS notification (success, started, every-run, note). This exists for exactly one problem: on an agent monitor, one run is one task you asked for, so asking the agent 'what's 2+2' produces a start and a success notification exactly like a 56-minute deploy does. If you have routed success/started/every-run/note to a destination, you WILL be paged for trivial runs unless you set this. IT NEVER SUPPRESSES A FAILURE. down, fail, recovery and blocked are alert-class and are never affected by this field, however short the run — a run that failed in two seconds is exactly what you need to hear about, and this field cannot silence that, structurally, no matter how it is set. It also never suppresses 'started': a run's duration does not exist yet the moment it begins, so started is always reported regardless of this floor. And it never suppresses an event whose duration could not be measured at all (e.g. a bare success with no preceding start ping) — an unknown duration always means 'notify', never 'suppress'. Default: unset, which means no floor and is exactly how every monitor behaved before this field existed. Range 60-31536000. Not supported on http monitors: an http probe has no start/success pair, so its run duration is never measured and the floor could never apply (the API returns 400 NOTIFY_MIN_RUN_NOT_SUPPORTED). On an upsert (existing slug), omitting this clears the monitor's notify_min_run_s and turns the notification duration floor back off — pass the current value to keep it."}, 'probe_interval_s': {'type': 'number', 'description': "http monitors only: how often to probe, in seconds. Required when monitor_type='http'. Range 30-86400."}, 'blocked_timeout_s': {'type': 'number', 'description': "Maximum seconds a run may sit in the 'blocked' state (an agent reported it is waiting on a human) before a 'blocked' incident opens. UNSET DOES NOT MEAN WAIT FOREVER: omitting this does not disable the timeout, it falls back to the default, which is 24 HOURS — an agent still blocked 24 hours after reporting so, with this field never set, gets a 'blocked' incident regardless. Lower it to be paged sooner when a stuck approval is urgent; raise it for work that legitimately waits on a human for longer than a day. This is distinct from the non-incident 'blocked' notification a route on the 'blocked' event type delivers (see set_route), which is held 10 minutes and sent only if the run is still blocked then, at most once per blocked stretch of a run; this field governs the separate incident that opens only if the wait outlives the timeout. Accepted on every monitor_type: unlike max_runtime_s/step_timeout_s it has no run-scoped precondition an http monitor could fail, so there is nothing to reject. On an upsert (existing slug), omitting this clears the monitor's blocked_timeout_s and falls back to the 24h default — pass the current value to keep it."}, 'failure_threshold': {'type': 'number', 'description': "Number of consecutive failures required before an incident opens. Default 1 (open on the very first failure). This is how you stop a single transient blip from paging someone: set 2-5 on a job that fails occasionally for reasons that resolve themselves, and no incident opens until that many runs in a row have failed. Any success resets the count to zero. It gates the 'fail' cause ONLY — silence (a missed ping), overrun, never_started and runaway are time- or rate-based, so a consecutive count means nothing for them and they are never delayed by it. A run that ends on the model provider's API error (server_error, overloaded, rate_limit) neither counts toward it nor resets it: those open an 'upstream' incident once 3 runs in a row end that way, or this many when it is higher. Range 1-100. On an upsert (existing slug), omitting this resets the monitor's threshold to 1 — pass the current value to keep it."}, 'probe_expected_body': {'type': 'string', 'description': 'http monitors only: a substring that MUST appear in the response body for the probe to count as healthy. THIS IS THE DIFFERENCE BETWEEN \'the server answered\' AND \'the app works\': a broken app that renders an error page still returns 200, passes a status-only check, and leaves the monitor green. Match on something only a healthy response contains, e.g. \'"status":"ok"\'. Substring match, not a regex, and case-sensitive. Default: empty, meaning the body is not inspected at all.'}, 'probe_expected_status': {'type': 'number', 'description': "http monitors only: the EXACT HTTP status code that counts as healthy. Default 200; any other code fails the probe. Set it when the healthy answer is not 200 — 204 for a no-content health endpoint, or 301 when what you are checking is that a redirect still exists (pair that with probe_follow_redirects=false, or the probe will follow it and see the destination's status instead)."}, 'probe_follow_redirects': {'type': 'boolean', 'description': 'http monitors only: whether the probe follows 3xx redirects. Default false. Leaving it false is usually what you want: the redirect itself is then compared against probe_expected_status like any other response, so a site that starts redirecting to a login wall, a parking page or an outage notice is CAUGHT rather than silently followed to a healthy-looking 200. Set true only when the URL you are checking is legitimately a redirect to the thing you actually care about.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['title'], 'properties': {'slug': {'type': 'string', 'description': 'Optional URL slug, which is what appears in the public link (/status/<slug>). Must match ^[a-z0-9][a-z0-9-]{1,48}[a-z0-9]$ (3-50 chars, lowercase alphanumeric and hyphens, starting and ending alphanumeric). Slugs are GLOBALLY unique across all projects, not just yours, so a desirable one may be taken — that returns 409. OMIT IT unless the user asked for a specific URL: a random unguessable slug is then generated, which is also the safer default for a public page.'}, 'title': {'type': 'string', 'description': "Human-readable page title, e.g. 'Acme API Status'. Shown at the top of the page, and to anyone the page is shared with."}, 'check_ids': {'type': 'string', 'description': 'Comma-separated monitor UUIDs to show on the page, in no particular order. Get them from list_monitors. Every id must belong to this project — an unknown or cross-project id returns 400 and nothing is saved. An empty value is legal and produces a page with no monitors on it.'}, 'visibility': {'type': 'string', 'description': "'private' (default) or 'public'. 'public' means the page is served at a guessable-free but UNAUTHENTICATED URL: anyone with the link sees the title, the name of every monitor on it, and its up/down history. Monitor names are frequently internal ('billing-reconciler', 'acme-corp-nightly-sync'), so treat this as publishing them. Choose 'private' unless the user has actually asked for a page other people can see. The free tier allows exactly ONE public page per project; a second returns 403."}}}
Schéma d’entrée
{'type': 'object', 'required': ['check_id', 'rid', 'assertions'], 'properties': {'rid': {'type': 'string', 'description': "The run id exactly as sent on this run's /start ping — the same rid used on every step and the terminal ping."}, 'check_id': {'type': 'string', 'description': 'Monitor UUID (from create_monitor or list_monitors).'}, 'assertions': {'type': 'string', 'description': 'The run\'s complete set of expectations, declared ONCE at the start of the run -- criteria the ping BODY of THIS run\'s eventual success ping must satisfy when the run closes, checked instead of letting the run grade itself. IMMUTABLE: a second call for the same rid is rejected with a conflict error and the first declaration stands unchanged -- there is no way to edit, add to, or replace it once made, so decide the whole set before you start work. Declaring nothing is allowed and always has been: simply never call this tool for a run, and the monitor\'s own check-level assertions (if any) stay in force unchanged. INCLUDE AT LEAST ONE POSITIVE CRITERION -- a \'contains\', \'matches\' or \'json_path\' entry -- in every declaration. A declaration made ENTIRELY of \'not_contains\' entries is self-satisfying on empty output: a run that produces nothing at all still passes, because there is nothing for the pattern to find. That is precisely the evasion this feature exists to close, so a purely negative declaration defeats its own purpose. A \'matches\' entry only counts as positive if its pattern REJECTS an empty body: \'.*\', \'(?s).*\' and \'^$\' all accept one and are validated as perfectly legal patterns, so a declaration resting on one of those is no better than a purely negative declaration. Supply a JSON ARRAY as a string, e.g. \'[{"kind":"json_path","path":"result.rows_processed","op":"gt","value":"0"}]\'. Fields per entry: kind (required), value, path, op -- no name; a run\'s declared criteria have none, unlike a monitor\'s own output assertions. kind is one of \'contains\' (body contains value as a substring), \'not_contains\' (body does not contain it), \'matches\' (body matches value as a Go RE2 regexp, max 1000 bytes), or \'json_path\' (parse the body as JSON, read the value at path, compare it against value with op). contains/not_contains/matches require value; json_path requires path and op. path is a DOTTED path only (\'a.b.c\') -- the query syntax of a real JSONPath library (\'[\', \'*\', \'$\') is rejected. op is one of \'eq\', \'ne\', \'gt\', \'gte\', \'lt\', \'lte\'. At most 20 assertions per run. A malformed entry (uncompilable regexp, a path carrying query syntax, an unknown kind or op) is rejected before anything is written, and nothing is stored if any entry fails.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Agent UUID (from register_agent or list_agents).'}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Destination (channel) UUID. Get it from list_destinations.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Monitor UUID.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['monitor_id', 'event_type'], 'properties': {'event_type': {'enum': ['down', 'recovery', 'fail', 'every-run', 'success', 'started', 'blocked', 'note'], 'type': 'string', 'description': 'The event type to unroute: down, recovery, fail, every-run, success, started, blocked or note.'}, 'monitor_id': {'type': 'string', 'description': 'Monitor (check) UUID.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Status page UUID, from list_status_pages.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['sources'], 'properties': {'sources': {'type': 'string', 'description': 'The complete scan result: a JSON ARRAY supplied as a string, one entry per scheduled job, e.g. \'[{"source_kind":"crontab","source_ref":"/etc/cron.d/backup:/usr/local/bin/backup.sh","name":"nightly backup","schedule_cron":"0 3 * * *","tz":"Europe/Berlin"}]\'. Fields per entry: source_kind and source_ref (both REQUIRED — an entry missing either cannot be matched against an existing monitor and would be re-created on every scan), name (optional display name; falls back to source_ref), schedule_cron (optional 5-field cron expression, sent only when you actually read one), tz (the IANA zone that cron fires in — REQUIRED for a crontab or systemd-timer entry carrying a schedule_cron, read from the host, never guessed), and suggested_expect_every_s (optional; state the silence floor outright when you know the real cadence better than the cron expression does — it WINS over the value derived from the cron). Send the WHOLE scan in one call: this is a diff, so a source you leave out is reported as orphaned rather than ignored. Send \'[]\' to report that the scan found nothing — every discovered monitor is then listed as orphaned, and none of them is deleted. At most 1000 entries per call, each source_kind/source_ref pair at most once (a duplicate is rejected outright, not merged), and the project\'s 100-monitor cap is applied to the whole batch at once — if the batch would exceed it, NOTHING is created. Nothing is written unless every entry validates: one bad entry rejects the entire payload and leaves no monitors behind.'}}}
Schéma d’entrée
{'type': 'object', 'required': [], 'properties': {'tag': {'type': 'string', 'description': "Optional tag to filter monitors by, e.g. 'agent:claude'. Only monitors carrying this tag (and their routes/templates) are exported."}, 'include': {'type': 'string', 'description': 'Optional comma-separated subset of monitors,destinations,routes,templates,status_pages. Omit to export everything.'}, 'monitor_slug': {'type': 'string', 'description': 'Optional slug to export a single monitor by. Combines with tag if both are given.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Agent UUID (from register_agent or list_agents).'}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Agent UUID or slug (from list_agents).'}, 'range': {'enum': ['24h', '7d', '30d'], 'type': 'string', 'description': 'Time window: 24h, 7d (the default) or 30d. It covers every UTC day that overlaps it, so 24h spans two days.'}, 'direction': {'enum': ['out', 'in', 'all'], 'type': 'string', 'description': 'out (the default): what this agent calls. in: who calls it. all: both.'}}}
Schéma d’entrée
{'type': 'object', 'required': [], 'properties': {'id': {'type': 'string', 'description': 'Agent UUID or slug (from list_agents). Omit for every agent in the project.'}, 'range': {'enum': ['24h', '7d', '30d'], 'type': 'string', 'description': 'Time window: 24h, 7d (the default) or 30d. It covers every UTC day that overlaps it, so 24h spans two days.'}, 'project': {'type': 'string', 'description': "Only usage from traced runs of this project; metrics-only usage has no project and is left out. days then hold one row per UTC day with model and provider empty (a run's totals are not split by model); by_agent is not narrowed."}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Monitor UUID.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['incident_id'], 'properties': {'incident_id': {'type': 'number', 'description': "The incident's numeric id, from list_incidents, list_open_incidents or add_incident_note."}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Monitor UUID.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Monitor UUID (from create_monitor or list_monitors).'}, 'tool': {'enum': ['claude-code', 'codex'], 'type': 'string', 'description': "Which tool's install `hook_install` carries: claude-code (the default) or codex. Everything else in the result is the same."}}}
Schéma d’entrée
{'type': 'object', 'required': ['id', 'rid'], 'properties': {'id': {'type': 'string', 'description': 'Monitor UUID.'}, 'rid': {'type': 'string', 'description': 'Run id as sent on the ping.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Monitor UUID.'}, 'limit': {'type': 'number', 'description': 'Max runs to return (default 20, max 100).'}}}
Schéma d’entrée
{'type': 'object', 'required': ['monitor_id'], 'properties': {'monitor_id': {'type': 'string', 'description': 'Monitor UUID (from create_monitor or list_monitors).'}}}
Schéma d’entrée
{'type': 'object', 'required': ['monitor_id'], 'properties': {'tool': {'enum': ['claude-code', 'codex', 'gemini', 'cursor', 'python', 'node', 'otel-sdk', 'collector'], 'type': 'string', 'description': 'Which tool will send the traces: claude-code, codex, gemini, cursor, python, node, otel-sdk or collector. Omit to get every block.'}, 'monitor_id': {'type': 'string', 'description': 'Monitor UUID (from create_monitor or list_monitors).'}}}
Schéma d’entrée
{'type': 'object', 'required': [], 'properties': {}}
Schéma d’entrée
{'type': 'object', 'required': [], 'properties': {}}
Schéma d’entrée
{'type': 'object', 'required': [], 'properties': {'limit': {'type': 'number', 'description': 'Max deliveries to return (default 20, max 100).'}, 'status': {'type': 'string', 'description': 'Restrict to one delivery status: pending, delivered, dead, or suppressed.'}, 'monitor': {'type': 'string', 'description': "Restrict to one monitor's deliveries (UUID)."}}}
Schéma d’entrée
{'type': 'object', 'required': [], 'properties': {'kind': {'enum': ['model', 'tool', 'http', 'database', 'queue', 'rpc', 'agent'], 'type': 'string', 'description': 'Only one kind of dependency: model, tool, http, database, queue, rpc or agent. Omit for all.'}, 'range': {'enum': ['24h', '7d', '30d'], 'type': 'string', 'description': 'Time window: 24h, 7d (the default) or 30d. It covers every UTC day that overlaps it, so 24h spans two days.'}}}
Schéma d’entrée
{'type': 'object', 'required': [], 'properties': {}}
Schéma d’entrée
{'type': 'object', 'required': [], 'properties': {}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Monitor UUID.'}, 'limit': {'type': 'number', 'description': 'Max incidents to return (default 50, max 200).'}}}
Schéma d’entrée
{'type': 'object', 'required': [], 'properties': {'tag': {'type': 'string', 'description': "Optional tag to filter by, e.g. 'agent:claude'. Returns only monitors that have this tag."}}}
Schéma d’entrée
{'type': 'object', 'required': ['agent_id'], 'properties': {'limit': {'type': 'number', 'description': 'Max incidents to return (default 50, max 200). Newest first, so a small limit drops the oldest open incidents, not the newest.'}, 'agent_id': {'type': 'string', 'description': 'Agent UUID (from register_agent or list_agents). The inbox covers every monitor this agent owns.'}}}
Schéma d’entrée
{'type': 'object', 'required': [], 'properties': {'q': {'type': 'string', 'description': 'Only runs whose run id (rid) contains this text.'}, 'agent': {'type': 'string', 'description': 'Only runs of this agent (agent UUID or slug).'}, 'limit': {'type': 'number', 'description': 'Runs per page (default 20, max 100).'}, 'model': {'type': 'string', 'description': 'Only runs that called this model, by exact model id.'}, 'since': {'type': 'string', 'description': 'RFC 3339 start of the window, e.g. 2026-09-01T00:00:00Z. Default 7 days ago; at most 90 days back.'}, 'until': {'type': 'string', 'description': 'RFC 3339 end of the window. Default now.'}, 'cursor': {'type': 'string', 'description': 'next_cursor from the previous page, verbatim.'}, 'traced': {'type': 'boolean', 'description': 'true: only runs that hold spans. Omit or false for every run.'}, 'monitor': {'type': 'string', 'description': "Only this monitor's runs (monitor UUID)."}, 'outcome': {'enum': ['succeeded', 'failed', 'cancelled', 'blocked', 'running', 'unfinished'], 'type': 'string', 'description': 'Only runs with this outcome.'}, 'project': {'type': 'string', 'description': 'Only runs whose Claude Code hook start carried this project label: the name of the folder the session worked in.'}, 'trace_id': {'type': 'string', 'description': 'Only the run a trace became: 8 to 32 hex digits of its trace id. Finds runs made from a trace with no run id, not runs that named their own.'}, 'has_error': {'type': 'boolean', 'description': 'true: only runs with a span that reported an error. false: only runs without one.'}, 'operation': {'type': 'string', 'description': 'Only runs with a span whose name starts with this text.'}, 'dependency': {'type': 'string', 'description': 'Only runs that called this dependency, by its exact name as get_agent_dependencies reports it (e.g. api.github.com).'}, 'min_cost_usd': {'type': 'string', 'description': 'Only runs whose cost is at least this many US dollars, as a decimal string, e.g. "0.25".'}, 'min_duration_ms': {'type': 'number', 'description': 'Only runs that took at least this many milliseconds.'}}}
Schéma d’entrée
{'type': 'object', 'required': [], 'properties': {}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Monitor UUID.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['api_key_id'], 'properties': {'api_key_id': {'type': 'string', 'description': 'UUID of the key to regenerate. Get it from list_api_keys.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['name'], 'properties': {'name': {'type': 'string', 'description': "Human-readable agent name, e.g. 'Deploy Bot'. Used to derive the agent's slug."}, 'description': {'type': 'string', 'description': 'Optional free-text description of what this agent does. Omit for none.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Monitor UUID.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['api_key_id'], 'properties': {'api_key_id': {'type': 'string', 'description': 'UUID of the key to revoke. Get it from list_api_keys.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['id', 'event_type', 'template'], 'properties': {'id': {'type': 'string', 'description': 'Monitor UUID.'}, 'cause': {'type': 'string', 'description': "Optional cause for a per-cause override (e.g. 'silence', 'overrun', 'never_started', 'stalled', 'runaway', 'upstream'). Omit or leave empty for an event-type-wide template."}, 'template': {'type': 'string', 'description': 'Template text with {variable} placeholders. Empty string resets to the built-in default.'}, 'event_type': {'type': 'string', 'description': "Event type: 'down', 'recovery', 'fail', 'every-run', 'success', 'started', 'blocked', 'note'."}}}
Schéma d’entrée
{'type': 'object', 'required': ['monitor_id', 'event_type'], 'properties': {'event_type': {'type': 'string', 'description': "One of eight: down (alert opened), recovery (alert cleared), fail (explicit failure ping), every-run (one notification per completed run, success or failure), success (fires only when a run completes successfully), started (fires when a run begins), blocked (an agent reported it is waiting on a human; held 10 minutes and sent only if that run is still blocked then, at most once per blocked stretch of a run: a note, step, success, fail or cancel for the same run inside the 10 minutes cancels it, a title does not; a blocked ping sent without a run id belongs to no run, so any non-blocked ping on the monitor ends its stretch: a start, log, note, success, fail or cancel of any run, or a step of the monitor's current run (any step when no run is current); this is separate from the 'blocked' INCIDENT that opens later only if the wait outlives blocked_timeout_s, see create_monitor/update_monitor), note (a free-form annotation ping — never itself opens or clears an incident). Prefer down/recovery/fail: they fire only on a state change. every-run, success, started, and note are not state changes and are bounded only by how often the monitor runs (or how often the agent chooses to send them), so they can be very chatty, and none of them is flap-damped. started is the chattiest of the bunch for CI-fed monitors: GitHub maps both the workflow_run 'requested' and 'in_progress' webhook events to a start signal, so a single CI run can emit more than one started event — this was observed in production, where a real run logged two starts seconds apart. every-run, success, started, and note share one separate per-channel rate cap (60/hour by default), so together they can no longer use up the budget that down/fail/recovery/blocked need — but a chatty route on any one of the four can silently suppress its own notifications, and its sibling informational types' notifications, once it exceeds that shared cap. blocked is deliberately NOT in that shared group even though it is agent-reported rather than system-derived: a blocked agent needs a human, so it draws on the protected down/fail/recovery budget instead, precisely so it cannot be starved by chatty every-run/success/started/note traffic. Route informational types to a low-stakes destination, not to the one that pages someone."}, 'monitor_id': {'type': 'string', 'description': 'Monitor (check) UUID.'}, 'channel_ids': {'type': 'string', 'description': 'Comma-separated destination (channel) UUIDs to notify. Empty string clears the route.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Monitor UUID.'}, 'clear': {'type': 'boolean', 'description': 'Set true to remove the active maintenance window.'}, 'until': {'type': 'string', 'description': 'RFC 3339 end timestamp. Use this OR duration OR clear.'}, 'duration': {'type': 'string', 'description': "Go duration string, e.g. '1h' or '24h'. Use this OR until OR clear."}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Destination (channel) UUID. Get it from list_destinations or create_destination.'}, 'resend_verification': {'type': 'boolean', 'description': 'Set true to re-send the email confirmation link INSTEAD of a test alert. Email destinations only — any other kind returns 400. Safe to repeat, and idempotent: on an already-verified destination it reports verified and sends nothing rather than mailing the user again.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['id', 'name'], 'properties': {'id': {'type': 'string', 'description': 'Agent UUID (from register_agent or list_agents).'}, 'name': {'type': 'string', 'description': "Human-readable agent name, e.g. 'Deploy Bot'."}, 'slug': {'type': 'string', 'description': "New slug for the agent, e.g. 'reddit-bot': 3-50 characters, lowercase letters, digits and hyphens. Omit to keep the current slug, which is the default: pass it only when the person asks to change the slug. Saved links, Terraform references and trace sources that name the old slug stop matching this agent, unless they also equal its name (case-insensitive), past traced runs included; a slug equal to a source another agent receives by name or adoption takes that source's traces."}, 'description': {'type': 'string', 'description': "Free-text description of what this agent does. Omit to leave the agent's current description unchanged — THIS IS THE DEFAULT AND SAFE CHOICE for a name-only rename. Pass an explicit empty string to clear an existing description back to none."}}}
Schéma d’entrée
{'type': 'object', 'required': ['destination_id'], 'properties': {'name': {'type': 'string', 'description': 'New human-readable label. Omit to leave unchanged.'}, 'config': {'type': 'object', 'properties': {}, 'description': 'Replacement config for the destination\'s existing kind — one of webhook, telegram, discord, slack, ntfy, pushover, msteams, googlechat, email. Shape must match the kind: {"url":…,"secret":…} for webhook, {"bot_token":…,"chat_id":…} for telegram, {"webhook_url":…} for slack/discord/msteams/googlechat, {"topic_url":…} for ntfy, {"token":…,"user_key":…} for pushover, {"address":…} for email. Omit to leave unchanged. A URL you supply is re-checked against the kind\'s allowed hosts and must be https; a destination created before that rule keeps working until you send a new config for it. The config must name only the fields listed for its kind, each exactly once. The host rule narrows a branded destination to the vendor\'s own platform; it does NOT prove the endpoint belongs to the person or project that owns the destination.'}, 'destination_id': {'type': 'string', 'description': 'UUID of the destination to update. Get it from list_destinations.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['id', 'name'], 'properties': {'id': {'type': 'string', 'description': 'Monitor UUID.'}, 'tz': {'type': 'string', 'description': 'IANA timezone for cron evaluation.'}, 'name': {'type': 'string', 'description': 'Human-readable monitor name.'}, 'tags': {'type': 'string', 'description': "Comma-separated labels to set on this monitor, e.g. 'agent:claude,env:prod'. Replaces existing tags. Max 20 tags, each max 50 chars."}, 'guards': {'type': 'string', 'description': 'Metric guards: CEILINGS on a number the job reports about itself, checked on every ping. An assertion catches a run that did nothing; a guard catches the opposite — an agent that loops, retries and burns money. Each guard reads one number out of the ping body at a dotted path, rolls it up across a trailing window, and opens an incident with cause \'runaway\' when the total EXCEEDS the ceiling (equal does not trip). Supply a JSON ARRAY as a string, e.g. \'[{"name":"daily spend","path":"cost.usd","window_s":86400,"ceiling":50,"aggregation":"sum"}]\'. REPLACE-THE-SET: the array you send becomes the monitor\'s complete guard set — it is NOT merged with what is already there. Omit the argument entirely to leave the current guards untouched; pass \'[]\' to remove all of them. Fields per entry, all required: name (appears on the incident, and is the only thing that tells a tripped guard apart from the fixed pings-per-hour runaway ceiling), path (DOTTED path into the ping body parsed as JSON — \'cost.usd\'; the query syntax of a real JSONPath library (\'[\', \'*\', \'$\') is rejected, exactly as for an assertion\'s path), window_s (trailing window in seconds), ceiling (number), aggregation (one of \'sum\', \'max\', \'avg\'). Pings whose body is missing, is not JSON, or carries nothing numeric at that path are SKIPPED, not counted as zero — so a `start` ping never drags an average down. At most 5 guards per monitor, and window_s at most 604800 seconds (7 days). Both caps are cost, not policy: a guard re-aggregates every ping body in its window on every ping, so the per-ping work is linear in BOTH the window and the number of guards (measured: 4.2 ms/ping at a 1-hour window, 390 ms/ping at 30 days). A window longer than the 90-day ping retention would also aggregate over already-pruned rows and quietly under-report. A malformed entry is rejected before anything is written and names the offending guard.'}, 'grace_s': {'type': 'number', 'description': 'Grace period in seconds.'}, 'agent_id': {'type': 'string', 'description': "Attach this monitor to an agent from the registry, by the agent's id OR its slug (both are returned by register_agent). Omit for a monitor with no owning agent. Naming an agent that does not exist is an error — 400 UNKNOWN_AGENT — it is NEVER created implicitly; call register_agent first to get a valid agent_id. Omit to leave the monitor's current attachment (or lack of one) unchanged."}, 'period_s': {'type': 'number', 'description': "Ping interval in seconds (for schedule_kind='simple')."}, 'ci_branch': {'type': 'string', 'description': "CI filter: only count runs on this branch, e.g. 'main'. REQUIRES ci_provider, and the API enforces it: without a CI binding the request is refused with 400 FIELD_NOT_IN_SHAPE rather than accepted and discarded. Subject to the SAME upsert exception as ci_workflow — create_monitor on an existing slug never writes this filter; use update_monitor. WITHOUT IT a run on ANY branch — a feature branch, a fork's pull request — reports to this monitor, so somebody else's broken branch marks your monitor down. Set it to the branch whose health you actually care about, which is almost always the default branch. Omit to leave the current filter unchanged; pass an explicit JSON null to remove it. An EMPTY STRING also leaves it unchanged — that is a deliberate API compatibility rule, not a bug, so an empty string cannot be used to clear the filter."}, 'cron_expr': {'type': 'string', 'description': "5-field cron expression (for schedule_kind='cron')."}, 'probe_url': {'type': 'string', 'description': "http monitors only: the absolute http/https URL to probe. Required when monitor_type='http'. The host is resolved at write time and rejected if it resolves only to private/link-local addresses. Omit to leave unchanged."}, 'assertions': {'type': 'string', 'description': 'Output assertions: conditions the ping BODY of a successful run must satisfy, checked on every success ping. This is how you catch the job that exits zero having done nothing — a backup that wrote no rows, an export that produced an empty file. When an assertion fails, the success ping opens an incident with cause \'assertion\' naming the assertion that did not hold, exactly as a real failure would. Supply a JSON ARRAY as a string, e.g. \'[{"name":"rows written","kind":"json_path","path":"result.rows_processed","op":"gt","value":"0"}]\'. REPLACE-THE-SET: the array you send becomes the monitor\'s complete assertion set — it is NOT merged with what is already there. Omit the argument entirely to leave the current assertions untouched; pass \'[]\' to remove all of them. Fields per entry: name (required, appears in the alert), kind (required), value, path, op. kind is one of \'contains\' (body contains value as a substring), \'not_contains\' (body does not contain it), \'matches\' (body matches value as a Go RE2 regexp, max 1000 bytes), or \'json_path\' (parse the body as JSON, read the value at path, compare it against value with op). contains/not_contains/matches require value; json_path requires path and op and ignores them otherwise. path is a DOTTED path only (\'a.b.c\') — the query syntax of a real JSONPath library (\'[\', \'*\', \'$\') is rejected. op is one of \'eq\', \'ne\', \'gt\', \'gte\', \'lt\', \'lte\'. Comparison rule for json_path: when BOTH the value read from the body and the value you supplied parse as numbers the comparison is numeric, otherwise both sides are compared as strings — so with op \'gt\', value \'3\' beats \'12.5\' lexically but loses numerically, and \'rows_processed gt 0\' means what it looks like it means. At most 20 assertions per monitor. A malformed entry (uncompilable regexp, a path carrying query syntax, an unknown kind or op) is rejected before anything is written and names the offending assertion.'}, 'ci_workflow': {'type': 'string', 'description': "CI filter: only count runs of the workflow / pipeline / job with this exact name. REQUIRES ci_provider, and the API enforces it: without a CI binding this filter has nowhere to be stored, so the request is refused with 400 FIELD_NOT_IN_SHAPE rather than accepted and discarded. Note that monitor_type='ci' does NOT bind anything on its own — ci_provider does. ONE EXCEPTION, and it is on the path agents use most, so do not rely on the enforcement here: create_monitor on a slug that ALREADY EXISTS is an upsert, and the upsert never writes this filter. With ci_provider in the same call the request is accepted and the filter is silently discarded; without it the request is refused, and doing what the error advises — adding ci_provider — reaches the discarding case instead. Set this filter with update_monitor, which does persist it. WITHOUT IT, EVERY workflow in the repository reports to this monitor — so one unrelated failing workflow opens an incident against a job that is perfectly healthy, and a green run of a different workflow clears an incident the real job never recovered from. Set it whenever the repository has more than one workflow. Omit to leave the current filter unchanged; pass an explicit JSON null to remove it. An EMPTY STRING also leaves it unchanged — that is a deliberate API compatibility rule, not a bug, so an empty string cannot be used to clear the filter."}, 'monitor_from': {'type': 'string', 'description': "DORMANT UNTIL: an RFC 3339 timestamp before which no deadline is computed and no incident can open — the monitor is fully configured but not yet armed. Use it when you provision ahead of the work: a monitor for a job that does not start running until next Monday is otherwise 'late' from the moment you create it, which is a false alert on day one. The first-run deadline is seeded as monitor_from + grace_s. Default: unset, meaning deadlines start immediately. Example: '2026-01-01T00:00:00Z'. Omit to leave the monitor's current value unchanged."}, 'probe_method': {'type': 'string', 'description': "http monitors only: the HTTP method the probe sends. One of 'GET', 'HEAD', 'POST'. Default 'GET'. Use 'HEAD' for a cheap liveness check when the body does not matter — but note it returns no body, so probe_expected_body cannot match anything. Omit to leave unchanged."}, 'max_runtime_s': {'type': 'number', 'description': "Maximum seconds a single run may take before it is reported overdue (the 'overrun' rule), measured from the run's start ping. Omit to fall back to grace_s. This is how a long job avoids being flagged overdue while still being detected quickly if it goes silent: e.g. grace_s=600 with max_runtime_s=14400 alerts 10 minutes after a missed ping but tolerates a 4-hour run. It replaces grace_s for the overrun deadline ONLY — the silence rule and the first-run deadline still use grace_s. Range 60-31536000. Not supported on http monitors: a probe has no start/success pair, so the overrun rule can never fire and the API returns 400 MAX_RUNTIME_NOT_SUPPORTED (use probe_timeout_s to bound a single probe). Omit to leave the monitor's current value unchanged; pass 0 to clear it and fall back to grace_s."}, 'schedule_kind': {'type': 'string', 'description': "'simple', 'cron', or 'on_demand'. NOT ACCEPTED on an http monitor, together with period_s, cron_expr and tz: its schedule is derived from probe_interval_s, so the API refuses all four with 400 FIELD_NOT_IN_SHAPE. 'on_demand' means no cadence at all: no period_s, no cron_expr — the API returns 400 if either is supplied — and, by default, NO ABSENCE DEADLINES ARE ARMED BETWEEN RUNS. What this trades away: nothing tells you if the agent is never invoked again; silence between runs is invisible unless you opt in to expect_every_s. What it buys: a healthy agent that nobody happens to invoke for a week never generates a false 'late' or 'down' for simply not having been asked to run. Only run-scoped detection still applies once a run starts — max_runtime_s (overrun), step_timeout_s (stall), blocked_timeout_s (stuck on a human) — because those are anchored to a run's own start ping, not to a cadence. IMPORTANT: if you would be alarmed to find this agent silent for hours, set expect_every_s as well — it is the silence floor, and it is the only thing that makes an on_demand monitor detect absence at all. Choose 'simple'/'cron' when the agent is supposed to run on a cadence; choose 'on_demand' when invocation is inherently irregular and a quiet stretch between runs is expected, not a symptom."}, 'trace_content': {'enum': ['dropped', 'redacted'], 'type': 'string', 'description': "What this monitor's traces keep of prompt, command and tool content. 'dropped' (the default) removes it; 'redacted' keeps it, with every secret-shaped value redacted when it arrives. Only a person should choose 'redacted': never set it on your own initiative, only when the person you work for has asked for content to be stored. Omit to leave the monitor's current value unchanged."}, 'expect_every_s': {'type': 'number', 'description': "SILENCE FLOOR in seconds: open a 'silence' incident if NO ping of any kind — success, start, fail, step — has arrived within this window, regardless of the schedule. It is anchored on the monitor's last activity, not on a cadence, which is what makes it the ONLY absence rule an 'on_demand' monitor can have: that schedule_kind arms nothing between runs, so without this field an on_demand monitor reads 'up' forever no matter how long the agent stays dark. Set it on any on_demand agent monitor you would be alarmed to find silent — that is what it is for. It does NOT fire mid-run: while a run is in flight (a start ping is outstanding) the floor stands down entirely and the run clock owns detection (max_runtime_s, step_timeout_s), so a legitimate 4-hour run that reports nothing is still not an incident. A 'blocked' ping also pauses it, bounded by blocked_timeout_s. On 'simple'/'cron' monitors it is a backstop rather than the main rule: it joins the existing deadline as whichever is SOONER, so it can tighten detection under a long cadence (a daily cron has a ~25-hour blind window) but can never loosen it. Default: unset, which means no floor and is exactly how every monitor behaved before this field existed. Range 60-31536000. Accepted on every monitor_type and every schedule_kind. Omit to leave the monitor's current value unchanged; pass 0 to clear it and turn the silence floor off."}, 'step_timeout_s': {'type': 'number', 'description': "Progress budget in seconds: how long an armed run may go without reporting a step before a 'stalled' incident opens (the stall rule). The clock is anchored on the LATER of the run's start ping and its most recent step, so a run that wedges before its first step is caught too. Reach for this when 'still running' and 'still making progress' are different things — a long agent loop, a multi-stage pipeline, a migration. max_runtime_s alone tells you nothing until the whole budget expires; step_timeout_s=300 on a 4-hour budget tells you within five minutes, and names the last step that reported. To use it the run must report steps: call get_ping_instructions and use curl_step (POST <ping_url>/step?rid=<run-id>&step=<name>). A monitor with step_timeout_s set whose job never reports a step will open a stalled incident on EVERY run — set the field and instrument the job in the same change. Stall detection needs your job to call /step: the Claude Code hook and the reporting prompt do not send steps, so leave step_timeout_s unset for them. Default: unset, which disables stall detection entirely; a monitor that sets nothing behaves exactly as it did before this field existed. Range 10-86400. Two constraints. (1) It must be strictly LESS than the effective run budget, COALESCE(max_runtime_s, grace_s), or the API returns 400 STEP_TIMEOUT_EXCEEDS_BUDGET — at or above the budget the run overruns first, so the stall rule could never fire. (2) Not supported on http monitors: a probe never arms a run and has no /step endpoint to call, so the API returns 400 STEP_TIMEOUT_NOT_SUPPORTED. A step resets the stall clock ONLY — it never extends max_runtime_s, so an agent that reports progress forever still overruns. Omit to leave the monitor's current value unchanged; pass 0 to clear it and disable stall detection."}, 'probe_timeout_s': {'type': 'number', 'description': "http monitors only: how many seconds a single probe may take before it counts as a failure. Range 1-30, default 10. This is the http equivalent of max_runtime_s, which http monitors reject: it is the only way to say 'answering, but far too slowly to be healthy'. Omit to leave unchanged."}, 'runaway_ceiling': {'type': 'number', 'description': "PING-RATE CEILING: the maximum number of pings this monitor may receive in a rolling one-hour window. Exceeding it opens a 'runaway' incident. This is the rule that catches a job or agent stuck in a LOOP — the failure every other rule misses, because a looping agent is pinging enthusiastically and therefore reads 'up' the whole time it is burning tokens or money. Set it a little above the monitor's real cadence: a job that runs every 15 minutes sends about 4 pings/hour, so 20 absorbs retries and still catches a loop. It is RATE-based, so failure_threshold does not gate it and neither does any run budget. Default: unset, which disables the runaway rule entirely. Omit to leave the monitor's current value unchanged; pass 0 to clear it and turn the runaway rule off."}, 'notify_min_run_s': {'type': 'number', 'description': "NOTIFICATION DURATION FLOOR in seconds: a run SHORTER than this does not produce an INFO-CLASS notification (success, started, every-run, note). This exists for exactly one problem: on an agent monitor, one run is one task you asked for, so asking the agent 'what's 2+2' produces a start and a success notification exactly like a 56-minute deploy does. If you have routed success/started/every-run/note to a destination, you WILL be paged for trivial runs unless you set this. IT NEVER SUPPRESSES A FAILURE. down, fail, recovery and blocked are alert-class and are never affected by this field, however short the run — a run that failed in two seconds is exactly what you need to hear about, and this field cannot silence that, structurally, no matter how it is set. It also never suppresses 'started': a run's duration does not exist yet the moment it begins, so started is always reported regardless of this floor. And it never suppresses an event whose duration could not be measured at all (e.g. a bare success with no preceding start ping) — an unknown duration always means 'notify', never 'suppress'. Default: unset, which means no floor and is exactly how every monitor behaved before this field existed. Range 60-31536000. Not supported on http monitors: an http probe has no start/success pair, so its run duration is never measured and the floor could never apply (the API returns 400 NOTIFY_MIN_RUN_NOT_SUPPORTED). Omit to leave the monitor's current value unchanged; pass 0 to clear it and turn the notification duration floor off."}, 'probe_interval_s': {'type': 'number', 'description': "http monitors only: how often to probe, in seconds. Required when monitor_type='http'. Range 30-86400. Omit to leave unchanged."}, 'blocked_timeout_s': {'type': 'number', 'description': "Maximum seconds a run may sit in the 'blocked' state (an agent reported it is waiting on a human) before a 'blocked' incident opens. UNSET DOES NOT MEAN WAIT FOREVER: omitting this does not disable the timeout, it falls back to the default, which is 24 HOURS — an agent still blocked 24 hours after reporting so, with this field never set, gets a 'blocked' incident regardless. Lower it to be paged sooner when a stuck approval is urgent; raise it for work that legitimately waits on a human for longer than a day. This is distinct from the non-incident 'blocked' notification a route on the 'blocked' event type delivers (see set_route), which is held 10 minutes and sent only if the run is still blocked then, at most once per blocked stretch of a run; this field governs the separate incident that opens only if the wait outlives the timeout. Accepted on every monitor_type: unlike max_runtime_s/step_timeout_s it has no run-scoped precondition an http monitor could fail, so there is nothing to reject. Omit to leave the monitor's current value unchanged; pass 0 to clear it and fall back to the 24h default."}, 'failure_threshold': {'type': 'number', 'description': "Number of consecutive failures required before an incident opens. Default 1 (open on the very first failure). This is how you stop a single transient blip from paging someone: set 2-5 on a job that fails occasionally for reasons that resolve themselves, and no incident opens until that many runs in a row have failed. Any success resets the count to zero. It gates the 'fail' cause ONLY — silence (a missed ping), overrun, never_started and runaway are time- or rate-based, so a consecutive count means nothing for them and they are never delayed by it. A run that ends on the model provider's API error (server_error, overloaded, rate_limit) neither counts toward it nor resets it: those open an 'upstream' incident once 3 runs in a row end that way, or this many when it is higher. Range 1-100. Omit to leave the monitor's current threshold unchanged."}, 'probe_expected_body': {'type': 'string', 'description': 'http monitors only: a substring that MUST appear in the response body for the probe to count as healthy. THIS IS THE DIFFERENCE BETWEEN \'the server answered\' AND \'the app works\': a broken app that renders an error page still returns 200, passes a status-only check, and leaves the monitor green. Match on something only a healthy response contains, e.g. \'"status":"ok"\'. Substring match, not a regex, and case-sensitive. Default: empty, meaning the body is not inspected at all. Omit to leave unchanged; pass an explicit JSON null to stop inspecting the body. An empty string leaves it unchanged, so it cannot be cleared that way.'}, 'probe_expected_status': {'type': 'number', 'description': "http monitors only: the EXACT HTTP status code that counts as healthy. Default 200; any other code fails the probe. Set it when the healthy answer is not 200 — 204 for a no-content health endpoint, or 301 when what you are checking is that a redirect still exists (pair that with probe_follow_redirects=false, or the probe will follow it and see the destination's status instead). Omit to leave unchanged."}, 'probe_follow_redirects': {'type': 'boolean', 'description': 'http monitors only: whether the probe follows 3xx redirects. Default false. Leaving it false is usually what you want: the redirect itself is then compared against probe_expected_status like any other response, so a site that starts redirecting to a login wall, a parking page or an outage notice is CAUGHT rather than silently followed to a healthy-looking 200. Set true only when the URL you are checking is legitimately a redirect to the thing you actually care about. Omit to leave unchanged; pass false to turn following back off.'}}}
Schéma d’entrée
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Status page UUID, from list_status_pages.'}, 'slug': {'type': 'string', 'description': 'New URL slug. Omit to leave unchanged — which is almost always right, because changing it breaks every link already handed out. Same format rules and same global uniqueness as on create; a taken slug returns 409.'}, 'title': {'type': 'string', 'description': 'New page title. Omit to leave unchanged.'}, 'check_ids': {'type': 'string', 'description': "Comma-separated monitor UUIDs to show on the page, in no particular order. Get them from list_monitors. Every id must belong to this project — an unknown or cross-project id returns 400 and nothing is saved. An empty value is legal and produces a page with no monitors on it. REPLACES the page's whole monitor set. Omit to leave the current set alone."}, 'visibility': {'type': 'string', 'description': "'private' (default) or 'public'. 'public' means the page is served at a guessable-free but UNAUTHENTICATED URL: anyone with the link sees the title, the name of every monitor on it, and its up/down history. Monitor names are frequently internal ('billing-reconciler', 'acme-corp-nightly-sync'), so treat this as publishing them. Choose 'private' unless the user has actually asked for a page other people can see. The free tier allows exactly ONE public page per project; a second returns 403. Omit to leave unchanged. Switching a page from private to public publishes every monitor name already on it."}}}
Modifications récentes des outils
Serveurs MCP similaires
shiply
Publishes and manages web applications with SQL databases, serverless functions, domains, email tests, client intake, contracts, …
Fine Structure
Builds and hosts full-stack applications from prompts and supports agent communication through WhatsApp and email.
mcp
Provisions and manages cloud instances, networks, storage, databases, backups, VPNs, application deployments, and autocoding-agen…
EchoRelay
Manages a hosted API relay runtime, including projects, endpoints, credentials, publishing, keys, delivery logs, and dead-letter …
Supero
Builds, validates, publishes, tests, and deploys multi-tenant applications with schemas, CRUD operations, external data connector…
Replicate
Provides access to Replicate models, versions, collections, hardware, predictions, and deployments for running and managing hoste…
Relin
Configures webhook sources and destinations, monitors delivery health, investigates failures, manages retries, and replays events.
Scalix Cloud
Provides managed databases, SQL tools, container builds, persistent Linux machines, domains, scheduled functions, storage, and pr…