Skip to content

Health and monitoring API ​

What this page covers ​

This page documents every HTTP and WebSocket endpoint in the health and monitoring areas of the mmune API: health, monitoring (alert rules, metrics, AI usage and budgets), watchdog, drift, healing, containment, incidents, the live-refresh cursor, notifications, ticketing, history stats, the agent dashboard, agent logs. For each endpoint you get the method, full path, purpose, required permission, parameters, response fields, status codes and a curl example. Five worked examples in Python (requests) and TypeScript (fetch) follow the reference.

Prerequisites: a running mmune backend, a bearer token for a user with at least the read permission (write for anything that changes state, role admin for the admin-only calls), and the base URL. Examples use https://mmune.example.com as the host and <TOKEN> as the token. Authentication, error format, pagination and rate-limit conventions are defined in API overview and only summarised here. Outbound webhook and event delivery details are in Webhooks and events.

Endpoint summary ​

All paths are relative to the host.

RouterMethodPathPurpose
healthGET/api/v1/health/Basic health summary of registered components
healthGET/api/v1/health/detailedRun health checks and return component, LLM and embedding status
healthGET/api/v1/health/llmLLM reachability probe, 503 when unreachable
healthGET/api/v1/health/sap-alertsList SAP drift alerts that were delivered
monitoringGET/api/v1/monitoring/alert-rulesList categorical alert rules
monitoringPOST/api/v1/monitoring/alert-rulesCreate an alert rule
monitoringPUT/api/v1/monitoring/alert-rules/{rule_id}Update an alert rule
monitoringDELETE/api/v1/monitoring/alert-rules/{rule_id}Delete an alert rule
monitoringGET/api/v1/monitoring/metricsPrometheus text metrics
monitoringGET/api/v1/monitoring/summaryMetrics registry summary
monitoringGET/api/v1/monitoring/ai-usageAggregated AI token and cost usage
monitoringGET/api/v1/monitoring/estate-loadAggregated query load mmune placed on client integrations
monitoringGET/api/v1/monitoring/ai-budgets/globalGlobal AI budget defaults
monitoringGET/api/v1/monitoring/ai-budgetsList per-integration budget overrides
monitoringGET/api/v1/monitoring/ai-budgets/{integration_id}Get one budget override
monitoringPUT/api/v1/monitoring/ai-budgets/{integration_id}Create or replace a budget override
monitoringDELETE/api/v1/monitoring/ai-budgets/{integration_id}Remove a budget override
monitoringGET/api/v1/monitoring/detection-metricsDetection precision, recall and latency
monitoringWebSocket/api/v1/monitoring/wsReal-time dashboard stream
monitoringGET/api/v1/monitoring/poll-intervalRead the autonomous polling interval
monitoringPOST/api/v1/monitoring/poll-intervalChange the autonomous polling interval
watchdogPOST/api/v1/watchdog/startStart the watchdog service
watchdogPOST/api/v1/watchdog/stopStop the watchdog service
watchdogGET/api/v1/watchdog/statusWatchdog service counters
watchdogPOST/api/v1/watchdog/integrationsAdd an integration to watch
watchdogDELETE/api/v1/watchdog/integrations/{session_id}Remove a watch session
watchdogGET/api/v1/watchdog/cdc/connectors/{integration_id}Debezium connector status
watchdogGET/api/v1/watchdog/sessionsList watch sessions
watchdogGET/api/v1/watchdog/sessions/{session_id}One watch session
watchdogGET/api/v1/watchdog/zombiesConfirmed zombie and flow-stall findings
watchdogGET/api/v1/watchdog/flowFlow health for TIBCO integrations
watchdogPOST/api/v1/watchdog/integrations/{integration_id}/priorityMark an integration as priority
watchdogGET/api/v1/watchdog/manager/statusWatchdog manager instances
driftGET/api/v1/drift/reportsList drift findings from durable storage
driftGET/api/v1/drift/statsDrift detection statistics
driftGET/api/v1/drift/value-driftList value-drift findings
driftPOST/api/v1/drift/value-drift/muteMute value-drift alerts for one field
driftPOST/api/v1/drift/value-drift/unmuteUnmute one field
healingGET/api/v1/healing/configRead healing mode and auto-apply threshold
healingPUT/api/v1/healing/configChange healing mode and threshold
healingGET/api/v1/healing/connection-healthConnection health for one integration or overall statistics
healingPOST/api/v1/healing/connection-health/checkRun an active connection health check
healingGET/api/v1/healing/broken-connectionsList broken connections
healingPOST/api/v1/healing/suggest-remappingSuggest remappings for a broken field
healingPOST/api/v1/healing/apply-suggestionApply a remapping suggestion
healingGET/api/v1/healing/health-history/{integration_id}Connection health history
healingGET/api/v1/healing/statisticsHealing service statistics
containmentGET/api/v1/containment/configRead containment settings
containmentPUT/api/v1/containment/configChange containment settings
containmentGET/api/v1/containment/marksList active quarantine marks
containmentPOST/api/v1/containment/marks/{mark_id}/liftLift a quarantine mark
containmentPOST/api/v1/containment/incidents/{incident_id}/revalidateRe-validate an incident and lift marks that pass
incidentsGET/api/v1/incidentsList incidents, paginated
incidentsGET/api/v1/incidents/ownersList incident owners
incidentsPOST/api/v1/incidents/ownersCreate an owner
incidentsPUT/api/v1/incidents/owners/{owner_id}Update an owner
incidentsDELETE/api/v1/incidents/owners/{owner_id}Delete an owner
incidentsGET/api/v1/incidents/slaRead the SLA policy
incidentsPUT/api/v1/incidents/slaReplace the SLA policy
incidentsGET/api/v1/incidents/rca-configRead the root-cause analysis mode
incidentsPUT/api/v1/incidents/rca-configChange the root-cause analysis mode
incidentsGET/api/v1/incidents/{incident_id}Incident detail with member events
incidentsGET/api/v1/incidents/{incident_id}/rcaLatest root-cause analysis
incidentsPOST/api/v1/incidents/{incident_id}/rcaGenerate a new root-cause analysis
incidentsGET/api/v1/incidents/{incident_id}/runbook.mdLatest runbook as Markdown
incidentsPOST/api/v1/incidents/{incident_id}/acknowledgeAcknowledge an incident
incidentsPOST/api/v1/incidents/{incident_id}/outcomeLabel an incident true positive, false positive or inconclusive
incidentsPOST/api/v1/incidents/{incident_id}/resolveResolve an incident
eventsGET/api/v1/events/cursorLive-refresh cursor counters
notificationsGET/api/v1/notifications/channelsList notification channels
notificationsGET/api/v1/notifications/channels/{channel}One channel
notificationsPUT/api/v1/notifications/channels/{channel}Enable, disable or set the endpoint of a channel
notificationsGET/api/v1/notifications/deliveriesRecent delivery attempts
notificationsPOST/api/v1/notifications/channels/{channel}/testSend a test notification
ticketingGET/api/v1/ticketing/providersSupported ticketing systems
ticketingGET/api/v1/ticketing/configRead the ticketing configuration
ticketingPOST/api/v1/ticketing/configSave the ticketing configuration
ticketingDELETE/api/v1/ticketing/configRemove the ticketing configuration
ticketingPOST/api/v1/ticketing/testTest the ticketing connection
ticketingPOST/api/v1/ticketing/ticketsCreate a ticket
statsGET/api/v1/stats/historyDaily integration and issue counts
dashboardGET/api/v1/dashboard/overviewAgent system overview
dashboardGET/api/v1/dashboard/agents/{agent_type}/healthHealth of one agent type
dashboardGET/api/v1/dashboard/errorsRecent agent errors
dashboardGET/api/v1/dashboard/performance/trendsDaily agent performance trend
dashboardGET/api/v1/dashboard/learning/progressLearning accuracy and volume gate
logsGET/api/v1/logsRead agent logs
logsDELETE/api/v1/logsClear agent logs

Two related routes are not in the table. GET /health is a public liveness probe that returns {"status": "ok"} with no credential. GET /api/v1/monitoring/data returns the same dashboard payload that the WebSocket sends.

Conventions used on this page ​

Every request needs Authorization: Bearer <TOKEN>. Every request is checked before it reaches the endpoint, so a missing token (or an Authorization header that is not a Bearer token) returns 401 with {"detail": "Not authenticated"} on every endpoint below, except the public GET /health noted above. A token that is present but bad also returns 401, with a different detail: Invalid token (which is also what an expired token returns), Token has been revoked, or MFA verification required. mmune then requires the read permission for all methods, requires write for every method other than GET, HEAD and OPTIONS, and requires role admin for these calls: PUT /api/v1/healing/config, PUT /api/v1/containment/config and DELETE /api/v1/logs. A user whose role is admin passes every permission check. A refused request returns 403 with {"detail": {"code": "insufficient_permission", "required": "<read|write|admin>", "message": "..."}}. Some endpoints check the permission a second time and then return {"detail": "Permission 'write' required"}.

In the tables below, the Permission line states the effective requirement: read, write or admin. Request bodies are JSON with Content-Type: application/json. A body or query value that fails validation returns 422 with the standard validation error body. An unexpected failure returns 500 with {"detail": "<message>"}. Timestamps are ISO 8601 strings. See API overview for the full conventions.

Pagination applies to one endpoint here, GET /api/v1/incidents, which takes page (default 1) and page_size (default 10, capped at 100). Other list endpoints use a limit parameter and no cursor, as stated per endpoint.

The curl examples assume:

bash
export TOKEN=<TOKEN>

Statuses and enumerations ​

This section lists every status and enumeration exactly as the API accepts and returns it. Several different "health" vocabularies exist and they are not interchangeable.

Component health, used by /api/v1/health/ and /api/v1/health/detailed for each component: healthy, unhealthy, degraded, unknown. The summary field overall_status is computed from the critical components only (all components if none is critical). It is healthy when every one of them is healthy, otherwise unhealthy if any is unhealthy, otherwise degraded if any is degraded. It stays unknown when no check has run yet, and also when the results are a mix that matches none of those rules.

Connection health, used by /api/v1/healing/*: healthy, degraded, unhealthy, broken, unknown.

Connection error types: connection_timeout, authentication_failed, field_not_found, schema_mismatch, rate_limit_exceeded, server_error, network_error, unknown.

Watchdog session health (last_health_status), set by the watchdog on each session: unknown before the first check, then healthy, degraded or broken. A session becomes broken when any active finding has severity critical and degraded for other findings. The session status field defaults to active.

Watchdog service status: starting, running, stopping, stopped, error. Watchdog manager instance status: healthy, degraded, unhealthy, offline.

Severities. Findings, alert rules and drift reports accept and rank these values, lower-cased: info (0), low (1), warning (2), medium (2), high (3), critical (4). An empty severity is treated as medium. Incident SLA policy and incident severity buckets use four values: critical, high, medium, low. For SLA purposes warning maps to medium and info maps to low. Ticket priorities use critical, high, medium, low.

Incident status: open, acknowledged, resolved. Incident outcome: true_positive, false_positive, inconclusive. SLA state: on_track, ack_breached, resolve_breached. Resolution mode: manual when resolved through the API, auto when the system resolves the incident (recovery re-validation or ticket sync). Root-cause analysis mode: off, on_demand, auto_on_open.

Detection dimensions, the finding types used in alert-rule matching and incident dimensions: schema_drift, value_drift, watchdog_connection, zombie_flow, duplicate, security_pii, stale_feed, containment, job_failure, reconciliation. GET /api/v1/drift/reports returns only schema_drift, value_drift, stale_feed, zombie_flow and job_failure.

Live-refresh domains for GET /api/v1/events/cursor: findings, lineage, health.

Healing mode: alert (suggest only) and autonomous (auto-apply the top suggestion when its confidence meets the threshold). Remapping suggestion strategy: semantic_match, pattern_match, ml_prediction, historical. Containment mode: off, observe, active.

Notification channels: slack, webhook, pagerduty, servicenow, teams, email_digest. Delivery status: ok, failed, queued, skipped_unconfigured, skipped_disabled, skipped_deduped, skipped_contained, skipped_circuit_open. Circuit breaker state: closed, open, half_open. Channel config source: db, env, default.

Ticketing providers: jira, servicenow.

Alert-rule action types: route_channels, escalate_severity, suppress, create_ticket.

The only event-style read on this page is the cursor endpoint. See Webhooks and events for what leaves the system.


Health router ​

Base path /api/v1/health.

GET /api/v1/health/ ​

Returns the current health summary without running a new check.

Permission: read. Parameters: none.

Response 200:

FieldTypeDescription
timestampstring or nullcomponents.last_check_time
componentsobjectHealth summary: overall_status, total_components, last_check_time, metrics, and components (map of component name to result)

Because this endpoint does not run a new check, components.overall_status is unknown until a check has run. Use /detailed to force a run.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/health/

GET /api/v1/health/detailed ​

Runs the health checks, then returns the summary plus LLM and embedding status.

Permission: read. Parameters: none.

Response 200 is the health summary object (overall_status, total_components, last_check_time, metrics, components) with two more keys:

FieldTypeDescription
llmobjectLocal or cloud LLM status. Cloud mode: mode (cloud), provider, reachable, chat_model_loaded (null), embed_model_loaded (null). Local mode (when MMUNE_LOCAL_LLM_BASE_URL is set): mode (local), provider (ollama), endpoint, chat_model, embed_model, reachable, chat_model_loaded, embed_model_loaded, and either models_available (list) or error (string).
embeddingsobjectactive_model, stored_models (list), mismatch (bool), message, error. mismatch is true when stored embedding model names differ from the active provider.

The LLM and embedding checks report failures inside the body instead of returning an error status, so a 200 does not mean the LLM is up. Check llm.reachable.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/health/detailed

GET /api/v1/health/llm ​

Checks the LLM. Returns 200 when healthy and 503 when not, which suits load balancer and acceptance checks.

Permission: read. Parameters: none.

Healthy means: in local mode, the runtime is reachable and the configured chat model is loaded; in cloud mode, the AI registry has a genuine healthy provider (placeholder keys do not count).

Response 200 or 503: status (healthy or unhealthy) plus all fields of the llm object described under /detailed.

bash
curl -s -o /dev/null -w "%{http_code}\n" -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/health/llm

GET /api/v1/health/sap-alerts ​

Lists SAP drift alerts that were delivered, newest first.

Permission: read.

NameInTypeRequiredDescription
integration_idquerystringnoFilter to one integration
severityquerystringnoExact match, for example high
limitqueryintegernoDefault 100, maximum 500 (larger values return 422)

Response 200: an array of alerts, each with id, drift_id, integration_id, drift_type, object_name, field_name, severity, message, detected_at, delivered_at, delivery_channels, delivery_status. If the alerts cannot be read, the endpoint still returns 200 with {"error": "<message>", "alerts": []}, so check for an error key.

bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/health/sap-alerts?severity=high&limit=20"

Monitoring router ​

Base path /api/v1/monitoring.

GET /api/v1/monitoring/alert-rules ​

Lists categorical alert rules, sorted by priority then id. Permission: read. Parameters: none.

Response 200: rules (array of rule objects, see below) and count.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/monitoring/alert-rules

POST /api/v1/monitoring/alert-rules ​

Creates a rule. Rules are persisted and survive restarts. They are evaluated before notification fan-out and can route to channels, raise severity, suppress, or request a ticket. Permission: write.

NameTypeRequiredDefaultDescription
idstringnoderived from nameRule id. If omitted it is the lower-cased name with runs of characters outside a-z0-9_ replaced by _, then leading and trailing _ removed
namestringyesDisplay name
descriptionstringno""
priorityintegerno100Lower values sort first
enabledbooleannotrue
match.dimensionsarrayno[]Detection dimension values. Empty matches all
match.integration_idsarrayno[]Integration ids or *. Empty matches all
match.severitiesarrayno[]Empty matches all
match.object_key_patternstring or nullnonullGlob on the object key or table name
match.kindsarrayno[]drift_type, kind or event strings. Empty matches all
actionsarrayno[]Action objects, validated on write

Action objects:

typeFieldsNotes
route_channelschannels (list of notification channel names)Channel names must be in the notification channel list, otherwise 400
escalate_severityto_severity (a severity value) or bump_one_level (true)One of the two is required
suppresssuppress_outbound (default true), cooldown_seconds (optional integer)
create_ticketprovider (optional), project_override (optional)Both fields are optional

Response 201: the stored rule: id, name, description, priority, enabled, match (object with the five match fields), actions.

Status codes: 201, 400 (cannot derive an id from the name, unknown action type, bad action fields, unknown channel), 409 (rule id already exists).

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"name":"Page on critical schema drift","match":{"dimensions":["schema_drift"],"severities":["critical"]},"actions":[{"type":"route_channels","channels":["pagerduty"]}]}' \
  https://mmune.example.com/api/v1/monitoring/alert-rules

PUT /api/v1/monitoring/alert-rules/ ​

Updates fields on an existing rule. Only fields present in the body change. Permission: write.

Path: rule_id (string). Body fields, all optional: name, description, priority, enabled, match (same object as on create, replaces the match criteria), actions (validated as on create).

Response 200: the updated rule. Status codes: 200, 400 (invalid actions), 404 (rule not found).

bash
curl -s -X PUT -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"enabled":false}' https://mmune.example.com/api/v1/monitoring/alert-rules/page_on_critical_schema_drift

DELETE /api/v1/monitoring/alert-rules/ ​

Deletes a rule. Permission: write. Response 200: {"deleted": "<rule_id>"}. Status codes: 200, 404.

bash
curl -s -X DELETE -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/monitoring/alert-rules/page_on_critical_schema_drift

GET /api/v1/monitoring/metrics ​

Prometheus exposition format, content type text/plain; version=0.0.4; charset=utf-8. Permission: read. Parameters: none.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/monitoring/metrics

GET /api/v1/monitoring/summary ​

Summary of the metrics registry. Permission: read. Response 200 includes total_metrics and a components object with per-component metric counts.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/monitoring/summary

GET /api/v1/monitoring/ai-usage ​

Aggregated AI usage grouped by one or more fields. Permission: read.

NameInTypeRequiredDescription
daysqueryintegernoLookback window in days. Default 7. 0 means all time. Must be 0 or greater
group_byquerystringnoComma-separated fields from provider, model, component, integration_id, day. Default provider

Response 200: window_days, group_by (list), groups (array). Each group has group_key (map of field to string or null), event_count, input_tokens, output_tokens, total_tokens, estimated_cost_usd.

Status codes: 200, 400 (empty group_by or an unknown field; the message lists the allowed fields).

bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/monitoring/ai-usage?days=30&group_by=provider,day"

GET /api/v1/monitoring/estate-load ​

Counts the queries mmune has issued against each client integration. Permission: read.

NameInTypeRequiredDescription
daysqueryintegernoDefault 7. 0 means all time
group_byquerystringnoComma-separated fields from integration_id, component, query_class, day. Default integration_id

Response 200: window_days, group_by, groups. Each group has group_key, query_count, total_rows_read. Status codes: 200, 400 (same rules as ai-usage).

bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/monitoring/estate-load?group_by=integration_id,component"

GET /api/v1/monitoring/ai-budgets/global ​

The global budget defaults. These come from the environment variables MMUNE_AI_BUDGET_DAILY_USD, MMUNE_AI_BUDGET_MONTHLY_USD and MMUNE_ESTATE_LOAD_CEILING_PER_CYCLE and cannot be changed through the API. Permission: read.

Response 200: daily_usd (number or null), monthly_usd (number or null), estate_load_ceiling_per_cycle (integer or null). Null means unlimited.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/monitoring/ai-budgets/global

GET /api/v1/monitoring/ai-budgets ​

Lists all per-integration overrides. Permission: read. Response 200: an array of {integration_id, daily_usd, monthly_usd, estate_load_ceiling_per_cycle}.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/monitoring/ai-budgets

GET /api/v1/monitoring/ai-budgets/ ​

One override. Permission: read. Response 200: same four fields as the list entries. Status codes: 200, 404 (no override, meaning the integration uses the global defaults).

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/monitoring/ai-budgets/erp-prod

PUT /api/v1/monitoring/ai-budgets/ ​

Creates or replaces an override. This is replace semantics: any field you omit or send as null clears that override and falls back to the global default. Permission: write.

NameTypeRequiredDefaultDescription
daily_usdnumbernonullMinimum 0
monthly_usdnumbernonullMinimum 0
estate_load_ceiling_per_cycleintegernonullMinimum 0

Response 200: the stored override. Status codes: 200, 422.

bash
curl -s -X PUT -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"daily_usd":5.0,"monthly_usd":100.0}' https://mmune.example.com/api/v1/monitoring/ai-budgets/erp-prod

DELETE /api/v1/monitoring/ai-budgets/ ​

Removes the override. Permission: write. Response 200: {"deleted": "<integration_id>"}. Status codes: 200, 404.

bash
curl -s -X DELETE -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/monitoring/ai-budgets/erp-prod

GET /api/v1/monitoring/detection-metrics ​

Latest precision, recall and detection latency, computed by comparing labelled ground truth with the detections that fired. Read-only. Permission: read.

Response 200: overall and dimensions (map of dimension name to metrics). Each metrics object has dimension, total_labels, true_positives, false_positives, recall, precision, latency_p50_seconds, latency_p90_seconds, latency_max_seconds. The ratio and latency fields are null when there is nothing to compute them from.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/monitoring/detection-metrics

WebSocket /api/v1/monitoring/ws ​

Real-time monitoring stream. The server accepts the connection, sends the current dashboard payload immediately, registers the socket for later pushes, then reads and discards any text you send. If the dashboard is not available the server closes with code 1011. Permission: read (a WebSocket handshake needs only read).

Browsers cannot set an Authorization header on a WebSocket, so offer two subprotocols: mmune.bearer followed by the token. Non-browser clients may send the Authorization header instead.

javascript
const ws = new WebSocket("wss://mmune.example.com/api/v1/monitoring/ws", ["mmune.bearer", "<TOKEN>"]);
ws.onmessage = (event) => console.log(JSON.parse(event.data));

GET /api/v1/monitoring/poll-interval ​

The interval the autonomous monitor polls at. Permission: read.

Response 200: poll_interval (integer), unit (seconds), note (string). If no running monitor is found, the value comes from the environment variable MMUNE_AUTOPILOT_INTERVAL, then 120. Status codes: 200, 500.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/monitoring/poll-interval

POST /api/v1/monitoring/poll-interval ​

Changes the polling interval. Lower values increase load on client systems. Permission: write.

NameTypeRequiredDefaultDescription
intervalintegeryesSeconds, minimum 1
reasonstringnonullLogged with the change

Response 200: old_interval, new_interval, message. Status codes: 200, 404 (AutoPilot not initialized), 422, 500.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"interval":60,"reason":"active migration window"}' https://mmune.example.com/api/v1/monitoring/poll-interval

Watchdog router ​

Base path /api/v1/watchdog. The watchdog watches integrations for schema drift, value drift, stalled flows and connection loss.

POST /api/v1/watchdog/start ​

Starts the watchdog service. Permission: write. Parameters: none. Response 200: {"status": "started", "message": "..."}. Status codes: 200, 500.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/watchdog/start

POST /api/v1/watchdog/stop ​

Stops the watchdog service. Permission: write. Response 200: {"status": "stopped", "message": "..."}. Status codes: 200, 500.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/watchdog/stop

GET /api/v1/watchdog/status ​

Service counters. Permission: read. Response 200:

FieldTypeDescription
statusstringWatchdog service status (see enumerations)
sessionsintegerNumber of sessions
active_sessionsintegerSessions whose status is active
total_checksintegerSum of session check counts
total_driftsintegerSum of session drift counts
priority_integrationsintegerCount of priority integrations
circuit_breakers_openintegerIntegrations whose circuit breaker is open
bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/watchdog/status

POST /api/v1/watchdog/integrations ​

Adds an integration to the watch list. Permission: write.

NameTypeRequiredDefaultDescription
integration_idstringyesRegistered integration id
providerstringyesProvider name, for example postgres
configobjectnonullSession options. config.cdc_enabled: true turns on Postgres change data capture for this integration. A provider listed in MMUNE_CDC_ENABLED_PROVIDERS also enables it. config.check_interval is read back in session status.

Response 200: session_id, integration_id, message. The response also carries cdc_enabled (a boolean, false when CDC is off). When the session has CDC enabled it additionally carries cdc_connector_state and cdc_transport. Status codes: 200, 400 (CDC prerequisite not met, for example wal_level is not logical, or the provider does not support CDC; polling is left unchanged), 500.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"integration_id":"erp-prod","provider":"postgres","config":{"cdc_enabled":false}}' \
  https://mmune.example.com/api/v1/watchdog/integrations

DELETE /api/v1/watchdog/integrations/ ​

Removes a watch session. Permission: write. Response 200: {"message": "Integration removed from watch"}. Status codes: 200, 404 (Session not found), 500.

bash
curl -s -X DELETE -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/watchdog/integrations/<SESSION_ID>

GET /api/v1/watchdog/cdc/connectors/ ​

Debezium connector status for an integration. Although this is a GET, it needs the write permission. Permission: write.

Response 200: the connector status object. Status codes: 200, 404 (No CDC connector for integration <id>).

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/watchdog/cdc/connectors/erp-prod

GET /api/v1/watchdog/sessions ​

All watch sessions. Permission: read. Response 200: sessions (array of session objects) and count. Session object fields: session_id, integration_id, provider, status, last_health_status, started_at, last_check_at, last_check (same value as last_check_at), check_count, drift_count, check_interval.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/watchdog/sessions

GET /api/v1/watchdog/sessions/ ​

One session object (fields as above). Permission: read. Status codes: 200, 404.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/watchdog/sessions/<SESSION_ID>

GET /api/v1/watchdog/zombies ​

Confirmed zombie and flow-stall findings across all sessions. A zombie is a process or consumer that reports itself running but is not processing: a growing backlog with no active consumer, flat CPU with silent logs, or data no longer arriving at a sink. Permission: read.

Response 200: findings and count. Each finding is the stored finding fields merged with integration_id and provider. The list is sorted with critical severity first, then by descending confidence.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/watchdog/zombies

GET /api/v1/watchdog/flow ​

Flow and liveness summary for sessions whose provider is tibco. Permission: read. Response 200: integrations and count. Each entry has integration_id, health, last_flow_sample (timestamp or null), active_findings (integer) and findings (array).

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/watchdog/flow

POST /api/v1/watchdog/integrations/{integration_id}/priority ​

Marks an integration as priority or clears the mark. Permission: write.

NameInTypeRequiredDefaultDescription
integration_idpathstringyes
is_priorityquerybooleannotrueSet false to clear

Response 200: {"message": "Priority set to True for integration <id>"}. Status codes: 200, 500.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/watchdog/integrations/erp-prod/priority?is_priority=true"

GET /api/v1/watchdog/manager/status ​

Status of watchdog manager instances. Permission: read. Response 200: total_instances, healthy_instances, total_sessions, and instances (map of instance id to status, session_count, load, last_heartbeat).

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/watchdog/manager/status

Drift router ​

Base path /api/v1/drift.

GET /api/v1/drift/reports ​

Stored drift findings, newest first. They survive restarts for MMUNE_DRIFT_REPORTS_RETENTION_DAYS days (default 90). Resolved findings are omitted by default. Permission: read.

NameInTypeRequiredDefaultDescription
integration_idquerystringnoFilter to one integration
limitqueryintegerno100Maximum reports.
include_resolvedquerybooleannofalseInclude findings the watchdog has marked resolved

Response 200: drifts and reports (two keys holding the same array) and count. Each report has drift_id, id, integration_id, dimension, object_name, field_name, drift_type, reasoning, detected_at, resolved_at, severity (defaults to warning), semantic_equivalence_score, source_component, and when present impact_estimate and passthrough keys such as table, column, kind, detail, before, after, confidence, signals. Status codes: 200, 500.

bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/drift/reports?integration_id=erp-prod&limit=50"

GET /api/v1/drift/stats ​

Drift statistics. Permission: read. Response 200: total_changes, semantic_drifts, drift_by_type (map), avg_drift_score. Status codes: 200, 500.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/drift/stats

GET /api/v1/drift/value-drift ​

Value-drift findings (the value_drift dimension). Findings for a muted field are hidden immediately, and reappear when the field is unmuted. Permission: read.

NameInTypeRequiredDefaultDescription
integration_idquerystringno
include_mutedquerybooleannofalseInclude findings for muted fields
limitqueryintegerno1001 to 500

Response 200: findings and count. Each finding: id, integration_id, table, column, kind, detail, severity, before, after, detected_at, muted, impact_estimate. The limit applies before the mute filter, so a page can contain fewer than limit items.

bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/drift/value-drift?integration_id=erp-prod"

POST /api/v1/drift/value-drift/mute ​

Mutes a baselined field. Permission: write.

NameTypeRequiredDescription
integration_idstringyes
tablestringyes
columnstringyes

Response 200: integration_id, table, column, muted (true), muted_at. Status codes: 200, 404 (no baseline for that field).

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"integration_id":"erp-prod","table":"orders","column":"status"}' \
  https://mmune.example.com/api/v1/drift/value-drift/mute

POST /api/v1/drift/value-drift/unmute ​

Reverses a mute. Permission: write. Body as for mute. Response 200: same shape with muted false and muted_at null. Status codes: 200, 404.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"integration_id":"erp-prod","table":"orders","column":"status"}' \
  https://mmune.example.com/api/v1/drift/value-drift/unmute

Healing router ​

Base path /api/v1/healing. Healing watches connection health, detects broken connections and field mappings, and proposes or applies remappings. In alert mode it only proposes. In autonomous mode it applies the top suggestion when confidence meets the threshold.

GET /api/v1/healing/config ​

Permission: read. Response 200: mode (alert or autonomous) and auto_apply_threshold (number from 0 to 1). The effective value is the current setting, otherwise the saved configuration, otherwise the environment variables HEALING_MODE and HEALING_AUTO_APPLY_THRESHOLD, otherwise the defaults alert and 0.9.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/healing/config

PUT /api/v1/healing/config ​

Changes the estate-wide healing mode and threshold. Permission: admin.

NameTypeRequiredDescription
modestringnoalert or autonomous
auto_apply_thresholdnumberno0 to 1 inclusive

If you send only the threshold, the mode is left as is. The change is saved and survives restarts. Read the value back with GET to confirm it. Response 200: the same two fields as GET. Status codes: 200, 400 (mode must be 'alert' or 'autonomous'), 403, 422 (threshold out of range).

bash
curl -s -X PUT -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"mode":"autonomous","auto_apply_threshold":0.92}' https://mmune.example.com/api/v1/healing/config

GET /api/v1/healing/connection-health ​

Connection health for one integration, or overall statistics. Permission: read.

NameInTypeRequiredDescription
integration_idquerystringnoIf set, returns that integration's latest record

With integration_id, response 200: integration_id, provider, health, error_type, error_message, response_time_ms, checked_at. Without it, response 200: statistics (total_connections, broken_connections, health_distribution as a map of health value to count, broken_by_error_type as a map of every error type to count) and a message. Status codes: 200, 404 (no record for that integration), 500.

bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/healing/connection-health?integration_id=erp-prod"

POST /api/v1/healing/connection-health/check ​

Runs an active check. If the result is broken it also raises an alert event. Permission: write. The body fields are top-level JSON keys.

NameTypeRequiredDefaultDescription
integration_idstringyes
providerstringyes
test_connectionbooleannotrueSet false to skip the live connection test

Response 200: integration_id, provider, health, error_type, error_message, response_time_ms, checked_at. Status codes: 200, 422, 500.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"integration_id":"erp-prod","provider":"postgres"}' https://mmune.example.com/api/v1/healing/connection-health/check

GET /api/v1/healing/broken-connections ​

Broken connections, combining connection health history and live watchdog sessions. Permission: read. Query: integration_id (optional filter).

Response 200: broken_connections and count. Each entry:

FieldTypeDescription
connection_idstring
integration_idstring
providerstring
error_typestringOne of the error types
error_messagestring
first_detectedstring
last_detectedstring
failure_countinteger
is_permanentboolean
broken_mappingsarray
suggested_fixesarray
retry_attemptsintegerAutomatic recovery attempts so far
retry_schedule_lengthintegerNumber of scheduled retries, 3 (backoff of 60, 300 and 900 seconds)
next_retry_atstring or null
retry_exhaustedboolean
bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/healing/broken-connections

POST /api/v1/healing/suggest-remapping ​

Returns remapping suggestions for a field that no longer exists. Permission: write. Important side effect: when the healing mode is autonomous, the top suggestion has confidence at or above the threshold, and containment is not holding the integration, the same call applies that suggestion to your saved mappings (change attributed to healing_autonomous). If containment holds the integration because of an open incident, the suggestion is returned with its reasoning annotated and nothing is applied. The top 3 suggestions are also published as events.

NameTypeRequiredDescription
integration_idstringyes
providerstringyes
broken_fieldstringyesField that was not found
available_fieldsarray of stringsyesCandidate target fields

Response 200: an array of suggestions, each with suggestion_id, source_field, suggested_target_field, confidence, strategy, reasoning, business_intent, alternative_suggestions, created_at. Status codes: 200, 422, 500.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"integration_id":"crm-prod","provider":"salesforce","broken_field":"cust_email","available_fields":["email","email_address","phone"]}' \
  https://mmune.example.com/api/v1/healing/suggest-remapping

POST /api/v1/healing/apply-suggestion ​

Applies a suggestion, as an operator would. Permission: write. Saved mappings are changed only when source_field, suggested_target_field and provider are all present, with the change attributed to healing_apply. In every case an event is published.

NameTypeRequiredDescription
suggestion_idstringyes
integration_idstringyes
source_fieldstringnoBroken source field
suggested_target_fieldstringno
providerstringnoVendor name, for example salesforce

Response 200: status (success), message (Mapping updated or Event emitted (mapping details not provided)), suggestion_id, mapping_updated (boolean). A successful remap also triggers containment recovery for the integration. Status codes: 200, 422, 500.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"suggestion_id":"<SUGGESTION_ID>","integration_id":"crm-prod","source_field":"cust_email","suggested_target_field":"email","provider":"salesforce"}' \
  https://mmune.example.com/api/v1/healing/apply-suggestion

GET /api/v1/healing/health-history/ ​

Health records over time. Permission: read. Query: limit (integer, default 100). Response 200: integration_id, history (array of health, error_type, error_message, response_time_ms, checked_at) and count.

bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/healing/health-history/erp-prod?limit=20"

GET /api/v1/healing/statistics ​

Permission: read. Response 200: connection_health (same object as statistics under connection-health) and message.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/healing/statistics

Containment router ​

Base path /api/v1/containment. When an incident is severe enough, containment places quarantine marks on lineage nodes so that downstream consumers and autonomous healing hold off until the cause is resolved.

GET /api/v1/containment/config ​

Permission: read. Response 200: mode (off, observe or active), severity_threshold (string), max_depth (integer). Defaults come from MMUNE_CONTAINMENT_MODE (default off), MMUNE_CONTAINMENT_SEVERITY_THRESHOLD (default critical) and MMUNE_CONTAINMENT_MAX_DEPTH (default 10), unless a stored value exists.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/containment/config

PUT /api/v1/containment/config ​

Estate-wide change. Permission: admin.

NameTypeRequiredDescription
modestringnooff, observe or active
severity_thresholdstringnoA severity value. Lower-cased, not checked against the list
max_depthintegerno1 to 200

Response 200: the three config fields. Status codes: 200, 400 (mode must be 'off', 'observe', or 'active'), 403, 422.

bash
curl -s -X PUT -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"mode":"observe","severity_threshold":"critical","max_depth":10}' https://mmune.example.com/api/v1/containment/config

GET /api/v1/containment/marks ​

Active quarantine marks. Permission: read. Query (both optional): incident_id, node_key. Response 200: {"marks": [...]} where each mark has id, node_key, incident_id, role, reason, depth, created_at, observe (true when the mark was placed in observe mode and does not enforce).

bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/containment/marks?incident_id=<INCIDENT_ID>"

POST /api/v1/containment/marks/{mark_id}/lift ​

Lifts one active mark. Permission: write.

NameInTypeRequiredDescription
mark_idpathstringyes
lift_reasonbodystringyesAt least 1 character

Response 200: id, node_key, incident_id, lifted_at, lifted_by (the name of the actor who lifted the mark), lift_reason. Status codes: 200, 404 (Active mark not found), 422.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"lift_reason":"Upstream fix verified"}' https://mmune.example.com/api/v1/containment/marks/<MARK_ID>/lift

POST /api/v1/containment/incidents/{incident_id}/revalidate ​

Runs the recovery re-validation sweep for an incident. Marks that pass are lifted and the incident can resolve automatically. Permission: write. No body.

Response 200: incident_id, verified, total, still_failing, truncated, lifted_node_keys, failed_node_keys, incident_resolved, summary, notification_sent. Status codes: 200, 404 (Incident not found).

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/containment/incidents/<INCIDENT_ID>/revalidate

Incidents router ​

Base path /api/v1/incidents. An incident groups related detection events (its members) that share an origin. Static paths such as /owners, /sla and /rca-config are matched before /{incident_id}, so those words are not valid incident ids.

GET /api/v1/incidents ​

Lists incidents, most recently active first. Permission: read.

NameInTypeRequiredDefaultDescription
pagequeryintegerno11-based
page_sizequeryintegerno10Capped at 100
statusquerystringnoopen, acknowledged or resolved. Not validated; an unknown value returns an empty list
severityquerystringnoExact match
integration_idquerystringnoMatches the incident's origin integration

Response 200: incidents (array of incident summaries), total, page, page_size.

Incident summary fields:

FieldTypeDescription
idstring
statusstringopen, acknowledged, resolved
severitystring
titlestring
origin_integration_idstring
origin_object_keystring or null
origin_node_keystring or nullLineage node key
dimensionsarray of stringsDetection dimensions of the member events
member_countinteger
first_detected_atstring
last_activity_atstring
ticket_provider, ticket_id, ticket_url, ticket_statusstring or nullLinked external ticket
ticket_last_synced_atstring or null
ack_deadline_at, resolve_deadline_atstring or nullSLA deadlines
sla_statestringon_track by default
ownerobject or nullResolved owner, fields as in the owners section
outcomestring or nullOperator label
outcome_labeled_at, outcome_labeled_by, outcome_notestring or null
bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/incidents?status=open&severity=critical&page=1&page_size=25"

GET /api/v1/incidents/ ​

Incident detail. Permission: read. Response 200: all summary fields plus:

FieldTypeDescription
member_event_idsarray of integers
member_eventsarrayOldest first. Each has id, dimension, integration_id, object_key, severity, detected_at, source_component, payload
impact_estimateobject or null
acknowledged_at, acknowledged_bystring or null
resolved_at, resolved_bystring or null
resolutionstring or nullFree text
resolution_modestring or nullmanual or auto

Status codes: 200, 404 (Incident not found).

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/incidents/<INCIDENT_ID>

POST /api/v1/incidents/{incident_id}/acknowledge ​

Sets status to acknowledged and records who and when (the caller's email, or user id if no email). No body. Permission: write. Response 200: incident detail. Status codes: 200, 400 (Cannot acknowledge a resolved incident), 404. Acknowledging an already acknowledged incident succeeds and overwrites the acknowledgement time and user.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/incidents/<INCIDENT_ID>/acknowledge

POST /api/v1/incidents/{incident_id}/resolve ​

Resolves an incident with resolution_mode manual. If a ticket is linked, mmune adds a comment to it and, when the ticketing configuration has close_ticket_on_resolve true, tries to move it to done. Permission: write.

NameTypeRequiredDescription
resolutionstringyesAt least 1 character

Response 200: incident detail. Status codes: 200, 400 (Incident already resolved), 404, 422.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"resolution":"Column restored upstream, mapping verified"}' \
  https://mmune.example.com/api/v1/incidents/<INCIDENT_ID>/resolve

POST /api/v1/incidents/{incident_id}/outcome ​

Labels the incident for later accuracy measurement. It writes only to the incident, never to the detection ground-truth labels, and changes no alert thresholds. Permission: write.

NameTypeRequiredDescription
outcomestringyestrue_positive, false_positive or inconclusive
notestringnoUp to 2000 characters. Whitespace-only becomes null

Response 200: incident detail. Status codes: 200, 404, 422 (invalid outcome or note too long).

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"outcome":"true_positive","note":"Confirmed with the data owner"}' \
  https://mmune.example.com/api/v1/incidents/<INCIDENT_ID>/outcome

GET /api/v1/incidents/{incident_id}/rca ​

Latest root-cause analysis. Permission: read. Response 200: incident_id, generated_at, generated_by, narrative, candidates (array), evidence_bundle (object), runbook (id, generated_at, generated_by, header, steps, footer). Status codes: 200, 404 (none generated yet).

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/incidents/<INCIDENT_ID>/rca

POST /api/v1/incidents/{incident_id}/rca ​

Generates (or regenerates) the analysis and runbook. No body. Permission: write. Response 200: same shape as the GET. Status codes: 200, 404 (unknown incident), 409 (a generation is already running for this incident).

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/incidents/<INCIDENT_ID>/rca

GET /api/v1/incidents/{incident_id}/runbook.md ​

The latest runbook rendered as Markdown (text/markdown). Permission: read. Status codes: 200, 404 (no analysis generated yet).

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/incidents/<INCIDENT_ID>/runbook.md

GET /api/v1/incidents/rca-config ​

Permission: read. Response 200: mode (off, on_demand or auto_on_open) and allow_heuristic_auto (boolean).

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/incidents/rca-config

PUT /api/v1/incidents/rca-config ​

Permission: write. Body, both optional: mode (off, on_demand, auto_on_open) and allow_heuristic_auto (boolean). Response 200: both fields. Status codes: 200, 400 (invalid mode).

bash
curl -s -X PUT -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"mode":"on_demand"}' https://mmune.example.com/api/v1/incidents/rca-config

GET /api/v1/incidents/sla ​

The SLA policy. Permission: read. Response 200: ack_minutes and resolve_minutes, each a map keyed by critical, high, medium, low. Defaults: acknowledge 15, 60, 240 and 1440 minutes; resolve 240, 1440, 4320 and 10080 minutes (in that severity order).

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/incidents/sla

PUT /api/v1/incidents/sla ​

Replaces the policy. Permission: write. Both maps are required and each must contain all four severity keys, otherwise 400 (Missing severity bucket '<key>' ...).

NameTypeRequiredDescription
ack_minutesobject of integersyesKeys critical, high, medium, low
resolve_minutesobject of integersyesSame keys

Response 200: the stored policy.

bash
curl -s -X PUT -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"ack_minutes":{"critical":10,"high":60,"medium":240,"low":1440},"resolve_minutes":{"critical":240,"high":1440,"medium":4320,"low":10080}}' \
  https://mmune.example.com/api/v1/incidents/sla

GET /api/v1/incidents/owners ​

The ownership registry used to route incidents to a team. Permission: read. Response 200: owners (sorted by priority then id), count, known_channels (the notification channel list). Owner fields: id, name, team, contact, notify_channels, integration_id_pattern, object_key_pattern, priority.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/incidents/owners

POST /api/v1/incidents/owners ​

Creates an owner. Permission: write.

NameTypeRequiredDefaultDescription
idstringnoderived from nameSame derivation as alert rule ids, including stripping leading and trailing _
namestringyes
teamstringno""
contactstringno""
notify_channelsarray of stringsno[]Must be known channel names
integration_id_patternstringnonullPattern for the origin integration id
object_key_patternstringnonullPattern for the object key
priorityintegerno100Lower values are evaluated first

Response 201: the owner. Status codes: 201, 400 (no id derivable, unknown channel), 409 (id exists).

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"name":"Data Platform","team":"platform","contact":"data-platform@example.com","notify_channels":["slack"],"integration_id_pattern":"erp-*"}' \
  https://mmune.example.com/api/v1/incidents/owners

PUT /api/v1/incidents/owners/ ​

Partial update; only fields present in the body change. Permission: write. Body fields are the same as create, minus id, all optional. Response 200: the owner. Status codes: 200, 400, 404.

bash
curl -s -X PUT -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"contact":"oncall@example.com"}' https://mmune.example.com/api/v1/incidents/owners/data_platform

DELETE /api/v1/incidents/owners/ ​

Permission: write. Response 200: {"deleted": "<owner_id>"}. Status codes: 200, 404.

bash
curl -s -X DELETE -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/incidents/owners/data_platform

Events router ​

Base path /api/v1/events.

GET /api/v1/events/cursor ​

Returns three counters. A client polls roughly every 10 seconds and refetches a view only when the counter for that view's domain has changed since the last poll, which avoids refetching unchanged lists. Permission: read. The counters carry no data other than opaque integers.

Response 200: findings, lineage, health, each an integer. Treat any change in a counter as a cue to refetch.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/events/cursor

Notifications router ​

Base path /api/v1/notifications. The set of channels is fixed (see the enumeration list). You configure whether each is enabled and its endpoint. For outbound payload formats, see Webhooks and events.

Where the endpoint_url field points depends on the channel:

ChannelWhat endpoint_url holdsEnvironment fallback
slackIncoming webhook URLMMUNE_SLACK_WEBHOOK_URL
webhookGeneric webhook URLMMUNE_ALERT_WEBHOOK_URL
pagerdutyEvents API v2 routing key (not a URL)MMUNE_PAGERDUTY_ROUTING_KEY
teamsWebhook URLMMUNE_TEAMS_WEBHOOK_URL
email_digestSMTP host. Recipients and secrets come from environment or stored credentialsMMUNE_SMTP_HOST
servicenowNot used. The ServiceNow channel uses the ticketing configuration for instance URL and credentialsnone

Resolution order for each channel: a saved database row, then the environment, then the default. Every channel is enabled by default except servicenow, which is off until enabled so that configuring ticketing does not start creating incidents unannounced. The environment variable MMUNE_NOTIFICATION_<CHANNEL>_ENABLED can also set the enabled flag.

GET /api/v1/notifications/channels ​

Permission: read. Response 200: {"channels": [...]} with one entry per known channel. Each entry:

FieldTypeDescription
channelstring
enabledboolean
configuredbooleanTrue only if the channel is enabled and could actually send (for example it has an endpoint)
endpoint_url_setbooleanWhether an endpoint is stored or inherited. The value itself is never returned
sourcestringdb, env or default
circuit_breakerobjectstate (closed, open, half_open) and consecutive_failures
bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/notifications/channels

GET /api/v1/notifications/channels/ ​

One channel entry, same shape. Permission: read. For a name that is not a known channel the call returns a default entry rather than 404.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/notifications/channels/slack

PUT /api/v1/notifications/channels/ ​

Saves the channel configuration to the database. This is a full replace of the two fields: omitting endpoint_url stores null. Permission: write.

NameTypeRequiredDefaultDescription
enabledbooleannotrue
endpoint_urlstring or nullnonullMeaning depends on the channel, see the table above

Response 200: {"message": "Channel configuration saved", "config": <channel entry>}. Status codes: 200, 400 (unknown channel; the message lists the known names), 422.

bash
curl -s -X PUT -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"enabled":true,"endpoint_url":"https://hooks.slack.example.com/services/<WEBHOOK_PATH>"}' \
  https://mmune.example.com/api/v1/notifications/channels/slack

POST /api/v1/notifications/channels/{channel}/test ​

Sends a test notification through one channel. Permission: write.

NameTypeRequiredDefaultDescription
messagestringnommune notification testAt least 1 character
severitystringnolow

Response 200: channel, status (a delivery status), attempts, error, ok. A channel that is registered but not configured returns 200 with a skipped status rather than an error. Status codes: 200, 400 (No channel named '<channel>'), 422.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"message":"Test from the API","severity":"low"}' https://mmune.example.com/api/v1/notifications/channels/slack/test

GET /api/v1/notifications/deliveries ​

Recent delivery attempts. Permission: read.

NameInTypeRequiredDefaultDescription
limitqueryintegerno50Capped at 200
channelquerystringno
statusquerystringnoA delivery status
alert_idquerystringno

Response 200: {"deliveries": [...]}, each with id, alert_id, channel, status, attempts, error, created_at.

bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/notifications/deliveries?channel=slack&status=failed&limit=20"

Ticketing router ​

Base path /api/v1/ticketing. One ticketing system is configured per deployment. Saving a new configuration replaces the previous one.

GET /api/v1/ticketing/providers ​

Permission: read. Response 200: {"providers": [{"provider": "jira", "label": "Jira", "description": "..."}, {"provider": "servicenow", "label": "ServiceNow", "description": "..."}]}. These are the only two systems supported.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/ticketing/providers

GET /api/v1/ticketing/config ​

Permission: read. Response 200: {"configured": false, "config": null} when nothing is saved, otherwise {"configured": true, "config": {...}} with provider, base_url, default_project, default_table, has_credentials (boolean) and close_ticket_on_resolve. Credentials are never returned.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/ticketing/config

POST /api/v1/ticketing/config ​

Saves the configuration. Permission: write.

NameTypeRequiredDefaultDescription
providerstringyesjira or servicenow
base_urlstringyesAt least 1 character. Jira site URL or ServiceNow instance URL
credentialsobjectno{}Keys depend on the provider, see below
default_projectstringnonullJira project key
default_tablestringnonullServiceNow table. Falls back to incident
close_ticket_on_resolvebooleannofalseWhen an incident is resolved, also move the linked ticket to done

Credential keys:

ProviderKeyNotes
jiraemail (or username)Account email
jiraapi_token (or token)API token
jiraproject_keyUsed only if default_project is not set
servicenowusername, passwordBasic authentication
servicenowclient_id, client_secret, refresh_tokenOAuth

Jira issues are created with type Task, labels mmune-<alert_source> (or mmune-alert), and priority mapped critical to Highest, high to High, medium to Medium, low to Low. ServiceNow records are created in the configured table with urgency 1 for critical and high, 2 for medium, 3 for low, and impact 2. When close_ticket_on_resolve is true, Jira closes only if exactly one transition leads to a "done" status category. ServiceNow sets state 6.

Response 200: {"message": "Ticketing configuration saved", "config": {...}} with the same fields as GET. Status codes: 200, 422.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"provider":"jira","base_url":"https://example.atlassian.net","credentials":{"email":"svc-mmune@example.com","api_token":"<JIRA_API_TOKEN>"},"default_project":"OPS","close_ticket_on_resolve":false}' \
  https://mmune.example.com/api/v1/ticketing/config

DELETE /api/v1/ticketing/config ​

Removes the stored configuration. Permission: write. Response 200: {"message": "Ticketing configuration removed"}.

bash
curl -s -X DELETE -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/ticketing/config

POST /api/v1/ticketing/test ​

Tests the connection with the saved configuration. Jira uses its health check; ServiceNow reads one sys_user row. The result is recorded against the stored configuration. Permission: write. No body. Response 200: {"success": true} or {"success": false}. Jira returns false when its health check is not healthy. Status codes: 200, 400 (Ticketing connection test failed: ..., including when nothing is configured).

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/ticketing/test

POST /api/v1/ticketing/tickets ​

Creates a ticket in the configured system. Permission: write.

NameTypeRequiredDefaultDescription
titlestringyesAt least 1 character
descriptionstringyesAt least 1 character
prioritystringnomediumcritical, high, medium, low
alert_idstringnonull
alert_sourcestringnonullUsed in the Jira label
metadataobjectno{}Passed as Jira custom fields, or stored as text in u_mmune_metadata on ServiceNow
incident_idstringnonullIf set, the incident's root-cause analysis is appended to the description and the ticket is linked to the incident (ticket_provider, ticket_id, ticket_url)

Response 200: ticket_id, ticket_url, provider. Status codes: 200, 400 (Failed to create ticket: ...), 404 (no ticketing configured, or the linked incident was not found), 422.

bash
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"title":"Schema drift on orders","description":"Column status changed type","priority":"high","incident_id":"<INCIDENT_ID>"}' \
  https://mmune.example.com/api/v1/ticketing/tickets

Stats router ​

GET /api/v1/stats/history ​

Daily history of two real counts: registered integrations (api_counts) and currently broken connections (issue_counts). The call also writes or updates today's point first. History starts at first deployment and is never back-filled, so a new install returns few points. Permission: read (the call records today's point but is a GET).

NameInTypeRequiredDefaultDescription
rangequerystringno7d7d, 30d or 1y. Any other value is treated as 7d

Response 200: api_counts and issue_counts, each an array of {"date": "YYYY-MM-DD", "value": <integer>}.

bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/stats/history?range=30d"

Dashboard router ​

Base path /api/v1/dashboard. These endpoints report on mmune's own agents.

GET /api/v1/dashboard/overview ​

Permission: read. Response 200: system_metrics (total_agents, active_agents, total_executions, successful_executions, failed_executions, average_success_rate, total_errors, uptime_percentage), agent_health (array of agent health objects) and timestamp.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/dashboard/overview

GET /api/v1/dashboard/agents/{agent_type}/health ​

Permission: read. Query: days (integer, 1 to 90, default 7). Response 200: agent_type, status (healthy, degraded, unhealthy or unknown), last_execution, success_rate, average_duration, error_count, warning_count.

bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/dashboard/agents/discovery/health?days=14"

GET /api/v1/dashboard/errors ​

Failed agent executions in the window. Permission: read.

NameInTypeRequiredDefaultDescription
agent_typequerystringnoFilter
daysqueryintegerno71 to 90
limitqueryintegerno501 to 200

Response 200: an array of errors, each with execution_id, agent_type, timestamp, errors (array of messages), context_id.

bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/dashboard/errors?days=3&limit=20"

Permission: read. Query: agent_type (optional), days (1 to 365, default 30). Response 200: agent_type (all when unfiltered), period_days, trends (array of date, total_executions, success_rate as a percentage, average_duration in seconds).

bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/dashboard/performance/trends?days=30"

GET /api/v1/dashboard/learning/progress ​

Accuracy of the learning system and whether enough labelled data exists to train. Permission: read. Query: project_id (optional; without it, counts use the global pool).

Response 200: accuracy_metrics (object), project_id, and volume_gate. volume_gate has mapping_ready (boolean), incident_ready (boolean), counts (object) and deficits (object of integers), or is null if the gate could not be computed. The readiness thresholds are 100 usable mapping feedback items of which 30 are corrections, and 50 labelled incidents (true or false positive) across at least 2 dimensions.

bash
curl -s -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/dashboard/learning/progress

Logs router ​

Base path /api/v1.

GET /api/v1/logs ​

Global agent logs, returned oldest first within the window. Permission: read.

NameInTypeRequiredDefaultDescription
limitqueryintegerno100Maximum 1000 (larger returns 422). The newest limit entries are selected
after_idqueryintegerno0Return only entries with an id greater than this. Use the last id you saw to poll

Response 200: an array of {id, timestamp, level, agent_type, message, job_id}.

bash
curl -s -H "Authorization: Bearer $TOKEN" "https://mmune.example.com/api/v1/logs?limit=200&after_id=0"

DELETE /api/v1/logs ​

Clears all agent logs. Permission: admin. Response 200: {"status": "cleared"}. Status codes: 200, 403.

bash
curl -s -X DELETE -H "Authorization: Bearer $TOKEN" https://mmune.example.com/api/v1/logs

Worked examples ​

Each example is a complete script. The helper is defined once here and reused.

Python helper:

python
import os
import requests

BASE = "https://mmune.example.com/api/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['MMUNE_TOKEN']}"}


def api(method, path, body=None, **params):
    response = requests.request(
        method, f"{BASE}{path}", headers=HEADERS, json=body, params=params or None, timeout=30
    )
    response.raise_for_status()
    return response.json()

TypeScript helper (Node 18 or later, or a browser):

typescript
const BASE = "https://mmune.example.com/api/v1";
const TOKEN = process.env.MMUNE_TOKEN;

async function api<T = any>(
  method: string,
  path: string,
  body?: unknown,
  params?: Record<string, string | number | boolean>,
): Promise<T> {
  const query = params
    ? "?" + new URLSearchParams(Object.entries(params).map(([k, v]) => [k, String(v)])).toString()
    : "";
  const response = await fetch(`${BASE}${path}${query}`, {
    method,
    headers: { Authorization: `Bearer ${TOKEN}`, "Content-Type": "application/json" },
    body: body === undefined ? undefined : JSON.stringify(body),
  });
  if (!response.ok) {
    throw new Error(`${method} ${path} failed: ${response.status} ${await response.text()}`);
  }
  return (await response.json()) as T;
}

1. Poll overall health and list alerts ​

Read overall_status from /health/detailed. For alerts, poll the cheap cursor and refetch drift findings and open incidents only when the findings counter moves.

python
import time

last_findings = None
while True:
    health = api("GET", "/health/detailed")
    print("overall:", health["overall_status"], "| llm reachable:", health["llm"]["reachable"])

    cursor = api("GET", "/events/cursor")
    if cursor["findings"] != last_findings:
        last_findings = cursor["findings"]
        drift = api("GET", "/drift/reports", limit=50)["reports"]
        open_incidents = api("GET", "/incidents", status="open", page_size=50)["incidents"]
        for item in drift:
            print(item["severity"], item["integration_id"], item["drift_type"], item["object_name"])
        print(len(open_incidents), "open incidents")
    time.sleep(10)
typescript
let lastFindings: number | undefined;
while (true) {
  const health = await api("GET", "/health/detailed");
  console.log("overall:", health.overall_status, "| llm reachable:", health.llm.reachable);

  const cursor = await api<{ findings: number }>("GET", "/events/cursor");
  if (cursor.findings !== lastFindings) {
    lastFindings = cursor.findings;
    const drift = (await api("GET", "/drift/reports", undefined, { limit: 50 })).reports;
    const open = (await api("GET", "/incidents", undefined, { status: "open", page_size: 50 })).incidents;
    for (const item of drift) console.log(item.severity, item.integration_id, item.drift_type, item.object_name);
    console.log(open.length, "open incidents");
  }
  await new Promise((resolve) => setTimeout(resolve, 10_000));
}

2. List, acknowledge and resolve incidents ​

python
page = api("GET", "/incidents", status="open", severity="critical", page=1, page_size=25)
for incident in page["incidents"]:
    print(incident["id"], incident["title"], incident["sla_state"])

incident_id = page["incidents"][0]["id"]
api("POST", f"/incidents/{incident_id}/acknowledge")
detail = api("GET", f"/incidents/{incident_id}")
print(detail["status"], detail["acknowledged_by"], len(detail["member_events"]), "member events")

api("POST", f"/incidents/{incident_id}/outcome", {"outcome": "true_positive", "note": "Confirmed with data owner"})
resolved = api("POST", f"/incidents/{incident_id}/resolve", {"resolution": "Mapping restored and verified"})
print(resolved["status"], resolved["resolution_mode"])  # resolved manual
typescript
const page = await api("GET", "/incidents", undefined, { status: "open", severity: "critical", page: 1, page_size: 25 });
for (const incident of page.incidents) console.log(incident.id, incident.title, incident.sla_state);

const incidentId: string = page.incidents[0].id;
await api("POST", `/incidents/${incidentId}/acknowledge`);
const detail = await api("GET", `/incidents/${incidentId}`);
console.log(detail.status, detail.acknowledged_by, detail.member_events.length, "member events");

await api("POST", `/incidents/${incidentId}/outcome`, { outcome: "true_positive", note: "Confirmed with data owner" });
const resolved = await api("POST", `/incidents/${incidentId}/resolve`, { resolution: "Mapping restored and verified" });
console.log(resolved.status, resolved.resolution_mode); // resolved manual

3. Switch healing mode and read what healing did ​

Switching the mode needs an admin token. Reading is read. The reads below show broken connections, connection health history and the containment marks that resulted.

python
print(api("GET", "/healing/config"))  # {'mode': 'alert', 'auto_apply_threshold': 0.9}
api("PUT", "/healing/config", {"mode": "autonomous", "auto_apply_threshold": 0.92})

for broken in api("GET", "/healing/broken-connections")["broken_connections"]:
    print(broken["integration_id"], broken["error_type"], "retries:", broken["retry_attempts"], "/", broken["retry_schedule_length"])
    history = api("GET", f"/healing/health-history/{broken['integration_id']}", limit=10)["history"]
    print("  last health:", history[-1]["health"] if history else "none")

print(api("GET", "/containment/marks")["marks"])  # quarantine marks created by containment

api("PUT", "/healing/config", {"mode": "alert"})  # back to suggest-only
typescript
console.log(await api("GET", "/healing/config")); // { mode: 'alert', auto_apply_threshold: 0.9 }
await api("PUT", "/healing/config", { mode: "autonomous", auto_apply_threshold: 0.92 });

const broken = (await api("GET", "/healing/broken-connections")).broken_connections;
for (const b of broken) {
  console.log(b.integration_id, b.error_type, "retries:", b.retry_attempts, "/", b.retry_schedule_length);
  const history = (await api("GET", `/healing/health-history/${b.integration_id}`, undefined, { limit: 10 })).history;
  console.log("  last health:", history.length ? history[history.length - 1].health : "none");
}

console.log((await api("GET", "/containment/marks")).marks);

await api("PUT", "/healing/config", { mode: "alert" });

4. Configure a ticketing integration and create a ticket ​

Jira and ServiceNow are the only supported systems. This example configures Jira, tests it, and files a ticket linked to an incident.

python
print([p["provider"] for p in api("GET", "/ticketing/providers")["providers"]])  # ['jira', 'servicenow']

api("POST", "/ticketing/config", {
    "provider": "jira",
    "base_url": "https://example.atlassian.net",
    "credentials": {"email": "svc-mmune@example.com", "api_token": os.environ["JIRA_API_TOKEN"]},
    "default_project": "OPS",
    "close_ticket_on_resolve": True,
})
print(api("GET", "/ticketing/config"))  # has_credentials is true, the token is not returned

assert api("POST", "/ticketing/test")["success"] is True

ticket = api("POST", "/ticketing/tickets", {
    "title": "Schema drift on orders",
    "description": "Column status changed type",
    "priority": "high",
    "incident_id": incident_id,
})
print(ticket["ticket_id"], ticket["ticket_url"])
typescript
console.log((await api("GET", "/ticketing/providers")).providers.map((p: any) => p.provider));

await api("POST", "/ticketing/config", {
  provider: "jira",
  base_url: "https://example.atlassian.net",
  credentials: { email: "svc-mmune@example.com", api_token: process.env.JIRA_API_TOKEN },
  default_project: "OPS",
  close_ticket_on_resolve: true,
});
console.log(await api("GET", "/ticketing/config"));

const test = await api("POST", "/ticketing/test");
if (!test.success) throw new Error("Ticketing connection test did not succeed");

const ticket = await api("POST", "/ticketing/tickets", {
  title: "Schema drift on orders",
  description: "Column status changed type",
  priority: "high",
  incident_id: incidentId,
});
console.log(ticket.ticket_id, ticket.ticket_url);

For ServiceNow, send "provider": "servicenow", the instance URL as base_url, credentials with username and password (or the OAuth keys), and optionally default_table.

5. Configure notification channels ​

python
for channel in api("GET", "/notifications/channels")["channels"]:
    print(channel["channel"], "enabled" if channel["enabled"] else "disabled", "configured" if channel["configured"] else "not configured", channel["source"])

api("PUT", "/notifications/channels/slack", {
    "enabled": True,
    "endpoint_url": os.environ["SLACK_WEBHOOK_URL"],
})

result = api("POST", "/notifications/channels/slack/test", {"message": "Hello from the API", "severity": "low"})
print(result["status"], result["ok"], result["error"])

failed = api("GET", "/notifications/deliveries", channel="slack", status="failed", limit=20)["deliveries"]
print(len(failed), "failed deliveries")
typescript
const { channels } = await api("GET", "/notifications/channels");
for (const c of channels) console.log(c.channel, c.enabled ? "enabled" : "disabled", c.configured ? "configured" : "not configured", c.source);

await api("PUT", "/notifications/channels/slack", {
  enabled: true,
  endpoint_url: process.env.SLACK_WEBHOOK_URL,
});

const result = await api("POST", "/notifications/channels/slack/test", { message: "Hello from the API", severity: "low" });
console.log(result.status, result.ok, result.error);

const failed = (await api("GET", "/notifications/deliveries", undefined, { channel: "slack", status: "failed", limit: 20 })).deliveries;
console.log(failed.length, "failed deliveries");

To route specific findings to a channel, create an alert rule with a route_channels action (see POST /api/v1/monitoring/alert-rules), or give an incident owner notify_channels.