Appearance
mmune WebSocket and streaming endpoints
What this page covers
This page documents the endpoints that push data to the client instead of answering once: one WebSocket endpoint and one Server-Sent Events (SSE) stream. For each it gives the path, how to authenticate, the messages in both directions, how the connection ends, and what a client should do to reconnect. It also covers the two small HTTP route groups that go with long-running agent executions (state checkpoints and cancellation), because a client that watches an execution usually needs them too.
Prerequisites: a running mmune deployment, a valid access token for an account with the read permission, and the conventions from API overview. Login, token lifetime, refresh and the error envelopes are described there and are not repeated here. The examples use https://mmune.example.com as the host (so wss://mmune.example.com for WebSockets) and <TOKEN> for the access token. The Python examples use the websockets and requests packages. The TypeScript examples use the standard WebSocket (a browser, or Node 22 and later) and fetch.
Endpoint summary
| Endpoint | Transport | Direction | Purpose |
|---|---|---|---|
/api/v1/monitoring/ws | WebSocket | server to client, once per second | Live dashboard snapshot |
POST /api/v1/analysis/chat/stream | SSE over a POST response | server to client | Streamed chat answer |
The rate limits in API overview apply to HTTP requests. The chat stream is an ordinary HTTP request and counts as one.
Authentication for WebSockets
The same checks that protect the HTTP API protect WebSocket handshakes. A handshake needs a valid Bearer JWT and the read permission, so a read-only account is enough. See API overview for how to obtain a token.
The token can be sent in one of two ways.
| Method | How | Who uses it |
|---|---|---|
| Subprotocol | Offer exactly two subprotocols, in this order: mmune.bearer and then the token. On the wire this is Sec-WebSocket-Protocol: mmune.bearer, <TOKEN>. | Browsers, because the browser WebSocket API cannot set an Authorization header. The mmune web console does this. |
| Header | Authorization: Bearer <TOKEN> on the upgrade request. | Server-side clients that can set headers. |
If an Authorization header is present it is used and the subprotocol is not consulted. If the header is present but not of the form Bearer <token>, the handshake is refused; there is no fall back to the subprotocol. The subprotocol form is only recognised when exactly two protocols are offered and the first is mmune.bearer. When it is used, the server answers the handshake by selecting mmune.bearer, which browsers require before they will keep the socket open.
A token in a query string (?token=...) is not accepted. Do not put tokens in URLs.
The token is checked at the handshake. Open each new connection with a valid token, and refresh it first if the old one is near expiry (see the refresh section of API overview).
Handshake failures
When the token is missing, malformed, expired, revoked, or lacks read, the server refuses the upgrade with HTTP 403 and does not switch protocols. Unlike the HTTP API, the reason is not returned in a body: the client only sees a failed handshake. To find the reason, call GET /api/v1/auth/validate-token with the same token and read the 401 messages listed in API overview. In the browser the failure surfaces as an error event followed by a close event with code 1006, with no way to read the HTTP status.
With the Python websockets package a refused handshake raises InvalidStatus (or InvalidStatusCode in older versions), and the response status on it is 403.
Monitoring dashboard feed
GET upgrade at /api/v1/monitoring/ws.
Behavior
The server accepts the connection, tries to send one snapshot straight away, and then sends the whole snapshot to every connected socket once per second. There is no filtering and no subscription message; every client gets every frame.
The server reads text frames from the client and ignores them. Sending messages is pointless, but harmless. There is no application-level heartbeat or ping on this endpoint. Frames arrive every second while the feed is running, which doubles as a liveness signal.
Message schema, server to client
Every frame is a single JSON object. The periodic frames have this shape (values are illustrative):
json
{
"timestamp": "2026-10-04T09:15:02.481233",
"system_metrics": {
"timestamp": "2026-10-04 09:15:01.902114",
"cpu_usage": 12.4,
"memory_usage": 48.1,
"disk_usage": 37.9,
"network_io": {
"bytes_sent": 918273645,
"bytes_recv": 1827364509,
"packets_sent": 1203948,
"packets_recv": 1938475
},
"process_count": 214,
"load_average": [0.42, 0.51, 0.47],
"temperature": null
},
"migration_progress": {},
"database_health": {
"connection_pool_size": 20,
"active_connections": 3,
"idle_connections": 17,
"query_latency_avg": 4.2,
"query_latency_p95": 11.8,
"transactions_per_second": 6.5,
"slow_queries_count": 0,
"deadlocks_count": 0
},
"application_health": {
"service_status": {"api": "healthy", "database": "healthy"},
"response_times": {"api": 0.031},
"error_rates": {"api": 0.0},
"request_rates": {"api": 2.4},
"memory_leaks": [],
"thread_pool_utilization": {"default": 0.1}
},
"capacity_metrics": {
"current_load": 0.12,
"capacity_utilization": 0.18,
"scaling_threshold": 0.8,
"recommended_scaling_action": "none",
"resource_bottlenecks": [],
"forecasted_demand": {"next_hour": 0.14}
},
"activity": [
{
"id": "a1",
"message": "Introspection finished",
"detail": "orders_db: 42 tables",
"timestamp": "2026-10-04 09:14:58.220118",
"type": "info"
}
],
"active_agents": 2,
"discovered_assets": 14,
"new_assets": 1,
"security_risks": 0,
"risk_score": 100,
"security_threats": [],
"compliance_status": [
{"name": "GDPR", "status": "not_assessed"},
{"name": "HIPAA", "status": "not_assessed"},
{"name": "SOC2", "status": "not_assessed"},
{"name": "PCI-DSS", "status": "not_assessed"}
],
"security_scans": []
}The top-level key set is timestamp, system_metrics, migration_progress, database_health, application_health, capacity_metrics, activity, active_agents, discovered_assets, new_assets, security_risks, risk_score, security_threats, compliance_status and security_scans. The metric objects are null until the first sample has been collected. migration_progress is an object keyed by migration id (empty when nothing is running). activity holds up to the last 10 entries and can be an empty array.
Field notes that matter to a parser:
- The top-level
timestampis ISO 8601 with aT. Timestamps nested insidesystem_metricsandactivityhave a space instead of theTand no time zone suffix (2026-10-04 09:15:01.902114). Both are UTC. Parse nested timestamps leniently. compliance_statusentries are{"name", "status"}and, once a real assessment exists, alsocontrols_passedandcontrols_total. A framework that has not been assessed has the statusnot_assessed; the feed never asserts compliance from a default.risk_scorestarts at100and is updated by security scans.migration_progressentries are keyed by id and containmigration_id,status,progress_percentage,current_phase,rows_processed,total_rows,processing_rate,estimated_completion,errors_countandwarnings_count.
The sub-objects update at different intervals. The socket sends once per second regardless, so consecutive frames often carry identical sub-objects. Compare values yourself if you only want changes.
The first frame
The server tries to send a first snapshot immediately after accepting. That snapshot has the same keys plus an actions array (pending system action items with id, title, description, priority, created_at and status), which the periodic frames do not carry. The periodic frames never include actions; fetch GET /api/v1/monitoring/data if you need it again. The first snapshot is not guaranteed to arrive: if it is skipped, the first frame you receive is the first periodic one, within about a second. Write clients so they do not depend on the first frame being the one with actions.
Closing
| Situation | What the client sees |
|---|---|
Missing or bad token, or no read permission | Handshake refused with HTTP 403; no socket is opened (see above). |
| The feed is unavailable on the server | The socket is accepted and then closed with code 1011. |
| Unexpected server error | The server closes with code 1011. |
| The backend shuts down | The server closes every connected socket. |
| Client closes | The server stops sending to it. |
Reconnection
The server keeps no per-client state and does not replay frames. A reconnect simply starts receiving the current snapshot again, so a client needs no resume token and loses nothing that matters. Reconnect with a short, growing delay, and fetch a fresh token first if the old one is near expiry (see the refresh section of API overview). Your client should reconnect when the connection drops.
In a standard install this is the only WebSocket path forwarded to the backend. Connections that stay silent for 60 seconds are closed by the standard install's web proxy; because the server sends a frame every second, that does not happen in normal operation.
Python client (websockets)
python
import asyncio
import json
import websockets
WS_URL = "wss://mmune.example.com/api/v1/monitoring/ws"
async def watch_dashboard(token: str) -> None:
delay = 1
while True:
try:
async with websockets.connect(
WS_URL, subprotocols=["mmune.bearer", token]
) as ws:
delay = 1
async for raw in ws:
frame = json.loads(raw)
metrics = frame.get("system_metrics") or {}
print(
frame["timestamp"],
"cpu",
metrics.get("cpu_usage"),
"assets",
frame["discovered_assets"],
)
except websockets.ConnectionClosed as exc:
print(f"closed: {exc.code} {exc.reason}")
except OSError as exc:
print(f"connect failed: {exc}")
await asyncio.sleep(delay)
delay = min(delay * 2, 30)
# asyncio.run(watch_dashboard("<TOKEN>"))This offers the token as a subprotocol, which works the same on every websockets release. To send a header instead, pass additional_headers={"Authorization": "Bearer <TOKEN>"} (the argument is called extra_headers before websockets 14). A refused handshake raises InvalidStatus, which is not caught above on purpose: retrying a 403 with the same token will not help, so let it surface and refresh the token.
TypeScript client
ts
const WS_URL = "wss://mmune.example.com/api/v1/monitoring/ws";
export function watchDashboard(
getToken: () => Promise<string>,
onFrame: (frame: Record<string, unknown>) => void,
): () => void {
let stopped = false;
let socket: WebSocket | undefined;
let delay = 1000;
const connect = async () => {
const token = await getToken();
socket = new WebSocket(WS_URL, ["mmune.bearer", token]);
socket.onopen = () => {
delay = 1000;
};
socket.onmessage = (event) => onFrame(JSON.parse(event.data));
socket.onclose = () => {
if (stopped) return;
setTimeout(connect, delay);
delay = Math.min(delay * 2, 30000);
};
};
void connect();
return () => {
stopped = true;
socket?.close();
};
}Streaming chat (SSE)
POST /api/v1/analysis/chat/stream. It is the streaming variant of POST /api/v1/analysis/chat, which returns the same answer as one JSON body.
Authentication
An ordinary Bearer header, as in API overview. The route needs an authenticated, active user and read is enough: it is one of the mutating calls that a read-only token may make. Action tools inside chat check the caller's write permission themselves. Because the request is plain HTTP, authentication failures are normal 401 and 403 JSON responses, sent before any stream starts.
This is a POST, so the browser EventSource API cannot be used (it only issues GET requests and cannot send a body or header). Use fetch and read the response body as a stream, as the web console does.
Request
json
{
"project_id": "default",
"question": "How many integrations are registered and which are degraded?",
"history": [
{"role": "user", "content": "What is mmune watching?"},
{"role": "assistant", "content": "It is watching 4 registered integrations."}
]
}project_id and question are required strings. history is optional, a list of objects with role and content strings.
Response
The status is 200 and the content type is text/event-stream. Each event is one data: line holding a JSON object, followed by a blank line. There are no event: names, no id: fields and no keep-alive comments; the event kind is the type field inside the JSON.
type | Fields | Meaning |
|---|---|---|
meta | used_tools (array of strings), path (fast or slow) | Always first. Names the data tools used and which route answered. fast means a deterministic skill answered with no language model; slow means the model path. |
delta | text | Part of the answer text. Concatenate in order. |
block | block (object) | A structured payload, for example a table: {"type": "table", "title": "...", "columns": [...], "rows": [[...]]}. Zero or more, only on the fast path. |
done | path, console_links | Always last on success. console_links is an array of {"label", "route", "tab", "focus_node_id"} pointing at the web console page that shows the underlying data. |
error | message | The stream failed after it started. Nothing follows. |
On the fast path the stream is meta, one delta with the summary, any block events, then done. On the slow path the model first decides which tools to call and the server runs them without streaming, so meta appears after that delay. On the slow path the answer usually arrives as a single delta rather than word by word. If no real language model is configured, the slow path returns a short delta with console navigation hints instead of generated text; it never fabricates an answer.
A complete fast-path exchange on the wire:
text
data: {"type": "meta", "used_tools": ["list_integrations"], "path": "fast"}
data: {"type": "delta", "text": "4 integrations are registered; 1 is degraded."}
data: {"type": "block", "block": {"type": "table", "title": "Integrations", "columns": ["name", "status"], "rows": [["orders_db", "healthy"], ["billing_db", "degraded"]]}}
data: {"type": "done", "path": "fast", "console_links": [{"label": "Data Explorer", "route": "/data", "tab": "integrations", "focus_node_id": null}]}The tool name and values above are illustrative. A failure after the stream has started looks like this, and ends the stream:
text
data: {"type": "meta", "used_tools": [], "path": "slow"}
data: {"type": "error", "message": "<provider error text>"}Closing, timeouts and retries
The stream ends when the server finishes the response. Always stop reading after done or error. If the connection drops before either arrives, the answer is incomplete; the server holds no state to resume from, so send the question again. Retrying repeats any model cost, as with any POST (see the idempotency note in API overview).
Chat requests can take minutes on a local model. A standard install allows 660 seconds for API responses. A proxy in front of mmune may buffer the response and deliver the events together, so do not rely on seeing them arrive one by one. In practice most answers are a handful of events, so this mostly affects latency to first byte.
Python client (requests)
python
import json
import requests
BASE_URL = "https://mmune.example.com"
def ask(token: str, question: str, project_id: str = "default") -> str:
resp = requests.post(
f"{BASE_URL}/api/v1/analysis/chat/stream",
headers={"Authorization": f"Bearer {token}"},
json={"project_id": project_id, "question": question},
stream=True,
timeout=(10, 700), # connect, then read; above the 660 s response limit
)
resp.raise_for_status()
parts: list[str] = []
for line in resp.iter_lines(decode_unicode=True):
if not line or not line.startswith("data: "):
continue
event = json.loads(line[len("data: "):])
if event["type"] == "delta":
parts.append(event["text"])
elif event["type"] == "error":
raise RuntimeError(event["message"])
elif event["type"] == "done":
break
return "".join(parts)
print(ask("<TOKEN>", "Which integrations are degraded?"))TypeScript client
ts
type ChatEvent =
| { type: "meta"; used_tools: string[]; path: "fast" | "slow" }
| { type: "delta"; text: string }
| { type: "block"; block: Record<string, unknown> }
| { type: "done"; path: "fast" | "slow"; console_links: unknown[] }
| { type: "error"; message: string };
export async function ask(
token: string,
question: string,
onEvent: (event: ChatEvent) => void,
projectId = "default",
): Promise<void> {
const res = await fetch("https://mmune.example.com/api/v1/analysis/chat/stream", {
method: "POST",
headers: {
Authorization: `Bearer ${token}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ project_id: projectId, question }),
});
if (!res.ok || !res.body) throw new Error(`${res.status} ${await res.text()}`);
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
for (;;) {
const { done, value } = await reader.read();
if (done) return;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split("\n");
buffer = lines.pop() ?? "";
for (const line of lines) {
if (!line.startsWith("data: ")) continue;
const event = JSON.parse(line.slice(6)) as ChatEvent;
onEvent(event);
if (event.type === "done" || event.type === "error") return;
}
}
}Agent state and cancellation routes
These are ordinary HTTP routes, not streams. They are the controls that go with a long-running agent execution identified by an execution_id. Errors use the shapes in API overview. Routes marked write need the write permission; the others need read.
| Method | Path | Permission | Purpose |
|---|---|---|---|
| GET | /api/v1/agent/state/{execution_id} | read | Latest checkpoint for the execution. |
| GET | /api/v1/agent/state/{execution_id}/can-resume | read | Whether the execution can be resumed. |
| POST | /api/v1/agent/state/{execution_id}/resume | write | Mark the execution resumed and return its state. |
| DELETE | /api/v1/agent/state/{execution_id} | write | Delete all checkpoints for the execution. |
| POST | /api/v1/agent/cancel/{execution_id} | write | Request cancellation. |
| GET | /api/v1/agent/cancel-status/{execution_id} | read | Whether cancellation has been requested. |
Checkpoint read, GET /api/v1/agent/state/{execution_id}:
json
{
"success": true,
"message": "State retrieved successfully",
"state": {
"state_id": "3b1e...",
"execution_id": "exec-9c21",
"agent_name": "discovery_agent",
"agent_type": "discovery",
"status": "running",
"checkpoint_number": 3,
"state_data": {},
"context_data": {},
"progress_data": {},
"created_at": "2026-10-04T09:10:00",
"updated_at": "2026-10-04T09:14:12",
"resumed_count": 0,
"last_error": null,
"error_count": 0,
"can_resume": true,
"resume_message": null
}
}An unknown execution id returns 404 with {"detail": "State not found"}. The contents of state_data, context_data and progress_data depend on the agent. Resume check:
json
{ "execution_id": "exec-9c21", "can_resume": true }Resume returns the same envelope as the checkpoint read, with a message such as Execution resumed (resume count: 1). If the execution is completed, cancelled or has no state, it answers 400 with {"detail": "Execution cannot be resumed (completed, cancelled, or no state found)"}. Deleting checkpoints returns {"success": true, "message": "State checkpoints deleted", "state": null}.
Cancel takes an optional reason (default Cancelled by user):
bash
curl -s -X POST "https://mmune.example.com/api/v1/agent/cancel/exec-9c21" \
-H "Authorization: Bearer <TOKEN>" \
-H 'Content-Type: application/json' \
-d '{"reason": "Wrong target"}'json
{
"success": true,
"message": "Execution cancelled: Wrong target",
"execution_id": "exec-9c21",
"cancelled": true
}The cancel route requests cancellation of the execution and, if a checkpoint exists, marks it cancelled and not resumable. Cancellation is best effort. Status:
json
{ "execution_id": "exec-9c21", "is_cancelled": true, "reason": "Wrong target" }reason is null when the execution has not been cancelled. The /api/v1/workflows/templates and /api/v1/workflows/from-template routes are plain request and response routes, documented with the workflow endpoints in Agents and workflows.