Appearance
Agents, workflows, orchestration and RAG API
What this page covers
This page documents the HTTP endpoints that expose mmune's agent framework: agent execution history, the orchestrator (plan, execute, poll, download), workflow definitions and templates, task cancellation, agent state checkpoints, document retrieval (RAG), and the chat endpoints the Workspace Chat tab calls. WebSocket progress routes are documented in WebSockets and streaming and are not repeated here.
Prerequisites: a running mmune installation, a user account, and a Bearer token from POST /api/v1/auth/login. Authentication, the error body shape, rate limits and pagination conventions are defined once in API overview and only the deviations are noted below. Example host: https://mmune.example.com. Replace <TOKEN> with your access token.
Conventions used on this page
Every route on this page requires a valid Bearer token whose user has the read permission (or the admin role). Any method other than GET, HEAD and OPTIONS also needs the write permission (or admin), with four exceptions on this page that are read-sufficient: POST /api/v1/analysis/chat, POST /api/v1/analysis/chat/stream, POST /api/v1/rag/query and POST /api/v1/orchestration/plan. The "Permission" line for each endpoint below is the effective requirement.
A missing or bad token returns 401. A missing permission returns 403 with detail set to an object {"code": "insufficient_permission", "required": "read|write|admin", "message": "..."}. Request bodies that fail schema validation return 422 as described in the overview. Those three are not repeated in every status code list.
IDs: agent execution IDs and orchestration job IDs are strings.
Endpoint index
| Method | Path | Purpose |
|---|---|---|
| POST | /api/v1/analysis/chat | Ask a question about the deployment, get one JSON answer |
| POST | /api/v1/analysis/chat/stream | Same question, answer streamed as server-sent events |
| GET | /api/v1/agents/executions | List agent execution history with filters |
| GET | /api/v1/agents/executions/{execution_id} | Full record for one agent execution |
| GET | /api/v1/agents/statistics | Aggregate success, failure and latency figures |
| POST | /api/v1/orchestration/plan | Ask the orchestrator to turn a goal into a step plan |
| POST | /api/v1/orchestration/execute | Run a plan as a background job |
| GET | /api/v1/orchestration/status/{execution_id} | Poll a job's status, logs and results |
| GET | /api/v1/orchestration/artifacts/{execution_id}/download | Download a completed job's results as JSON |
| POST | /api/v1/workflows/register | Register a workflow definition |
| GET | /api/v1/workflows/templates | List workflow templates |
| GET | /api/v1/workflows/templates/{template_id} | Get one workflow template summary |
| POST | /api/v1/workflows/from-template | Build and register a workflow from a template |
| POST | /api/v1/agent/cancel/{execution_id} | Request cancellation of an agent execution |
| GET | /api/v1/agent/cancel-status/{execution_id} | Read the cancellation flag |
| GET | /api/v1/agent/state/{execution_id} | Latest state checkpoint for an execution |
| GET | /api/v1/agent/state/{execution_id}/can-resume | Whether the execution can be resumed |
| POST | /api/v1/agent/state/{execution_id}/resume | Mark an execution resumed and return its checkpoint |
| DELETE | /api/v1/agent/state/{execution_id} | Delete all checkpoints for an execution |
| POST | /api/v1/rag/ingest | Ingest a text document |
| POST | /api/v1/rag/ingest/file | Ingest an uploaded file |
| POST | /api/v1/rag/query | Retrieve context and generate an answer |
| GET | /api/v1/rag/context | Retrieve context chunks only |
| GET | /api/v1/rag/documents | List ingested documents |
| GET | /api/v1/rag/documents/{document_name} | Status of one document |
| DELETE | /api/v1/rag/documents/{document_name} | Delete a document and its chunks |
LLM provider dependence and latency
Several endpoints here call a language model. What you get back depends on which provider is configured and healthy, so build callers to expect slow answers and degraded answers.
A provider counts as available when it is configured with a real key and its last health check passed. Providers are configured at install time through environment variables such as ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY and MMUNE_LOCAL_LLM_BASE_URL plus MMUNE_LOCAL_LLM_MODEL. A key set to the literal placeholder is treated as unset. See configuration.
What each endpoint does when no healthy real provider exists:
| Endpoint | Behavior without a healthy real LLM |
|---|---|
POST /analysis/chat and /chat/stream | Questions the deterministic skill router recognises are still answered (path: "fast", no LLM involved). Everything else returns HTTP 200 with a fixed message saying no model is configured, or that the configured model is temporarily unavailable, plus links to console pages. It is not an error status, so check path and read answer. |
POST /orchestration/plan and /execute | 503 with a message listing the provider variables to set. |
POST /rag/ingest and /ingest/file | 422 No embedding-capable AI provider available... when no provider can produce embeddings. |
POST /rag/query | If no embedding provider exists, retrieval returns no chunks and generation runs on an empty context, so the answer will say it does not know. If generation has no provider the call fails with 500. |
GET /rag/context | Returns an empty list when no embedding provider exists. |
Latency is bounded per model call. Cloud provider calls time out after 30 seconds by default. Calls to a local model time out after 120 seconds. A request that makes several model calls (a tool-using chat answer, a plan, several RAG calls) can take longer than any single timeout. On CPU-only local inference, set client timeouts well above 2 minutes for chat and plan calls.
An agent run has its own default time limit of 300 seconds.
Agent roster as seen through the API
There is no endpoint that lists agents. You see an agent through its agent_type string, which appears in GET /api/v1/agents/executions, in orchestration plan steps, and in workflow task definitions.
| agent_type | What it does | Where you meet it in the API |
|---|---|---|
discovery | Scans hosts and ports for databases. A step may carry probe_url and probe_key params to run the scan on a remote probe instead of locally. | Orchestration step actions scan_network or scan. |
analysis | Schema analysis, complexity scoring, relationship analysis and documentation generation for several database types. | Orchestration actions analyze_schema or analyze. The chat endpoints also report agent_type analysis. |
mapping | Field mapping with semantic similarity and pattern recognition, producing mapping suggestions. | Orchestration actions generate_mapping, map or suggest. |
security | Identifies PII and PHI and recommends controls. | Orchestration actions scan_database or scan. |
validation | Validates data and mappings: integrity checks, anomaly detection, business rules, CDC quality. | Appears in execution history and as a task agent_type in workflow templates. The orchestrator cannot delegate to it. |
execution | Runs generated migration scripts: script executor, batch processor, migration validator and monitor. | Appears in execution history and as a workflow task agent_type. The orchestrator cannot delegate to it. |
orchestration | Uses the LLM to break a goal into steps, then runs the steps as a dependency graph, running independent steps concurrently. | POST /orchestration/plan and /execute. |
The orchestrator only delegates to four agent types: discovery, analysis, mapping and security. A plan step with any other agent_type, or an action not listed above for its agent, makes the job end in FAILED with Unsupported agent/action or Unknown agent type.
Chat endpoints (Workspace Chat tab)
The Chat tab in the Workspace page sends planning requests to POST /api/v1/orchestration/plan and all other messages to POST /api/v1/analysis/chat/stream, rendering the events as they arrive. Both use the project ID default.
POST /api/v1/analysis/chat
Answers a question about the mmune deployment in one JSON response. A deterministic skill router handles common structured questions without calling an LLM (path: "fast"). Everything else goes to the LLM, which can call read-only data tools (path: "slow").
Permission: read. Action tools that the model may choose to call are gated per caller: they run only if the caller has write or is admin.
| Field (body) | Type | Required | Description |
|---|---|---|---|
project_id | string | yes | Required. The Chat tab sends default. |
question | string | yes | The question to answer. |
history | array of {role, content} | no | Earlier turns, both strings. Default null. |
Response 200:
| Field | Type | Description |
|---|---|---|
answer | string | Always present. Plain text. |
blocks | array of objects | Structured payloads such as tables. Empty for LLM answers. |
console_links | array of {label, route, tab, focus_node_id} | Where in the console to look. tab and focus_node_id may be null. |
path | string | fast for the deterministic path, slow for the LLM path or the no-model message. |
used_tools | array of strings | Data tools consulted. |
Status codes: 200, plus the common 401, 403, 422. Provider and tool failures do not return 500. A failed chat call comes back as HTTP 200 with a degraded answer, for example the fixed text I encountered an error while processing your question., so read the text rather than relying on the status code.
bash
curl -sS -X POST "https://mmune.example.com/api/v1/analysis/chat" \
-H "Authorization: Bearer <TOKEN>" \
-H "Content-Type: application/json" \
-d '{"project_id": "default", "question": "Which integrations are degraded right now?"}'POST /api/v1/analysis/chat/stream
Same inputs as /chat, answered as a text/event-stream. Each frame is a line data: <json> followed by a blank line. Permission: read, with the same per-caller gating of action tools.
Request body: identical to /chat.
Event types:
type | Other fields | Meaning |
|---|---|---|
meta | used_tools (array), path (fast or slow) | First frame. |
delta | text (string) | A piece of the answer. Append to what you have. On the slow path expect the answer as one delta rather than token by token. |
block | block (object) | A structured payload such as a table. Zero or more. |
done | path, console_links | Last frame on success. |
error | message (string) | Terminal failure. The HTTP status is still 200 because the stream has started. |
Status codes: 200 for any stream, plus the common 401, 403, 422.
bash
curl -sS -N -X POST "https://mmune.example.com/api/v1/analysis/chat/stream" \
-H "Authorization: Bearer <TOKEN>" \
-H "Content-Type: application/json" \
-d '{"project_id": "default", "question": "Summarise open drift alerts"}'Agents router (/api/v1/agents)
These three routes are read-only views over the agent execution history. They need the read permission.
GET /api/v1/agents/executions
Lists agent executions, filtered by the query parameters.
Permission: read.
| Field (query) | Type | Required | Default | Description |
|---|---|---|---|---|
agent_type | string | no | none | Filter by agent type, for example discovery. |
status | string | no | none | started, completed or failed. |
project_id | string | no | none | Filter by project. |
user_id | string | no | none | Filter by the user who triggered it. |
days | integer | no | 7 | Look back this many days. A value of 0 or less removes the date filter. |
limit | integer | no | 100 | 1 to 1000. |
offset | integer | no | 0 | 0 or more. Use with limit to page. |
Response 200 is a JSON array. Each item:
| Field | Type | Description |
|---|---|---|
id | string | Execution ID, usable in the detail and state endpoints. |
agent_id | string | Agent instance ID. |
agent_type | string | For example discovery, analysis, mapping. |
agent_name | string | Configured agent name. |
status | string | started, completed or failed. |
context_id, project_id, session_id, user_id | string or null | Correlation fields. |
started_at | string | ISO 8601 timestamp, empty string if missing. |
completed_at | string or null | ISO 8601 timestamp. |
duration_ms | integer or null | Wall-clock duration. |
error_message | string or null | Set on failure. |
retry_count | integer | Retries attempted. |
Status codes: 200.
bash
curl -sS "https://mmune.example.com/api/v1/agents/executions?agent_type=discovery&status=failed&days=1&limit=20" \
-H "Authorization: Bearer <TOKEN>"GET /api/v1/agents/executions/
Returns the full record for one execution, including input and output data.
Permission: read.
| Field (path) | Type | Required | Description |
|---|---|---|---|
execution_id | string | yes | The id from the list endpoint. |
Response 200 fields: id, agent_id, agent_type, agent_name, status, context_id, project_id, session_id, user_id, input_data, output_data, error_message, error_type, started_at, completed_at, duration_ms, retry_count, metrics, created_at. Timestamps are ISO 8601 strings or null.
Status codes: 200, 404 Execution not found.
bash
curl -sS "https://mmune.example.com/api/v1/agents/executions/<EXECUTION_ID>" \
-H "Authorization: Bearer <TOKEN>"GET /api/v1/agents/statistics
Aggregates executions over a window.
Permission: read.
| Field (query) | Type | Required | Default | Description |
|---|---|---|---|---|
agent_type | string | no | all types | Restrict to one agent type. |
days | integer | no | 7 | Window length in days. |
Response 200:
| Field | Type | Description |
|---|---|---|
total_executions | integer | Executions in the window. |
successful | integer | Status completed. |
failed | integer | Status failed. |
average_duration_ms | integer | Mean of executions that have a duration; 0 if none. |
success_rate | number | successful / total_executions, 0.0 when there are none. |
Status codes: 200.
bash
curl -sS "https://mmune.example.com/api/v1/agents/statistics?agent_type=mapping&days=30" \
-H "Authorization: Bearer <TOKEN>"Orchestration router (/api/v1/orchestration)
This is the flow behind the Workspace Workflow tab and the "plan" path of Chat: generate a plan from a goal, review it, execute it as a background job, poll the job, download the results. See LLM provider dependence for the 503 behavior and for which agents and actions a plan may use.
POST /api/v1/orchestration/plan
Generates a plan with the LLM. The plan is returned to you and is not stored or executed.
Permission: read (this route is on the read-sufficient list because it only plans).
| Field (body) | Type | Required | Default | Description |
|---|---|---|---|---|
goal | string | yes | none | What you want done, in plain language. |
project_id | string | yes | none | Project the plan belongs to. |
session_id | string | no | generated UUID | Correlation ID. |
Response 200 (the plan object):
| Field | Type | Description |
|---|---|---|
plan_id | string | UUID. |
goal | string | Echo of the goal. |
steps | array | Ordered steps, see below. |
created_at | string | ISO 8601 timestamp. |
status | string | created. |
Each step has step_id (string; the model's id if it gave one, else a UUID), description, agent_type, action, params (object), dependencies (array of step_id values that must finish first), status (pending) and result (null).
Status codes: 200, 503 no real AI provider configured, 500 on a provider error or when the model returns no usable steps (Planner produced no executable steps). The model, not mmune, chooses the steps, so review agent_type and action before you execute.
bash
curl -sS -X POST "https://mmune.example.com/api/v1/orchestration/plan" \
-H "Authorization: Bearer <TOKEN>" \
-H "Content-Type: application/json" \
-d '{"goal": "Discover databases on host 10.0.0.12 and analyse their schemas", "project_id": "default"}'POST /api/v1/orchestration/execute
Starts a plan as a background job and returns immediately. The job runs in the background.
Permission: write.
| Field (body) | Type | Required | Default | Description |
|---|---|---|---|---|
plan | object | yes | none | A plan as returned by /plan. It must contain plan_id and goal. steps is read with a default of an empty list, so a plan with no steps key is accepted and the job completes with status COMPLETED and empty results. Each step must contain description, agent_type, action, params, dependencies and step_id (or id). |
project_id | string | yes | none | Project ID. |
session_id | string | no | generated UUID | Correlation ID. |
A request without plan or project_id returns 422. If the plan itself is incomplete, for example a step is missing a required field, the job ends as FAILED with the reason in error and in the logs.
Response 200:
| Field | Type | Description |
|---|---|---|
message | string | Plan execution started. |
execution_id | string | Job ID, a UUID string. Use it for status and download. |
status | string | RUNNING. |
Status codes: 200, 422, 503 no real AI provider configured.
bash
curl -sS -X POST "https://mmune.example.com/api/v1/orchestration/execute" \
-H "Authorization: Bearer <TOKEN>" \
-H "Content-Type: application/json" \
-d '{"plan": {"plan_id": "6f1c2f0e-0000-4000-8000-000000000001", "goal": "Scan host", "steps": [{"step_id": "step_1", "description": "Scan for databases", "agent_type": "discovery", "action": "scan_network", "params": {"hosts": ["10.0.0.12"]}, "dependencies": []}]}, "project_id": "default"}'GET /api/v1/orchestration/status/
Returns job state and new log lines. Poll this until status is COMPLETED or FAILED.
Permission: read.
| Field | In | Type | Required | Default | Description |
|---|---|---|---|---|---|
execution_id | path | string | yes | none | Job ID from /execute. |
cursor | query | integer | no | 0 | Log position. Pass the previous next_cursor to receive only new lines. Must be 0 or more. |
Response 200:
| Field | Type | Description |
|---|---|---|
status | string | RUNNING, COMPLETED or FAILED. |
logs | array of strings | Log lines from cursor onward. |
next_cursor | integer | Cursor for the next poll. |
results | object or null | Updated after each finished step: steps (full step objects) and results (map of step_id to step output). On completion it also carries plan_id, goal, execution_id and completed_at. |
error | string or null | Set when FAILED. |
created_at, updated_at | string | Timestamps. |
metadata | object | plan_id, project_id, session_id. |
Status codes: 200, 404 Execution ID not found.
bash
curl -sS "https://mmune.example.com/api/v1/orchestration/status/<EXECUTION_ID>?cursor=0" \
-H "Authorization: Bearer <TOKEN>"GET /api/v1/orchestration/artifacts/{execution_id}/download
Downloads the results of a completed job as a JSON attachment.
Permission: read.
| Field | In | Type | Required | Default | Description |
|---|---|---|---|---|---|
execution_id | path | string | yes | none | Job ID. |
format | query | string | no | json | Only json is supported. |
Response 200 is application/json with Content-Disposition: attachment; filename=orchestration-artifacts-<id>-<YYYYMMDD>.json. Body fields: execution_id, status, created_at, updated_at, metadata, artifacts (the job's result object).
Status codes: 200, 400 when the job is not COMPLETED or the format is not json, 404 when the job is unknown or has no stored result.
bash
curl -sS -OJ "https://mmune.example.com/api/v1/orchestration/artifacts/<EXECUTION_ID>/download?format=json" \
-H "Authorization: Bearer <TOKEN>"Workflows router (/api/v1/workflows)
A workflow is a named set of tasks with dependencies. Each task names an agent_type. The routes below manage workflow definitions and templates. Work is run through the orchestration endpoints (plan, execute, status, download).
POST /api/v1/workflows/register
Registers a workflow definition with the engine. The definition is validated for missing dependencies and cycles.
Permission: write.
| Field (body) | Type | Required | Default | Description |
|---|---|---|---|---|
workflow_id | string | yes | none | Your identifier. Registering the same ID again replaces the definition. |
name | string | yes | none | Display name. |
description | string | yes | none | Description. |
tasks | array of objects | yes | none | Task definitions, below. |
initial_data | object | no | {} | Initial workflow data. |
version | string | no | 1.0 | Definition version. |
Each task object: task_id (required), name (required), agent_type (required), depends_on (array of task_id, default []), input_mapping (object, default {}), output_mapping (object, default {}), retry_count (integer, default 0), timeout_seconds (integer, default 300), optional (boolean, default false). Other keys are ignored. A task missing a required key returns 500.
Response 200: {"success": true, "message": "Workflow '<name>' registered successfully", "workflow_id": "<id>"}.
Status codes: 200, 400 when validation finds a missing dependency or a cycle, 500 otherwise.
bash
curl -sS -X POST "https://mmune.example.com/api/v1/workflows/register" \
-H "Authorization: Bearer <TOKEN>" \
-H "Content-Type: application/json" \
-d '{"workflow_id": "nightly-scan", "name": "Nightly scan", "description": "Discover then analyse", "tasks": [{"task_id": "discover", "name": "Discover", "agent_type": "discovery"}, {"task_id": "analyse", "name": "Analyse", "agent_type": "analysis", "depends_on": ["discover"]}]}'Workflow templates router
Templates are predefined. Creating from a template builds a new definition with a fresh UUID workflow_id and registers it with the engine.
Available templates:
template_id | Name | Tasks | Required body fields | Task agent types, in order |
|---|---|---|---|---|
full_migration | Full Migration | 6 | source_db, target_db, project_name | discovery, analysis, mapping, validation, execution, qa |
schema_analysis | Schema Analysis | 3 | source_db, project_name | discovery, analysis, documentation |
data_validation | Data Validation | 3 | source_db, target_db, project_name | qa, validation, documentation |
GET /api/v1/workflows/templates
Lists templates.
Permission: read.
Response 200: {"success": true, "count": 3, "templates": {"<template_id>": {"name", "description", "steps"}}}. In this response steps is the number of tasks.
Status codes: 200.
bash
curl -sS "https://mmune.example.com/api/v1/workflows/templates" \
-H "Authorization: Bearer <TOKEN>"GET /api/v1/workflows/templates/
Returns the summary entry for one template, the same object shown in the list.
Permission: read.
| Field (path) | Type | Required | Description |
|---|---|---|---|
template_id | string | yes | full_migration, schema_analysis or data_validation. |
Response 200: {"success": true, "template_id": "<id>", "template": {"name", "description", "steps"}}.
Status codes: 200, 404 Template not found.
bash
curl -sS "https://mmune.example.com/api/v1/workflows/templates/schema_analysis" \
-H "Authorization: Bearer <TOKEN>"POST /api/v1/workflows/from-template
Builds a workflow from a template and registers it.
Permission: write.
| Field (body) | Type | Required | Default | Description |
|---|---|---|---|---|
template_id | string | yes | none | One of the template IDs above. |
project_name | string | yes | none | Used in the workflow name, for example <project_name> - Schema Analysis. |
source_db | string | yes | none | Source connection string, copied into task config. |
target_db | string | conditional | null | Required for full_migration and data_validation. |
Connection strings you send are stored in the task config of the registered definition. Use a least-privilege read account, and avoid putting production passwords in a URL you do not want logged client side.
Response 200:
| Field | Type | Description |
|---|---|---|
success | boolean | true. |
message | string | Workflow created from template: <template_id>. |
workflow_id | string | New UUID. |
workflow | object | workflow_id, name, description, version, initial_data, and tasks where each task has task_id, name, step_type, agent_type, depends_on, config, retry_count, timeout_seconds, optional. |
Status codes: 200, 400 when target_db is missing for a template that needs it, 404 Template not found: <id>, 500.
bash
curl -sS -X POST "https://mmune.example.com/api/v1/workflows/from-template" \
-H "Authorization: Bearer <TOKEN>" \
-H "Content-Type: application/json" \
-d '{"template_id": "schema_analysis", "project_name": "CRM Docs", "source_db": "postgresql://reader@db.example.com:5432/crm"}'Cancellation and state routers
These routes act on agent execution IDs (the id values from /agents/executions), not on orchestration job IDs.
Cancellation is best effort.
State checkpoints are stored per execution. If none exist for an ID, the state routes report not found or cannot resume.
POST /api/v1/agent/cancel/
Requests cancellation of an agent execution.
Permission: write.
| Field | In | Type | Required | Default | Description |
|---|---|---|---|---|---|
execution_id | path | string | yes | none | Agent execution ID. |
reason | body | string | no | Cancelled by user | Recorded as the cancellation reason. |
The body object is required even when empty; send {} to accept the default reason.
Response 200: {"success": true, "message": "Execution cancelled: <reason>", "execution_id": "<id>", "cancelled": true}. Calling it twice is safe; the first reason is kept.
Status codes: 200, 500.
bash
curl -sS -X POST "https://mmune.example.com/api/v1/agent/cancel/<EXECUTION_ID>" \
-H "Authorization: Bearer <TOKEN>" \
-H "Content-Type: application/json" \
-d '{"reason": "Wrong source selected"}'GET /api/v1/agent/cancel-status/
Reads the cancellation flag.
Permission: read.
Response 200: {"execution_id": "<id>", "is_cancelled": false, "reason": null}. reason is a string once cancelled.
Status codes: 200, 500.
bash
curl -sS "https://mmune.example.com/api/v1/agent/cancel-status/<EXECUTION_ID>" \
-H "Authorization: Bearer <TOKEN>"GET /api/v1/agent/state/
Returns the latest checkpoint for an execution.
Permission: read.
Response 200:
| Field | Type | Description |
|---|---|---|
success | boolean | true. |
message | string | State retrieved successfully. |
state | object | state_id, execution_id, agent_name, agent_type, status, checkpoint_number, state_data, context_data, progress_data, created_at, updated_at, resumed_count, last_error, error_count, can_resume, resume_message. |
Status codes: 200, 404 State not found, 500.
bash
curl -sS "https://mmune.example.com/api/v1/agent/state/<EXECUTION_ID>" \
-H "Authorization: Bearer <TOKEN>"GET /api/v1/agent/state/{execution_id}/can-resume
Tells you whether a resume would be accepted. An execution can resume when a checkpoint exists, its can_resume flag is true, and its status is neither completed nor cancelled.
Permission: read.
Response 200: {"execution_id": "<id>", "can_resume": true}.
Status codes: 200, 500.
bash
curl -sS "https://mmune.example.com/api/v1/agent/state/<EXECUTION_ID>/can-resume" \
-H "Authorization: Bearer <TOKEN>"POST /api/v1/agent/state/{execution_id}/resume
Records a resume: increments resumed_count, sets status active, and returns the checkpoint. It does not itself start any work.
Permission: write.
Response 200: same shape as the get-state response, with message of Execution resumed (resume count: <n>).
Status codes: 200, 400 Execution cannot be resumed (completed, cancelled, or no state found), 500.
bash
curl -sS -X POST "https://mmune.example.com/api/v1/agent/state/<EXECUTION_ID>/resume" \
-H "Authorization: Bearer <TOKEN>"DELETE /api/v1/agent/state/
Deletes every checkpoint for the execution. It returns success even if there were none.
Permission: write.
Response 200: {"success": true, "message": "State checkpoints deleted", "state": null}.
Status codes: 200, 500.
bash
curl -sS -X DELETE "https://mmune.example.com/api/v1/agent/state/<EXECUTION_ID>" \
-H "Authorization: Bearer <TOKEN>"RAG router (/api/v1/rag)
Documents are split into chunks, embedded with an embedding-capable provider, and stored in the mmune database. Every user of a deployment sees the same document set.
Retrieval uses cosine similarity. Chunks below min_similarity are dropped. If a chunk was embedded by a different provider than the one now answering, retrieval still runs, but quality may be worse. Re-ingest after you change embedding providers.
POST /api/v1/rag/ingest
Ingests text. Re-ingesting under the same name replaces that document's chunks.
Permission: write.
| Field (body) | Type | Required | Default | Description |
|---|---|---|---|---|
name | string | yes | none | Unique document name; used as the key for updates and deletes. |
content | string | yes | none | Raw text. |
metadata | object | no | null | Stored with each chunk. |
chunk_size | integer | no | 1000 | 100 to 10000 characters. |
chunk_overlap | integer | no | 100 | 0 to 2000 characters. |
Response 200: {"status": "success", "document_name": "<name>", "chunks_created": <n>}. Ingestion makes one embedding call per chunk, so long documents on a slow local model take time.
Status codes: 200, 422 for no embedding provider or invalid content, 500.
bash
curl -sS -X POST "https://mmune.example.com/api/v1/rag/ingest" \
-H "Authorization: Bearer <TOKEN>" \
-H "Content-Type: application/json" \
-d '{"name": "failover-runbook", "content": "If the primary database is unreachable, check the replica lag before promoting.", "metadata": {"owner": "dba-team"}}'POST /api/v1/rag/ingest/file
Ingests an uploaded file as multipart/form-data. Plain text, Markdown, CSV and JSON are decoded as UTF-8 (invalid bytes are replaced). PDF and DOCX uploads need the optional document parsers to be installed in your deployment; without them the route returns 422.
Permission: write.
| Field | In | Type | Required | Default | Description |
|---|---|---|---|---|---|
file | form | file | yes | none | The upload. |
name | query | string | no | the filename | Document name. |
chunk_size | query | integer | no | 1000 | 100 to 10000. |
chunk_overlap | query | integer | no | 100 | 0 to 2000. |
Response 200: same as /ingest. The chunk metadata records source_filename and content_type.
Status codes: 200, 422 when text cannot be extracted, the file has no text, or no embedding provider is available, 500.
bash
curl -sS -X POST "https://mmune.example.com/api/v1/rag/ingest/file?name=failover-runbook" \
-H "Authorization: Bearer <TOKEN>" \
-F "file=@./failover-runbook.md"POST /api/v1/rag/query
Retrieves matching chunks and asks the LLM to answer using only that context. Temperature is 0.
Permission: read.
| Field (body) | Type | Required | Default | Description |
|---|---|---|---|---|
query | string | yes | none | The question. |
provider | string | no | registry choice | Preferred provider name for generation. |
model | string | no | RAG_GENERATION_MODEL, else the default provider's model, else gpt-4 | Model ID for generation. |
limit | integer | no | 5 | Chunks to retrieve, 1 to 50. |
min_similarity | number | no | 0.7 | 0.0 to 1.0. |
Response 200:
| Field | Type | Description |
|---|---|---|
answer | string | Generated answer. The prompt tells the model to say it does not know if the context lacks the answer. |
sources | array of {document_name, similarity, chunk_index} | Chunks that were used. Empty when nothing cleared min_similarity. |
provider | string or null | Provider that generated the answer. |
model | string | Model used. |
Status codes: 200, 500 on generation or provider failure. An empty sources array with a "don't know" answer is a normal 200, not an error.
bash
curl -sS -X POST "https://mmune.example.com/api/v1/rag/query" \
-H "Authorization: Bearer <TOKEN>" \
-H "Content-Type: application/json" \
-d '{"query": "What do I check before promoting the replica?", "limit": 5, "min_similarity": 0.7}'GET /api/v1/rag/context
Returns matching chunks without generating an answer. Use it to see what retrieval finds, or to feed your own model.
Permission: read.
| Field (query) | Type | Required | Default | Description |
|---|---|---|---|---|
query | string | yes | none | Text to search for. |
limit | integer | no | 5 | 1 to 50. |
min_similarity | number | no | 0.7 | 0.0 to 1.0. |
Response 200 is an array of chunks with chunk_id, document_name, document_type, source_url, content, chunk_index, metadata, embedding_provider, embedding_model, embedding_dimension, created_at, updated_at, and similarity (number).
Status codes: 200, 500.
bash
curl -sS -G "https://mmune.example.com/api/v1/rag/context" \
--data-urlencode "query=replica lag before promotion" \
-d "limit=3" \
-H "Authorization: Bearer <TOKEN>"GET /api/v1/rag/documents
Lists ingested documents. The list is not paginated.
Permission: read.
Response 200 is an array of {document_name, chunk_count, ingested_at, updated_at, embedding_provider, embedding_model}, ordered by name. A document embedded by more than one provider or model appears once per combination.
Status codes: 200.
bash
curl -sS "https://mmune.example.com/api/v1/rag/documents" \
-H "Authorization: Bearer <TOKEN>"GET /api/v1/rag/documents/
Status of one document.
Permission: read.
Response 200: {document_name, chunk_count, ingested_at, updated_at, embedding_provider, embedding_model, embedding_dimension}. chunk_count sums across providers; the provider, model and dimension fields come from the first group.
Status codes: 200, 404 Document '<name>' not found.
bash
curl -sS "https://mmune.example.com/api/v1/rag/documents/failover-runbook" \
-H "Authorization: Bearer <TOKEN>"DELETE /api/v1/rag/documents/
Deletes a document and all its chunks.
Permission: write.
Response 200: {"status": "deleted", "document_name": "<name>", "chunks_deleted": <n>}.
Status codes: 200, 404 Document '<name>' not found.
bash
curl -sS -X DELETE "https://mmune.example.com/api/v1/rag/documents/failover-runbook" \
-H "Authorization: Bearer <TOKEN>"Client examples
All examples need a token. POST /api/v1/auth/login takes {"email", "password"} and returns access_token. Accounts with MFA enabled receive a challenge instead; see API overview.
Shared helpers
Python (requests):
python
import requests
BASE = "https://mmune.example.com"
def login(email: str, password: str) -> str:
r = requests.post(
f"{BASE}/api/v1/auth/login",
json={"email": email, "password": password},
timeout=30,
)
r.raise_for_status()
return r.json()["access_token"]
def headers(token: str) -> dict:
return {"Authorization": f"Bearer {token}"}TypeScript (fetch, Node 18 or later, or any browser):
typescript
const BASE = "https://mmune.example.com";
async function login(email: string, password: string): Promise<string> {
const r = await fetch(`${BASE}/api/v1/auth/login`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ email, password }),
});
if (!r.ok) throw new Error(`login failed: ${r.status} ${await r.text()}`);
return (await r.json()).access_token as string;
}
async function api<T>(
token: string,
method: string,
path: string,
body?: unknown,
timeoutMs = 30_000,
): Promise<T> {
const r = await fetch(`${BASE}${path}`, {
method,
headers: {
Authorization: `Bearer ${token}`,
...(body === undefined ? {} : { "Content-Type": "application/json" }),
},
body: body === undefined ? undefined : JSON.stringify(body),
signal: AbortSignal.timeout(timeoutMs),
});
if (!r.ok) throw new Error(`${method} ${path} -> ${r.status} ${await r.text()}`);
return (await r.json()) as T;
}1. Ask the orchestrator a question through chat
This calls POST /api/v1/analysis/chat, the same agent the Chat tab uses for questions. Set the timeout above your provider's timeout, since the slow path can make several model calls. Check path and used_tools to see whether you got a deterministic answer or an LLM answer, and treat a slow-path answer that starts with "No AI model is configured" as the degraded case described earlier.
Python:
python
def ask(token: str, question: str, history: list | None = None) -> dict:
r = requests.post(
f"{BASE}/api/v1/analysis/chat",
headers=headers(token),
json={"project_id": "default", "question": question, "history": history},
timeout=180,
)
r.raise_for_status()
return r.json()
token = login("you@example.com", "<PASSWORD>")
reply = ask(token, "Which integrations are degraded right now?")
print(reply["answer"])
print("path:", reply["path"], "tools:", reply["used_tools"])
for link in reply["console_links"]:
print("see:", link["label"], link["route"], link.get("tab"))TypeScript:
typescript
interface ChatReply {
answer: string;
blocks: unknown[];
console_links: { label: string; route: string; tab?: string | null; focus_node_id?: string | null }[];
path: "fast" | "slow";
used_tools: string[];
}
const token = await login("you@example.com", "<PASSWORD>");
const reply = await api<ChatReply>(
token,
"POST",
"/api/v1/analysis/chat",
{ project_id: "default", question: "Which integrations are degraded right now?" },
180_000,
);
console.log(reply.answer, reply.path, reply.used_tools);If you want the plan flow instead, call POST /api/v1/orchestration/plan with {"goal", "project_id"} and read steps from the response.
2. Query the RAG endpoint
Ingest a document, then ask a question about it. Both calls depend on a configured provider: ingestion needs an embedding-capable one, and the answer needs a generation provider.
Python:
python
def rag_demo(token: str) -> dict:
h = headers(token)
ing = requests.post(
f"{BASE}/api/v1/rag/ingest",
headers=h,
json={
"name": "failover-runbook",
"content": "If the primary is unreachable, check replica lag before promoting.",
},
timeout=300, # one embedding call per chunk
)
ing.raise_for_status()
q = requests.post(
f"{BASE}/api/v1/rag/query",
headers=h,
json={"query": "What do I check before promoting the replica?", "limit": 5},
timeout=180,
)
q.raise_for_status()
body = q.json()
if not body["sources"]:
print("No chunk cleared min_similarity; the answer is not grounded in your documents.")
print(body["answer"])
for s in body["sources"]:
print(s["document_name"], s["chunk_index"], round(s["similarity"], 3))
return bodyTypeScript:
typescript
interface RagAnswer {
answer: string;
sources: { document_name: string; similarity: number; chunk_index: number }[];
provider: string | null;
model: string;
}
async function ragDemo(token: string): Promise<RagAnswer> {
await api(
token,
"POST",
"/api/v1/rag/ingest",
{
name: "failover-runbook",
content: "If the primary is unreachable, check replica lag before promoting.",
},
300_000,
);
const body = await api<RagAnswer>(
token,
"POST",
"/api/v1/rag/query",
{ query: "What do I check before promoting the replica?", limit: 5 },
180_000,
);
if (body.sources.length === 0) {
console.warn("No chunk cleared min_similarity; the answer is not grounded in your documents.");
}
console.log(body.answer, body.sources);
return body;
}Related pages
API overview for authentication, errors, rate limits and pagination. WebSockets and streaming for live progress streams.