Skip to content

Agents, workflows, orchestration and RAG API ​

What this page covers ​

This page documents the HTTP endpoints that expose mmune's agent framework: agent execution history, the orchestrator (plan, execute, poll, download), workflow definitions and templates, task cancellation, agent state checkpoints, document retrieval (RAG), and the chat endpoints the Workspace Chat tab calls. WebSocket progress routes are documented in WebSockets and streaming and are not repeated here.

Prerequisites: a running mmune installation, a user account, and a Bearer token from POST /api/v1/auth/login. Authentication, the error body shape, rate limits and pagination conventions are defined once in API overview and only the deviations are noted below. Example host: https://mmune.example.com. Replace <TOKEN> with your access token.

Conventions used on this page ​

Every route on this page requires a valid Bearer token whose user has the read permission (or the admin role). Any method other than GET, HEAD and OPTIONS also needs the write permission (or admin), with four exceptions on this page that are read-sufficient: POST /api/v1/analysis/chat, POST /api/v1/analysis/chat/stream, POST /api/v1/rag/query and POST /api/v1/orchestration/plan. The "Permission" line for each endpoint below is the effective requirement.

A missing or bad token returns 401. A missing permission returns 403 with detail set to an object {"code": "insufficient_permission", "required": "read|write|admin", "message": "..."}. Request bodies that fail schema validation return 422 as described in the overview. Those three are not repeated in every status code list.

IDs: agent execution IDs and orchestration job IDs are strings.

Endpoint index ​

MethodPathPurpose
POST/api/v1/analysis/chatAsk a question about the deployment, get one JSON answer
POST/api/v1/analysis/chat/streamSame question, answer streamed as server-sent events
GET/api/v1/agents/executionsList agent execution history with filters
GET/api/v1/agents/executions/{execution_id}Full record for one agent execution
GET/api/v1/agents/statisticsAggregate success, failure and latency figures
POST/api/v1/orchestration/planAsk the orchestrator to turn a goal into a step plan
POST/api/v1/orchestration/executeRun a plan as a background job
GET/api/v1/orchestration/status/{execution_id}Poll a job's status, logs and results
GET/api/v1/orchestration/artifacts/{execution_id}/downloadDownload a completed job's results as JSON
POST/api/v1/workflows/registerRegister a workflow definition
GET/api/v1/workflows/templatesList workflow templates
GET/api/v1/workflows/templates/{template_id}Get one workflow template summary
POST/api/v1/workflows/from-templateBuild and register a workflow from a template
POST/api/v1/agent/cancel/{execution_id}Request cancellation of an agent execution
GET/api/v1/agent/cancel-status/{execution_id}Read the cancellation flag
GET/api/v1/agent/state/{execution_id}Latest state checkpoint for an execution
GET/api/v1/agent/state/{execution_id}/can-resumeWhether the execution can be resumed
POST/api/v1/agent/state/{execution_id}/resumeMark an execution resumed and return its checkpoint
DELETE/api/v1/agent/state/{execution_id}Delete all checkpoints for an execution
POST/api/v1/rag/ingestIngest a text document
POST/api/v1/rag/ingest/fileIngest an uploaded file
POST/api/v1/rag/queryRetrieve context and generate an answer
GET/api/v1/rag/contextRetrieve context chunks only
GET/api/v1/rag/documentsList ingested documents
GET/api/v1/rag/documents/{document_name}Status of one document
DELETE/api/v1/rag/documents/{document_name}Delete a document and its chunks

LLM provider dependence and latency ​

Several endpoints here call a language model. What you get back depends on which provider is configured and healthy, so build callers to expect slow answers and degraded answers.

A provider counts as available when it is configured with a real key and its last health check passed. Providers are configured at install time through environment variables such as ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY and MMUNE_LOCAL_LLM_BASE_URL plus MMUNE_LOCAL_LLM_MODEL. A key set to the literal placeholder is treated as unset. See configuration.

What each endpoint does when no healthy real provider exists:

EndpointBehavior without a healthy real LLM
POST /analysis/chat and /chat/streamQuestions the deterministic skill router recognises are still answered (path: "fast", no LLM involved). Everything else returns HTTP 200 with a fixed message saying no model is configured, or that the configured model is temporarily unavailable, plus links to console pages. It is not an error status, so check path and read answer.
POST /orchestration/plan and /execute503 with a message listing the provider variables to set.
POST /rag/ingest and /ingest/file422 No embedding-capable AI provider available... when no provider can produce embeddings.
POST /rag/queryIf no embedding provider exists, retrieval returns no chunks and generation runs on an empty context, so the answer will say it does not know. If generation has no provider the call fails with 500.
GET /rag/contextReturns an empty list when no embedding provider exists.

Latency is bounded per model call. Cloud provider calls time out after 30 seconds by default. Calls to a local model time out after 120 seconds. A request that makes several model calls (a tool-using chat answer, a plan, several RAG calls) can take longer than any single timeout. On CPU-only local inference, set client timeouts well above 2 minutes for chat and plan calls.

An agent run has its own default time limit of 300 seconds.

Agent roster as seen through the API ​

There is no endpoint that lists agents. You see an agent through its agent_type string, which appears in GET /api/v1/agents/executions, in orchestration plan steps, and in workflow task definitions.

agent_typeWhat it doesWhere you meet it in the API
discoveryScans hosts and ports for databases. A step may carry probe_url and probe_key params to run the scan on a remote probe instead of locally.Orchestration step actions scan_network or scan.
analysisSchema analysis, complexity scoring, relationship analysis and documentation generation for several database types.Orchestration actions analyze_schema or analyze. The chat endpoints also report agent_type analysis.
mappingField mapping with semantic similarity and pattern recognition, producing mapping suggestions.Orchestration actions generate_mapping, map or suggest.
securityIdentifies PII and PHI and recommends controls.Orchestration actions scan_database or scan.
validationValidates data and mappings: integrity checks, anomaly detection, business rules, CDC quality.Appears in execution history and as a task agent_type in workflow templates. The orchestrator cannot delegate to it.
executionRuns generated migration scripts: script executor, batch processor, migration validator and monitor.Appears in execution history and as a workflow task agent_type. The orchestrator cannot delegate to it.
orchestrationUses the LLM to break a goal into steps, then runs the steps as a dependency graph, running independent steps concurrently.POST /orchestration/plan and /execute.

The orchestrator only delegates to four agent types: discovery, analysis, mapping and security. A plan step with any other agent_type, or an action not listed above for its agent, makes the job end in FAILED with Unsupported agent/action or Unknown agent type.


Chat endpoints (Workspace Chat tab) ​

The Chat tab in the Workspace page sends planning requests to POST /api/v1/orchestration/plan and all other messages to POST /api/v1/analysis/chat/stream, rendering the events as they arrive. Both use the project ID default.

POST /api/v1/analysis/chat ​

Answers a question about the mmune deployment in one JSON response. A deterministic skill router handles common structured questions without calling an LLM (path: "fast"). Everything else goes to the LLM, which can call read-only data tools (path: "slow").

Permission: read. Action tools that the model may choose to call are gated per caller: they run only if the caller has write or is admin.

Field (body)TypeRequiredDescription
project_idstringyesRequired. The Chat tab sends default.
questionstringyesThe question to answer.
historyarray of {role, content}noEarlier turns, both strings. Default null.

Response 200:

FieldTypeDescription
answerstringAlways present. Plain text.
blocksarray of objectsStructured payloads such as tables. Empty for LLM answers.
console_linksarray of {label, route, tab, focus_node_id}Where in the console to look. tab and focus_node_id may be null.
pathstringfast for the deterministic path, slow for the LLM path or the no-model message.
used_toolsarray of stringsData tools consulted.

Status codes: 200, plus the common 401, 403, 422. Provider and tool failures do not return 500. A failed chat call comes back as HTTP 200 with a degraded answer, for example the fixed text I encountered an error while processing your question., so read the text rather than relying on the status code.

bash
curl -sS -X POST "https://mmune.example.com/api/v1/analysis/chat" \
  -H "Authorization: Bearer <TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"project_id": "default", "question": "Which integrations are degraded right now?"}'

POST /api/v1/analysis/chat/stream ​

Same inputs as /chat, answered as a text/event-stream. Each frame is a line data: <json> followed by a blank line. Permission: read, with the same per-caller gating of action tools.

Request body: identical to /chat.

Event types:

typeOther fieldsMeaning
metaused_tools (array), path (fast or slow)First frame.
deltatext (string)A piece of the answer. Append to what you have. On the slow path expect the answer as one delta rather than token by token.
blockblock (object)A structured payload such as a table. Zero or more.
donepath, console_linksLast frame on success.
errormessage (string)Terminal failure. The HTTP status is still 200 because the stream has started.

Status codes: 200 for any stream, plus the common 401, 403, 422.

bash
curl -sS -N -X POST "https://mmune.example.com/api/v1/analysis/chat/stream" \
  -H "Authorization: Bearer <TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"project_id": "default", "question": "Summarise open drift alerts"}'

Agents router (/api/v1/agents) ​

These three routes are read-only views over the agent execution history. They need the read permission.

GET /api/v1/agents/executions ​

Lists agent executions, filtered by the query parameters.

Permission: read.

Field (query)TypeRequiredDefaultDescription
agent_typestringnononeFilter by agent type, for example discovery.
statusstringnononestarted, completed or failed.
project_idstringnononeFilter by project.
user_idstringnononeFilter by the user who triggered it.
daysintegerno7Look back this many days. A value of 0 or less removes the date filter.
limitintegerno1001 to 1000.
offsetintegerno00 or more. Use with limit to page.

Response 200 is a JSON array. Each item:

FieldTypeDescription
idstringExecution ID, usable in the detail and state endpoints.
agent_idstringAgent instance ID.
agent_typestringFor example discovery, analysis, mapping.
agent_namestringConfigured agent name.
statusstringstarted, completed or failed.
context_id, project_id, session_id, user_idstring or nullCorrelation fields.
started_atstringISO 8601 timestamp, empty string if missing.
completed_atstring or nullISO 8601 timestamp.
duration_msinteger or nullWall-clock duration.
error_messagestring or nullSet on failure.
retry_countintegerRetries attempted.

Status codes: 200.

bash
curl -sS "https://mmune.example.com/api/v1/agents/executions?agent_type=discovery&status=failed&days=1&limit=20" \
  -H "Authorization: Bearer <TOKEN>"

GET /api/v1/agents/executions/ ​

Returns the full record for one execution, including input and output data.

Permission: read.

Field (path)TypeRequiredDescription
execution_idstringyesThe id from the list endpoint.

Response 200 fields: id, agent_id, agent_type, agent_name, status, context_id, project_id, session_id, user_id, input_data, output_data, error_message, error_type, started_at, completed_at, duration_ms, retry_count, metrics, created_at. Timestamps are ISO 8601 strings or null.

Status codes: 200, 404 Execution not found.

bash
curl -sS "https://mmune.example.com/api/v1/agents/executions/<EXECUTION_ID>" \
  -H "Authorization: Bearer <TOKEN>"

GET /api/v1/agents/statistics ​

Aggregates executions over a window.

Permission: read.

Field (query)TypeRequiredDefaultDescription
agent_typestringnoall typesRestrict to one agent type.
daysintegerno7Window length in days.

Response 200:

FieldTypeDescription
total_executionsintegerExecutions in the window.
successfulintegerStatus completed.
failedintegerStatus failed.
average_duration_msintegerMean of executions that have a duration; 0 if none.
success_ratenumbersuccessful / total_executions, 0.0 when there are none.

Status codes: 200.

bash
curl -sS "https://mmune.example.com/api/v1/agents/statistics?agent_type=mapping&days=30" \
  -H "Authorization: Bearer <TOKEN>"

Orchestration router (/api/v1/orchestration) ​

This is the flow behind the Workspace Workflow tab and the "plan" path of Chat: generate a plan from a goal, review it, execute it as a background job, poll the job, download the results. See LLM provider dependence for the 503 behavior and for which agents and actions a plan may use.

POST /api/v1/orchestration/plan ​

Generates a plan with the LLM. The plan is returned to you and is not stored or executed.

Permission: read (this route is on the read-sufficient list because it only plans).

Field (body)TypeRequiredDefaultDescription
goalstringyesnoneWhat you want done, in plain language.
project_idstringyesnoneProject the plan belongs to.
session_idstringnogenerated UUIDCorrelation ID.

Response 200 (the plan object):

FieldTypeDescription
plan_idstringUUID.
goalstringEcho of the goal.
stepsarrayOrdered steps, see below.
created_atstringISO 8601 timestamp.
statusstringcreated.

Each step has step_id (string; the model's id if it gave one, else a UUID), description, agent_type, action, params (object), dependencies (array of step_id values that must finish first), status (pending) and result (null).

Status codes: 200, 503 no real AI provider configured, 500 on a provider error or when the model returns no usable steps (Planner produced no executable steps). The model, not mmune, chooses the steps, so review agent_type and action before you execute.

bash
curl -sS -X POST "https://mmune.example.com/api/v1/orchestration/plan" \
  -H "Authorization: Bearer <TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"goal": "Discover databases on host 10.0.0.12 and analyse their schemas", "project_id": "default"}'

POST /api/v1/orchestration/execute ​

Starts a plan as a background job and returns immediately. The job runs in the background.

Permission: write.

Field (body)TypeRequiredDefaultDescription
planobjectyesnoneA plan as returned by /plan. It must contain plan_id and goal. steps is read with a default of an empty list, so a plan with no steps key is accepted and the job completes with status COMPLETED and empty results. Each step must contain description, agent_type, action, params, dependencies and step_id (or id).
project_idstringyesnoneProject ID.
session_idstringnogenerated UUIDCorrelation ID.

A request without plan or project_id returns 422. If the plan itself is incomplete, for example a step is missing a required field, the job ends as FAILED with the reason in error and in the logs.

Response 200:

FieldTypeDescription
messagestringPlan execution started.
execution_idstringJob ID, a UUID string. Use it for status and download.
statusstringRUNNING.

Status codes: 200, 422, 503 no real AI provider configured.

bash
curl -sS -X POST "https://mmune.example.com/api/v1/orchestration/execute" \
  -H "Authorization: Bearer <TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"plan": {"plan_id": "6f1c2f0e-0000-4000-8000-000000000001", "goal": "Scan host", "steps": [{"step_id": "step_1", "description": "Scan for databases", "agent_type": "discovery", "action": "scan_network", "params": {"hosts": ["10.0.0.12"]}, "dependencies": []}]}, "project_id": "default"}'

GET /api/v1/orchestration/status/ ​

Returns job state and new log lines. Poll this until status is COMPLETED or FAILED.

Permission: read.

FieldInTypeRequiredDefaultDescription
execution_idpathstringyesnoneJob ID from /execute.
cursorqueryintegerno0Log position. Pass the previous next_cursor to receive only new lines. Must be 0 or more.

Response 200:

FieldTypeDescription
statusstringRUNNING, COMPLETED or FAILED.
logsarray of stringsLog lines from cursor onward.
next_cursorintegerCursor for the next poll.
resultsobject or nullUpdated after each finished step: steps (full step objects) and results (map of step_id to step output). On completion it also carries plan_id, goal, execution_id and completed_at.
errorstring or nullSet when FAILED.
created_at, updated_atstringTimestamps.
metadataobjectplan_id, project_id, session_id.

Status codes: 200, 404 Execution ID not found.

bash
curl -sS "https://mmune.example.com/api/v1/orchestration/status/<EXECUTION_ID>?cursor=0" \
  -H "Authorization: Bearer <TOKEN>"

GET /api/v1/orchestration/artifacts/{execution_id}/download ​

Downloads the results of a completed job as a JSON attachment.

Permission: read.

FieldInTypeRequiredDefaultDescription
execution_idpathstringyesnoneJob ID.
formatquerystringnojsonOnly json is supported.

Response 200 is application/json with Content-Disposition: attachment; filename=orchestration-artifacts-<id>-<YYYYMMDD>.json. Body fields: execution_id, status, created_at, updated_at, metadata, artifacts (the job's result object).

Status codes: 200, 400 when the job is not COMPLETED or the format is not json, 404 when the job is unknown or has no stored result.

bash
curl -sS -OJ "https://mmune.example.com/api/v1/orchestration/artifacts/<EXECUTION_ID>/download?format=json" \
  -H "Authorization: Bearer <TOKEN>"

Workflows router (/api/v1/workflows) ​

A workflow is a named set of tasks with dependencies. Each task names an agent_type. The routes below manage workflow definitions and templates. Work is run through the orchestration endpoints (plan, execute, status, download).

POST /api/v1/workflows/register ​

Registers a workflow definition with the engine. The definition is validated for missing dependencies and cycles.

Permission: write.

Field (body)TypeRequiredDefaultDescription
workflow_idstringyesnoneYour identifier. Registering the same ID again replaces the definition.
namestringyesnoneDisplay name.
descriptionstringyesnoneDescription.
tasksarray of objectsyesnoneTask definitions, below.
initial_dataobjectno{}Initial workflow data.
versionstringno1.0Definition version.

Each task object: task_id (required), name (required), agent_type (required), depends_on (array of task_id, default []), input_mapping (object, default {}), output_mapping (object, default {}), retry_count (integer, default 0), timeout_seconds (integer, default 300), optional (boolean, default false). Other keys are ignored. A task missing a required key returns 500.

Response 200: {"success": true, "message": "Workflow '<name>' registered successfully", "workflow_id": "<id>"}.

Status codes: 200, 400 when validation finds a missing dependency or a cycle, 500 otherwise.

bash
curl -sS -X POST "https://mmune.example.com/api/v1/workflows/register" \
  -H "Authorization: Bearer <TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"workflow_id": "nightly-scan", "name": "Nightly scan", "description": "Discover then analyse", "tasks": [{"task_id": "discover", "name": "Discover", "agent_type": "discovery"}, {"task_id": "analyse", "name": "Analyse", "agent_type": "analysis", "depends_on": ["discover"]}]}'

Workflow templates router ​

Templates are predefined. Creating from a template builds a new definition with a fresh UUID workflow_id and registers it with the engine.

Available templates:

template_idNameTasksRequired body fieldsTask agent types, in order
full_migrationFull Migration6source_db, target_db, project_namediscovery, analysis, mapping, validation, execution, qa
schema_analysisSchema Analysis3source_db, project_namediscovery, analysis, documentation
data_validationData Validation3source_db, target_db, project_nameqa, validation, documentation

GET /api/v1/workflows/templates ​

Lists templates.

Permission: read.

Response 200: {"success": true, "count": 3, "templates": {"<template_id>": {"name", "description", "steps"}}}. In this response steps is the number of tasks.

Status codes: 200.

bash
curl -sS "https://mmune.example.com/api/v1/workflows/templates" \
  -H "Authorization: Bearer <TOKEN>"

GET /api/v1/workflows/templates/ ​

Returns the summary entry for one template, the same object shown in the list.

Permission: read.

Field (path)TypeRequiredDescription
template_idstringyesfull_migration, schema_analysis or data_validation.

Response 200: {"success": true, "template_id": "<id>", "template": {"name", "description", "steps"}}.

Status codes: 200, 404 Template not found.

bash
curl -sS "https://mmune.example.com/api/v1/workflows/templates/schema_analysis" \
  -H "Authorization: Bearer <TOKEN>"

POST /api/v1/workflows/from-template ​

Builds a workflow from a template and registers it.

Permission: write.

Field (body)TypeRequiredDefaultDescription
template_idstringyesnoneOne of the template IDs above.
project_namestringyesnoneUsed in the workflow name, for example <project_name> - Schema Analysis.
source_dbstringyesnoneSource connection string, copied into task config.
target_dbstringconditionalnullRequired for full_migration and data_validation.

Connection strings you send are stored in the task config of the registered definition. Use a least-privilege read account, and avoid putting production passwords in a URL you do not want logged client side.

Response 200:

FieldTypeDescription
successbooleantrue.
messagestringWorkflow created from template: <template_id>.
workflow_idstringNew UUID.
workflowobjectworkflow_id, name, description, version, initial_data, and tasks where each task has task_id, name, step_type, agent_type, depends_on, config, retry_count, timeout_seconds, optional.

Status codes: 200, 400 when target_db is missing for a template that needs it, 404 Template not found: <id>, 500.

bash
curl -sS -X POST "https://mmune.example.com/api/v1/workflows/from-template" \
  -H "Authorization: Bearer <TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"template_id": "schema_analysis", "project_name": "CRM Docs", "source_db": "postgresql://reader@db.example.com:5432/crm"}'

Cancellation and state routers ​

These routes act on agent execution IDs (the id values from /agents/executions), not on orchestration job IDs.

Cancellation is best effort.

State checkpoints are stored per execution. If none exist for an ID, the state routes report not found or cannot resume.

POST /api/v1/agent/cancel/ ​

Requests cancellation of an agent execution.

Permission: write.

FieldInTypeRequiredDefaultDescription
execution_idpathstringyesnoneAgent execution ID.
reasonbodystringnoCancelled by userRecorded as the cancellation reason.

The body object is required even when empty; send {} to accept the default reason.

Response 200: {"success": true, "message": "Execution cancelled: <reason>", "execution_id": "<id>", "cancelled": true}. Calling it twice is safe; the first reason is kept.

Status codes: 200, 500.

bash
curl -sS -X POST "https://mmune.example.com/api/v1/agent/cancel/<EXECUTION_ID>" \
  -H "Authorization: Bearer <TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"reason": "Wrong source selected"}'

GET /api/v1/agent/cancel-status/ ​

Reads the cancellation flag.

Permission: read.

Response 200: {"execution_id": "<id>", "is_cancelled": false, "reason": null}. reason is a string once cancelled.

Status codes: 200, 500.

bash
curl -sS "https://mmune.example.com/api/v1/agent/cancel-status/<EXECUTION_ID>" \
  -H "Authorization: Bearer <TOKEN>"

GET /api/v1/agent/state/ ​

Returns the latest checkpoint for an execution.

Permission: read.

Response 200:

FieldTypeDescription
successbooleantrue.
messagestringState retrieved successfully.
stateobjectstate_id, execution_id, agent_name, agent_type, status, checkpoint_number, state_data, context_data, progress_data, created_at, updated_at, resumed_count, last_error, error_count, can_resume, resume_message.

Status codes: 200, 404 State not found, 500.

bash
curl -sS "https://mmune.example.com/api/v1/agent/state/<EXECUTION_ID>" \
  -H "Authorization: Bearer <TOKEN>"

GET /api/v1/agent/state/{execution_id}/can-resume ​

Tells you whether a resume would be accepted. An execution can resume when a checkpoint exists, its can_resume flag is true, and its status is neither completed nor cancelled.

Permission: read.

Response 200: {"execution_id": "<id>", "can_resume": true}.

Status codes: 200, 500.

bash
curl -sS "https://mmune.example.com/api/v1/agent/state/<EXECUTION_ID>/can-resume" \
  -H "Authorization: Bearer <TOKEN>"

POST /api/v1/agent/state/{execution_id}/resume ​

Records a resume: increments resumed_count, sets status active, and returns the checkpoint. It does not itself start any work.

Permission: write.

Response 200: same shape as the get-state response, with message of Execution resumed (resume count: <n>).

Status codes: 200, 400 Execution cannot be resumed (completed, cancelled, or no state found), 500.

bash
curl -sS -X POST "https://mmune.example.com/api/v1/agent/state/<EXECUTION_ID>/resume" \
  -H "Authorization: Bearer <TOKEN>"

DELETE /api/v1/agent/state/ ​

Deletes every checkpoint for the execution. It returns success even if there were none.

Permission: write.

Response 200: {"success": true, "message": "State checkpoints deleted", "state": null}.

Status codes: 200, 500.

bash
curl -sS -X DELETE "https://mmune.example.com/api/v1/agent/state/<EXECUTION_ID>" \
  -H "Authorization: Bearer <TOKEN>"

RAG router (/api/v1/rag) ​

Documents are split into chunks, embedded with an embedding-capable provider, and stored in the mmune database. Every user of a deployment sees the same document set.

Retrieval uses cosine similarity. Chunks below min_similarity are dropped. If a chunk was embedded by a different provider than the one now answering, retrieval still runs, but quality may be worse. Re-ingest after you change embedding providers.

POST /api/v1/rag/ingest ​

Ingests text. Re-ingesting under the same name replaces that document's chunks.

Permission: write.

Field (body)TypeRequiredDefaultDescription
namestringyesnoneUnique document name; used as the key for updates and deletes.
contentstringyesnoneRaw text.
metadataobjectnonullStored with each chunk.
chunk_sizeintegerno1000100 to 10000 characters.
chunk_overlapintegerno1000 to 2000 characters.

Response 200: {"status": "success", "document_name": "<name>", "chunks_created": <n>}. Ingestion makes one embedding call per chunk, so long documents on a slow local model take time.

Status codes: 200, 422 for no embedding provider or invalid content, 500.

bash
curl -sS -X POST "https://mmune.example.com/api/v1/rag/ingest" \
  -H "Authorization: Bearer <TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"name": "failover-runbook", "content": "If the primary database is unreachable, check the replica lag before promoting.", "metadata": {"owner": "dba-team"}}'

POST /api/v1/rag/ingest/file ​

Ingests an uploaded file as multipart/form-data. Plain text, Markdown, CSV and JSON are decoded as UTF-8 (invalid bytes are replaced). PDF and DOCX uploads need the optional document parsers to be installed in your deployment; without them the route returns 422.

Permission: write.

FieldInTypeRequiredDefaultDescription
fileformfileyesnoneThe upload.
namequerystringnothe filenameDocument name.
chunk_sizequeryintegerno1000100 to 10000.
chunk_overlapqueryintegerno1000 to 2000.

Response 200: same as /ingest. The chunk metadata records source_filename and content_type.

Status codes: 200, 422 when text cannot be extracted, the file has no text, or no embedding provider is available, 500.

bash
curl -sS -X POST "https://mmune.example.com/api/v1/rag/ingest/file?name=failover-runbook" \
  -H "Authorization: Bearer <TOKEN>" \
  -F "file=@./failover-runbook.md"

POST /api/v1/rag/query ​

Retrieves matching chunks and asks the LLM to answer using only that context. Temperature is 0.

Permission: read.

Field (body)TypeRequiredDefaultDescription
querystringyesnoneThe question.
providerstringnoregistry choicePreferred provider name for generation.
modelstringnoRAG_GENERATION_MODEL, else the default provider's model, else gpt-4Model ID for generation.
limitintegerno5Chunks to retrieve, 1 to 50.
min_similaritynumberno0.70.0 to 1.0.

Response 200:

FieldTypeDescription
answerstringGenerated answer. The prompt tells the model to say it does not know if the context lacks the answer.
sourcesarray of {document_name, similarity, chunk_index}Chunks that were used. Empty when nothing cleared min_similarity.
providerstring or nullProvider that generated the answer.
modelstringModel used.

Status codes: 200, 500 on generation or provider failure. An empty sources array with a "don't know" answer is a normal 200, not an error.

bash
curl -sS -X POST "https://mmune.example.com/api/v1/rag/query" \
  -H "Authorization: Bearer <TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"query": "What do I check before promoting the replica?", "limit": 5, "min_similarity": 0.7}'

GET /api/v1/rag/context ​

Returns matching chunks without generating an answer. Use it to see what retrieval finds, or to feed your own model.

Permission: read.

Field (query)TypeRequiredDefaultDescription
querystringyesnoneText to search for.
limitintegerno51 to 50.
min_similaritynumberno0.70.0 to 1.0.

Response 200 is an array of chunks with chunk_id, document_name, document_type, source_url, content, chunk_index, metadata, embedding_provider, embedding_model, embedding_dimension, created_at, updated_at, and similarity (number).

Status codes: 200, 500.

bash
curl -sS -G "https://mmune.example.com/api/v1/rag/context" \
  --data-urlencode "query=replica lag before promotion" \
  -d "limit=3" \
  -H "Authorization: Bearer <TOKEN>"

GET /api/v1/rag/documents ​

Lists ingested documents. The list is not paginated.

Permission: read.

Response 200 is an array of {document_name, chunk_count, ingested_at, updated_at, embedding_provider, embedding_model}, ordered by name. A document embedded by more than one provider or model appears once per combination.

Status codes: 200.

bash
curl -sS "https://mmune.example.com/api/v1/rag/documents" \
  -H "Authorization: Bearer <TOKEN>"

GET /api/v1/rag/documents/ ​

Status of one document.

Permission: read.

Response 200: {document_name, chunk_count, ingested_at, updated_at, embedding_provider, embedding_model, embedding_dimension}. chunk_count sums across providers; the provider, model and dimension fields come from the first group.

Status codes: 200, 404 Document '<name>' not found.

bash
curl -sS "https://mmune.example.com/api/v1/rag/documents/failover-runbook" \
  -H "Authorization: Bearer <TOKEN>"

DELETE /api/v1/rag/documents/ ​

Deletes a document and all its chunks.

Permission: write.

Response 200: {"status": "deleted", "document_name": "<name>", "chunks_deleted": <n>}.

Status codes: 200, 404 Document '<name>' not found.

bash
curl -sS -X DELETE "https://mmune.example.com/api/v1/rag/documents/failover-runbook" \
  -H "Authorization: Bearer <TOKEN>"

Client examples ​

All examples need a token. POST /api/v1/auth/login takes {"email", "password"} and returns access_token. Accounts with MFA enabled receive a challenge instead; see API overview.

Shared helpers ​

Python (requests):

python
import requests

BASE = "https://mmune.example.com"


def login(email: str, password: str) -> str:
    r = requests.post(
        f"{BASE}/api/v1/auth/login",
        json={"email": email, "password": password},
        timeout=30,
    )
    r.raise_for_status()
    return r.json()["access_token"]


def headers(token: str) -> dict:
    return {"Authorization": f"Bearer {token}"}

TypeScript (fetch, Node 18 or later, or any browser):

typescript
const BASE = "https://mmune.example.com";

async function login(email: string, password: string): Promise<string> {
  const r = await fetch(`${BASE}/api/v1/auth/login`, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ email, password }),
  });
  if (!r.ok) throw new Error(`login failed: ${r.status} ${await r.text()}`);
  return (await r.json()).access_token as string;
}

async function api<T>(
  token: string,
  method: string,
  path: string,
  body?: unknown,
  timeoutMs = 30_000,
): Promise<T> {
  const r = await fetch(`${BASE}${path}`, {
    method,
    headers: {
      Authorization: `Bearer ${token}`,
      ...(body === undefined ? {} : { "Content-Type": "application/json" }),
    },
    body: body === undefined ? undefined : JSON.stringify(body),
    signal: AbortSignal.timeout(timeoutMs),
  });
  if (!r.ok) throw new Error(`${method} ${path} -> ${r.status} ${await r.text()}`);
  return (await r.json()) as T;
}

1. Ask the orchestrator a question through chat ​

This calls POST /api/v1/analysis/chat, the same agent the Chat tab uses for questions. Set the timeout above your provider's timeout, since the slow path can make several model calls. Check path and used_tools to see whether you got a deterministic answer or an LLM answer, and treat a slow-path answer that starts with "No AI model is configured" as the degraded case described earlier.

Python:

python
def ask(token: str, question: str, history: list | None = None) -> dict:
    r = requests.post(
        f"{BASE}/api/v1/analysis/chat",
        headers=headers(token),
        json={"project_id": "default", "question": question, "history": history},
        timeout=180,
    )
    r.raise_for_status()
    return r.json()


token = login("you@example.com", "<PASSWORD>")
reply = ask(token, "Which integrations are degraded right now?")
print(reply["answer"])
print("path:", reply["path"], "tools:", reply["used_tools"])
for link in reply["console_links"]:
    print("see:", link["label"], link["route"], link.get("tab"))

TypeScript:

typescript
interface ChatReply {
  answer: string;
  blocks: unknown[];
  console_links: { label: string; route: string; tab?: string | null; focus_node_id?: string | null }[];
  path: "fast" | "slow";
  used_tools: string[];
}

const token = await login("you@example.com", "<PASSWORD>");
const reply = await api<ChatReply>(
  token,
  "POST",
  "/api/v1/analysis/chat",
  { project_id: "default", question: "Which integrations are degraded right now?" },
  180_000,
);
console.log(reply.answer, reply.path, reply.used_tools);

If you want the plan flow instead, call POST /api/v1/orchestration/plan with {"goal", "project_id"} and read steps from the response.

2. Query the RAG endpoint ​

Ingest a document, then ask a question about it. Both calls depend on a configured provider: ingestion needs an embedding-capable one, and the answer needs a generation provider.

Python:

python
def rag_demo(token: str) -> dict:
    h = headers(token)

    ing = requests.post(
        f"{BASE}/api/v1/rag/ingest",
        headers=h,
        json={
            "name": "failover-runbook",
            "content": "If the primary is unreachable, check replica lag before promoting.",
        },
        timeout=300,  # one embedding call per chunk
    )
    ing.raise_for_status()

    q = requests.post(
        f"{BASE}/api/v1/rag/query",
        headers=h,
        json={"query": "What do I check before promoting the replica?", "limit": 5},
        timeout=180,
    )
    q.raise_for_status()
    body = q.json()
    if not body["sources"]:
        print("No chunk cleared min_similarity; the answer is not grounded in your documents.")
    print(body["answer"])
    for s in body["sources"]:
        print(s["document_name"], s["chunk_index"], round(s["similarity"], 3))
    return body

TypeScript:

typescript
interface RagAnswer {
  answer: string;
  sources: { document_name: string; similarity: number; chunk_index: number }[];
  provider: string | null;
  model: string;
}

async function ragDemo(token: string): Promise<RagAnswer> {
  await api(
    token,
    "POST",
    "/api/v1/rag/ingest",
    {
      name: "failover-runbook",
      content: "If the primary is unreachable, check replica lag before promoting.",
    },
    300_000,
  );

  const body = await api<RagAnswer>(
    token,
    "POST",
    "/api/v1/rag/query",
    { query: "What do I check before promoting the replica?", limit: 5 },
    180_000,
  );
  if (body.sources.length === 0) {
    console.warn("No chunk cleared min_similarity; the answer is not grounded in your documents.");
  }
  console.log(body.answer, body.sources);
  return body;
}

API overview for authentication, errors, rate limits and pagination. WebSockets and streaming for live progress streams.