Multi-Turn AssetOps is an interactive multi-agent system for industrial operations and maintenance (O&M). It preserves tool outputs as structured, reusable evidence so later dialogue turns can continue an investigation without repeating expensive data retrieval and time-series analysis.
The system uses an evidence-aware supervisor and four domain specialists for data collection, time-series analysis, failure reasoning, and maintenance planning. Industrial tools are exposed through Model Context Protocol (MCP) servers and operate over IoT, work-order, failure-mode, and time-series data.
Industrial investigations are iterative. A user may first ask whether a chiller is behaving abnormally, then ask what caused the anomaly, and finally ask which explanation best fits the evidence. A conventional agent can repeat retrieval and analysis on every turn. Multi-Turn AssetOps instead stores each useful result as a typed artifact with:
- a persistent artifact ID;
- asset and time-window metadata;
- the producing specialist and source tools;
- a structured payload and concise summary;
- references to upstream artifacts.
The supervisor inspects these artifacts before every decision. It can reuse existing evidence, dispatch only the specialist needed to fill a gap, or answer directly when the evidence is sufficient.
flowchart LR
U[User request] --> S[Evidence-aware supervisor]
M[Conversation state] --> S
E[(Artifact store)] --> S
S -->|missing raw evidence| D[Data collection]
S -->|missing analytics| T[Time-series analysis]
S -->|missing diagnosis| F[Failure reasoning]
S -->|missing action plan| P[Maintenance planning]
S -->|evidence sufficient| R[Response]
D --> MCP[MCP servers]
T --> MCP
F --> MCP
P --> MCP
MCP -->|typed artifacts| E
E --> R
| Component | Responsibility | Typical artifacts |
|---|---|---|
| Supervisor | Assesses evidence, routes work, and decides when to stop | Routing decisions and final responses |
| Data Collection Specialist | Retrieves assets, sensors, histories, events, and work orders | inventory, dataset_ref |
| Time-Series Analysis Specialist | Computes statistics, anomaly detection, forecasting, and fine-tuning | analysis_result |
| Failure Reasoning Specialist | Maps observed behavior to failure modes and competing diagnoses | reasoning_result |
| Maintenance Planning Specialist | Produces recovery, verification, and maintenance actions | planning_result, draft_action |
The default MCP servers are:
iotfor sites, assets, sensors, and historical measurements;wofor work orders, events, alerts, and failure codes;tsfmfor forecasting, anomaly detection, and TSFM fine-tuning;fmsrfor failure-mode and sensor mappings;utilitiesfor CSV statistics, JSON reading, and date/time utilities.
The paper demonstration follows Chiller 6 at Site MAIN across three turns:
- Detect the problem. The system retrieves sensor data and performs anomaly analysis, creating reusable data and analysis artifacts.
- Explain the problem. The supervisor reuses those artifacts and requests only the missing failure reasoning and supporting statistics.
- Compare the causes. The accumulated evidence already supports the answer, so the supervisor responds without another tool call.
This illustrates the main execution policy: stored evidence guides both specialist selection and termination.
- Python 3.12+
uv- Docker, for the local CouchDB service
- credentials for IBM WatsonX, TokenRouter, or a LiteLLM-compatible endpoint
git clone https://github.com/Coderlicr/Multi-Turn-AssetOps.git
cd Multi-Turn-AssetOps
uv sync
cp .env.example .envConfigure one model provider in .env. For WatsonX:
DEFAULT_MODEL_ID=watsonx/meta-llama/llama-4-maverick-17b-128e-instruct-fp8
WATSONX_APIKEY=your-api-key
WATSONX_PROJECT_ID=your-project-id
WATSONX_URL=https://us-south.ml.cloud.ibm.comtokenrouter/<model> and litellm_proxy/<model> identifiers are also supported through TOKENROUTER_* and LITELLM_* credentials. See .env.example for all settings.
The repository includes work-order data, a Chiller 6 sample, and vibration data. The full evaluation dataset expects main.json under src/couchdb/sample_data/iot/.
docker compose -f src/couchdb/docker-compose.yaml up -d
uv run python src/couchdb/check_couchdb_data.pyThe default databases are chiller, workorder, and vibration at http://localhost:5984.
Run a single turn:
uv run supervisor-specialist \
--reference-date 2020-06-20 \
"Investigate unusually high temperature on Chiller 6 at Site MAIN over the past month."Start a multi-turn session:
uv run supervisor-specialist --multi-turn --reference-date 2020-06-20Example conversation:
> Is Chiller 6 behaving abnormally over the past month?
> What could be causing this during the same period?
> Which cause best fits the data?
Enable independent specialist tool calls in parallel:
uv run supervisor-specialist \
--parallel \
--reference-date 2020-06-20 \
"Compare Chiller 3, Chiller 4, Chiller 6, and Chiller 9 over the past month."Cross-turn state is controlled with --reuse-mode:
| Mode | Behavior |
|---|---|
full |
Reuses dialogue memory and structured artifacts across turns; default |
none |
Keeps natural-language dialogue memory but hides artifacts from earlier turns |
stateless |
Drops all cross-turn state |
Run the no-reuse ablation:
uv run supervisor-specialist \
--multi-turn \
--reuse-mode none \
--reference-date 2020-06-20uv run plan-execute \
--show-plan \
--show-history \
--reference-date 2020-06-20 \
"List the assets available at Site MAIN."The optional Stirrup runner exposes the same default MCP server set. Its code-execution track is enabled by default; use --no-code for tools-only execution.
uv run stirrup-agent \
--no-code \
--reference-date 2020-06-20 \
--model-id tokenrouter/<model-name> \
"List the assets available at Site MAIN."Run all 16 Supervisor–Specialist dialogues:
uv run python eval/run_eval.py \
--system supervisor-specialist \
--reuse-mode fullRun the no-reuse ablation:
uv run python eval/run_eval.py \
--system supervisor-specialist \
--reuse-mode noneRun selected dialogues or override the model:
uv run python eval/run_eval.py \
--dialogs 1 6 12 \
--model-id watsonx/meta-llama/llama-4-maverick-17b-128e-instruct-fp8The harness writes per-dialog responses, LLM/tool/CouchDB metrics, behavior logs, and summary.json under eval/results/<run-id>/. Profiling integrations for Weights & Biases and LangSmith are optional; configuration is documented in PROFILING.md.
Dedicated entry points are also available for the architecture-specific evaluation runners:
uv run python eval/run_plan_execute_eval.py
uv run python eval/run_supervisor_specialist_eval.py.
├── eval/ # 16 dialogue scenarios and evaluation runners
├── src/
│ ├── agent/
│ │ ├── plan_execute/ # Plan–Execute baseline
│ │ ├── stirrup_agent/ # Optional Stirrup runner
│ │ └── supervisor_specialist/ # Supervisor, specialists, graph, and artifact store
│ ├── couchdb/ # Local CouchDB setup and sample data
│ ├── llm/ # LiteLLM backend and provider routing
│ └── servers/ # MCP servers: IoT, WO, TSFM, FMSR, utilities, vibration
├── DESIGN.md # Scenario and system design notes
├── INSTRUCTIONS.md # Detailed usage documentation
├── PROFILING.md # Metrics and profiling setup
└── pyproject.toml # Dependencies and CLI entry points
Released under the Apache License 2.0.