Monitoring & Debugging
When a workflow misbehaves in production, you need a fast, repeatable path from "something is wrong" to "here is the failing step, here is why, here is the fix." This page is that runbook. It assumes you already understand the workflow engine, runners, and Lumen observability at the concept level — here we focus on the operational moves.
The two instruments you reach for, in order:
The debugging decision tree
Use this to pick your first move. Each leaf links to the matching section below.
- What's wrong?
- A run is still goingand I want to follow it
- A run failedand ended
- A run is stucknot progressing
- Many runs are slowor failing at once
- Runs depending on aconnection/runner fail
- watch_workflow → live transitions
- get_workflow_history → find failed statethen Lumen error group + trace
- list_stuck_workflows → retry_workflowforce_terminate as last resort
- Lumen dashboard → top error groupslatency/cost pulse
- Check connection / runner health+ SLA Health policies
- What's wrong?→A run is still going
- What's wrong?→A run failed
- What's wrong?→A run is stuck
- What's wrong?→Many runs are slow
- What's wrong?→Runs depending on a
- A run is still going→watch_workflow → live transitions
- A run failed→get_workflow_history → find failed state
- A run is stuck→list_stuck_workflows → retry_workflow
- Many runs are slow→Lumen dashboard → top error groups
- Runs depending on a→Check connection / runner health
| Symptom | First move | Section |
|---|---|---|
| Run in progress, want to follow it | watch_workflow | Watch a live run |
Run ended in FAILED | get_workflow_history, then Lumen trace | Read history & debug a failure |
| Run hasn't moved in a while | list_stuck_workflows → retry_workflow | Find & recover stuck workflows |
| Broad spike of failures/latency | Lumen dashboard → error groups | Debug with Lumen traces & logs |
| Failures tied to a cloud/runner | Connection + runner + SLA health | Check runner & SLA health |
Identify yourself first
Every run is scoped to your organization, resolved server-side from your token. Before any operation, confirm who you are — this is the mandatory first step for the workflow MCP and surfaces the org your run-state queries will be scoped to.
# via the MCP tool surface (whoami)
orkestia whoami
# → { organization_uuid: "...", scopes: ["workflow:read", "workflow:write", ...] }
organization_uuid in initial_data. The engine derives it from your token and scopes every list/watch/retry call automatically. If whoami shows the wrong org or missing scopes, fix auth before debugging — you may simply be looking at the wrong tenant's runs.Watch a live run
When a run is in progress and you want to follow its state machine in real time, watch_workflow streams transitions as they happen. This is the cheapest way to confirm a long-running DAG is advancing rather than wedged.
watch_workflow(workflow_id="wf_01J...")
You'll see each transition as the engine commits it — state name, the transition that fired, and any emitted output. For a virtual workflow / composition you'll see the constituent steps advance in dependency order. If the stream goes quiet for longer than the step's expected duration, the run is a stuck-workflow candidate — jump to recovery.
watch_workflow re-attaches to the same workflow_id.Read history and debug a failure
When a run has ended in a failure, get_workflow_history gives you the full ordered transition log — every state the run entered, the transition that moved it, timestamps, and the error captured at the failing state. Pair it with get_workflow_status for the one-line current verdict.
get_workflow_status(workflow_id="wf_01J...")
# → { state: "FAILED", failed_at: "...", error: "..." }
get_workflow_history(workflow_id="wf_01J...")
# → ordered list of { state, transition, entered_at, output | error }
Read the history bottom-up: the last non-terminal state before FAILED is where the work actually broke. Atomic workflows follow the 3-state pattern — PENDING → COMPLETED | FAILED — so a failure is unambiguous about which operation failed. For a DAG, the history tells you which node failed and which upstream nodes had already completed (so you know what not to redo).
Once you have the failing step and its error string, cross to Lumen for the root cause.
sequenceDiagram participant You participant Engine as Workflow engine participant Lumen You->>Engine: get_workflow_history(workflow_id) Engine-->>You: failing state + error string You->>Lumen: search logs by workflow_id / trace_id Lumen-->>You: grouped error + trace + prior resolution You->>Engine: retry_workflow(workflow_id)
Find and recover stuck workflows
A stuck workflow is one that has stopped progressing without reaching a terminal state — typically an auto_advance step that never fired, a step waiting on a dependency that will never arrive, or a worker that died mid-step. The engine surfaces these directly.
Step 1 — List them
list_stuck_workflows()
# → [ { workflow_id, workflow_type, current_state, stuck_since, ... }, ... ]
This returns runs the engine considers stalled for your org. Triage by workflow_type and stuck_since — a cluster of the same type stuck at the same state points to a systemic cause (bad input shape, a downstream connection outage) rather than one unlucky run.
Step 2 — Retry
For the common case — a transient failure (rate limit, brief connection blip, a runner that was reaped mid-job) — retry resumes the run from where it failed rather than starting over. Because the engine is event-sourced, retry replays from the last committed state, so already-completed steps are not repeated.
retry_workflow(workflow_id="wf_01J...")
Step 2b — Resolve a remediation gate
Not every non-terminal run is stuck. A DAG step that fails on a fixable precondition parks the run in remediation_pending instead of compensating — the run is waiting for you. Its state_data.remediation envelope (read it via get_workflow_status) names exactly what to fix, typically a fix workflow and its inputs. Apply the fix, then:
resolve_workflow(workflow_id="wf_01J...", resolution="remediated")
# the engine re-runs ONLY the failed step; completed steps keep their outputs
resolve_workflow(workflow_id="wf_01J...", resolution="denied")
# no fix — deferred compensation runs and the run fails terminally
Find gated runs with list_workflows filtered to state_name="remediation_pending".
Step 3 — Force-terminate (last resort)
If a run is genuinely unrecoverable — corrupt input, a step waiting on something that no longer exists, or a duplicate you need to clear — force_terminate_workflow drives it to a terminal state so it stops consuming queue and lock resources.
force_terminate_workflow(workflow_id="wf_01J...")
| Tool | Use when | Reversible? |
|---|---|---|
retry_workflow | Transient failure; want to resume from the failed step | Yes — re-runnable |
resolve_workflow | Run parked on the remediation gate; fix applied (or denied) | remediated resumes; denied compensates + fails |
force_terminate_workflow | Unrecoverable run; need to free resources | No — terminal |
Debug with Lumen traces and logs
Lumen is a separate API (https://lumen-api.orkestia.dev). Filter on fields you put on ingest (workflow_id, trace_id, project, level, …) — they are query params on GET /api/logs, not reconstructed later. Contract: Send data, Query API. App mapping: Use Lumen.
trace_id/span_id. Where those are missing, Lumen shows pseudo-traces grouped by workflow/error/fingerprint rather than true distributed spans.The triage path
1. Dashboard
Start at the project dashboard: top error groups, event feed, log pulse, and latency/cost signals. A broad spike shows here first.
2. Error groups
Failures are fingerprinted and grouped. Open the group to see occurrence count, severity, affected workflows, and whether it's new, regressed, or known.
Connecting the engine to Lumen
The link between the two instruments is identity. Carry these from get_workflow_history into Lumen's log filters:
| From the engine | Filter in Lumen | Gets you |
|---|---|---|
workflow_id | log search by workflow | Every log line that run emitted |
trace_id (if emitted) | trace view | The end-to-end span timeline |
failing state name | log filter by state + error fingerprint | The grouped error and its history |
workflow_type | error groups by type | Whether this is one run or a fleet-wide pattern |
curl -sG "https://lumen-api.orkestia.dev/api/logs" \
--data-urlencode "workflow_id=<uuid-from-the-run>" \
--data-urlencode "level=ERROR" \
--data-urlencode "processed=false" \
-H "Authorization: Bearer $ORKESTIA_TOKEN"
Default window is the last 24h (from/to to widen, max 90 days). limit max 500. Agents use Lumen MCP (mcp-lumen.orkestia.dev) — not the workflow MCP. lumk_ read keys are GET-only; resolve/rules need an org token.
Check runner and SLA health
When failures cluster around a particular cloud, region, or CI pipeline, the problem is often below the workflow — a degraded runner group or a drifted resource, not the workflow logic itself.
Runner health
Runner compute lives in your cloud; Orkestia is the control plane that provisions, registers, scales, and reaps it. When runner-dependent workflows fail:
- Confirm the runner group is
active(not stuck inprovisioningor silentlyarchived) — see Runner management. - Check whether drift detection has flagged the group; a runner that GitHub no longer sees, or compute that was deleted out-of-band in your account, manifests as workflows that can't acquire a runner.
- Verify the underlying cloud connection still grants the trust Orkestia needs to mint registration tokens and scale.
SLA health
For live health signals, rely on Lumen (latency/error pulse, slow-span p95/p99, error-group counts), which is the production observability surface.
Quick reference
| Goal | Tool | Notes |
|---|---|---|
| Confirm identity & org | whoami | Mandatory first step |
| Follow a live run | watch_workflow | Read-only, re-attachable |
| One-line verdict | get_workflow_status | Current state + error |
| Full transition log | get_workflow_history | Read bottom-up to the failing step |
| Find stalled runs | list_stuck_workflows | Triage by type + stuck_since |
| Resume a failed run | retry_workflow | Replays from last committed state |
| Resume a gated run | resolve_workflow | After fixing the named precondition |
| Kill an unrecoverable run | force_terminate_workflow | Destructive, no compensation |
| Root-cause a failure | Lumen logs / error groups / traces | Filter by workflow_id / trace_id |
Where to go next
Lumen concepts
How workflow-native observability is structured — logs, error groups, traces, triage memory.
For per-workflow detail — exact inputs, states, and recovery semantics of any individual workflow type — consult the live catalog at reference.orkestia.dev.
