Orkestia
Blog
Operations

Monitoring & Debugging

An operational runbook for watching runs, reading history, recovering stuck workflows, and tracing failures with Lumen

When a workflow misbehaves in production, you need a fast, repeatable path from "something is wrong" to "here is the failing step, here is why, here is the fix." This page is that runbook. It assumes you already understand the workflow engine, runners, and Lumen observability at the concept level — here we focus on the operational moves.

The two instruments you reach for, in order:

The workflow MCP / SDK

Run-state tooling — watch_workflow, get_workflow_history, list_stuck_workflows, retry_workflow. This is the engine's own view of what each run is doing right now.

Lumen

Workflow-native observability — logs, error groups, traces, spans, and triage memory, all attributed by workflow, actor, and state. This is why a run did what the engine reports.

The engine answers what happened (which state, which transition, which error string); Lumen answers why (the trace, the grouped error, the previous resolution). Most debugging sessions bounce between the two. Start with the engine for the failing step, jump to Lumen for the root cause.

The debugging decision tree

Use this to pick your first move. Each leaf links to the matching section below.

SymptomFirst moveSection
Run in progress, want to follow itwatch_workflowWatch a live run
Run ended in FAILEDget_workflow_history, then Lumen traceRead history & debug a failure
Run hasn't moved in a whilelist_stuck_workflows → retry_workflowFind & recover stuck workflows
Broad spike of failures/latencyLumen dashboard → error groupsDebug with Lumen traces & logs
Failures tied to a cloud/runnerConnection + runner + SLA healthCheck runner & SLA health

Identify yourself first

Every run is scoped to your organization, resolved server-side from your token. Before any operation, confirm who you are — this is the mandatory first step for the workflow MCP and surfaces the org your run-state queries will be scoped to.

# via the MCP tool surface (whoami)
orkestia whoami
# → { organization_uuid: "...", scopes: ["workflow:read", "workflow:write", ...] }
Never pass organization_uuid in initial_data. The engine derives it from your token and scopes every list/watch/retry call automatically. If whoami shows the wrong org or missing scopes, fix auth before debugging — you may simply be looking at the wrong tenant's runs.

Watch a live run

When a run is in progress and you want to follow its state machine in real time, watch_workflow streams transitions as they happen. This is the cheapest way to confirm a long-running DAG is advancing rather than wedged.

watch_workflow(workflow_id="wf_01J...")

You'll see each transition as the engine commits it — state name, the transition that fired, and any emitted output. For a virtual workflow / composition you'll see the constituent steps advance in dependency order. If the stream goes quiet for longer than the step's expected duration, the run is a stuck-workflow candidate — jump to recovery.

Watching is read-only and safe to leave running. It reflects committed engine state (the engine uses Postgres advisory locks and an event-sourced log), so what you see is the durable truth, not an optimistic UI guess. A pod restart or HPA scale event will not lose the run — watch_workflow re-attaches to the same workflow_id.

Read history and debug a failure

When a run has ended in a failure, get_workflow_history gives you the full ordered transition log — every state the run entered, the transition that moved it, timestamps, and the error captured at the failing state. Pair it with get_workflow_status for the one-line current verdict.

get_workflow_status(workflow_id="wf_01J...")
# → { state: "FAILED", failed_at: "...", error: "..." }

get_workflow_history(workflow_id="wf_01J...")
# → ordered list of { state, transition, entered_at, output | error }

Read the history bottom-up: the last non-terminal state before FAILED is where the work actually broke. Atomic workflows follow the 3-state pattern — PENDING → COMPLETED | FAILED — so a failure is unambiguous about which operation failed. For a DAG, the history tells you which node failed and which upstream nodes had already completed (so you know what not to redo).

Once you have the failing step and its error string, cross to Lumen for the root cause.

sequenceDiagram
  participant You
  participant Engine as Workflow engine
  participant Lumen
  You->>Engine: get_workflow_history(workflow_id)
  Engine-->>You: failing state + error string
  You->>Lumen: search logs by workflow_id / trace_id
  Lumen-->>You: grouped error + trace + prior resolution
  You->>Engine: retry_workflow(workflow_id)

Find and recover stuck workflows

A stuck workflow is one that has stopped progressing without reaching a terminal state — typically an auto_advance step that never fired, a step waiting on a dependency that will never arrive, or a worker that died mid-step. The engine surfaces these directly.

Step 1 — List them

list_stuck_workflows()
# → [ { workflow_id, workflow_type, current_state, stuck_since, ... }, ... ]

This returns runs the engine considers stalled for your org. Triage by workflow_type and stuck_since — a cluster of the same type stuck at the same state points to a systemic cause (bad input shape, a downstream connection outage) rather than one unlucky run.

Step 2 — Retry

For the common case — a transient failure (rate limit, brief connection blip, a runner that was reaped mid-job) — retry resumes the run from where it failed rather than starting over. Because the engine is event-sourced, retry replays from the last committed state, so already-completed steps are not repeated.

retry_workflow(workflow_id="wf_01J...")
Before retrying en masse, fix the underlying cause. Retrying a run whose input is malformed, or whose dependency is still down, just re-queues the same failure. Confirm the root cause in Lumen first (next section), then retry.

Step 2b — Resolve a remediation gate

Not every non-terminal run is stuck. A DAG step that fails on a fixable precondition parks the run in remediation_pending instead of compensating — the run is waiting for you. Its state_data.remediation envelope (read it via get_workflow_status) names exactly what to fix, typically a fix workflow and its inputs. Apply the fix, then:

resolve_workflow(workflow_id="wf_01J...", resolution="remediated")
# the engine re-runs ONLY the failed step; completed steps keep their outputs
resolve_workflow(workflow_id="wf_01J...", resolution="denied")
# no fix — deferred compensation runs and the run fails terminally

Find gated runs with list_workflows filtered to state_name="remediation_pending".

Step 3 — Force-terminate (last resort)

If a run is genuinely unrecoverable — corrupt input, a step waiting on something that no longer exists, or a duplicate you need to clear — force_terminate_workflow drives it to a terminal state so it stops consuming queue and lock resources.

force_terminate_workflow(workflow_id="wf_01J...")
Force-termination is destructive: it ends the run without completing remaining work and without compensating already-completed steps. Use it only after retry has failed and you've confirmed the run can't be salvaged. For workflows with side effects in your cloud, verify what already executed (via history) before terminating, so you know what state your account is left in.
ToolUse whenReversible?
retry_workflowTransient failure; want to resume from the failed stepYes — re-runnable
resolve_workflowRun parked on the remediation gate; fix applied (or denied)remediated resumes; denied compensates + fails
force_terminate_workflowUnrecoverable run; need to free resourcesNo — terminal

Debug with Lumen traces and logs

Lumen is a separate API (https://lumen-api.orkestia.dev). Filter on fields you put on ingest (workflow_id, trace_id, project, level, …) — they are query params on GET /api/logs, not reconstructed later. Contract: Send data, Query API. App mapping: Use Lumen.

Lumen is alpha. The self-serve onboarding/control plane (org provisioning, plan, retention) ships, but standalone packaging is still maturing, and trace fidelity depends on workflows emitting trace_id/span_id. Where those are missing, Lumen shows pseudo-traces grouped by workflow/error/fingerprint rather than true distributed spans.

The triage path

1. Dashboard

Start at the project dashboard: top error groups, event feed, log pulse, and latency/cost signals. A broad spike shows here first.

2. Error groups

Failures are fingerprinted and grouped. Open the group to see occurrence count, severity, affected workflows, and whether it's new, regressed, or known.

3. Traces & spans

Drill into the trace timeline for the failing run — slow spans, error spans, and latency stats reveal whether the failure is logic, a dependency, or a timeout.

4. Triage memory

Lumen surfaces previous resolutions for similar fingerprints. If this error has been fixed before, you reuse the fix instead of re-diagnosing.

Connecting the engine to Lumen

The link between the two instruments is identity. Carry these from get_workflow_history into Lumen's log filters:

From the engineFilter in LumenGets you
workflow_idlog search by workflowEvery log line that run emitted
trace_id (if emitted)trace viewThe end-to-end span timeline
failing state namelog filter by state + error fingerprintThe grouped error and its history
workflow_typeerror groups by typeWhether this is one run or a fleet-wide pattern
curl -sG "https://lumen-api.orkestia.dev/api/logs" \
  --data-urlencode "workflow_id=<uuid-from-the-run>" \
  --data-urlencode "level=ERROR" \
  --data-urlencode "processed=false" \
  -H "Authorization: Bearer $ORKESTIA_TOKEN"

Default window is the last 24h (from/to to widen, max 90 days). limit max 500. Agents use Lumen MCP (mcp-lumen.orkestia.dev) — not the workflow MCP. lumk_ read keys are GET-only; resolve/rules need an org token.

Drop/sample ingest rules can hide signal before it reaches the dashboard, and over-broad fingerprint rules can merge unrelated failures. If a failure you can see in engine history has no matching Lumen logs, check the project's ingest rules before assuming the instrumentation is broken.

Check runner and SLA health

When failures cluster around a particular cloud, region, or CI pipeline, the problem is often below the workflow — a degraded runner group or a drifted resource, not the workflow logic itself.

Runner health

Runner compute lives in your cloud; Orkestia is the control plane that provisions, registers, scales, and reaps it. When runner-dependent workflows fail:

  • Confirm the runner group is active (not stuck in provisioning or silently archived) — see Runner management.
  • Check whether drift detection has flagged the group; a runner that GitHub no longer sees, or compute that was deleted out-of-band in your account, manifests as workflows that can't acquire a runner.
  • Verify the underlying cloud connection still grants the trust Orkestia needs to mint registration tokens and scale.

SLA health

SLA Health (Verifiable Health) is a preview surface. It defines policies for uptime, latency, and error-rate over evaluation windows (5m, 10m, 1h, monthly) and turns observed health into shareable, tamper-evident reports. The backing metrics-ingestion and proof path are on the roadmap, so treat it as a preview of that experience rather than a production health source today.

For live health signals, rely on Lumen (latency/error pulse, slow-span p95/p99, error-group counts), which is the production observability surface.

Quick reference

GoalToolNotes
Confirm identity & orgwhoamiMandatory first step
Follow a live runwatch_workflowRead-only, re-attachable
One-line verdictget_workflow_statusCurrent state + error
Full transition logget_workflow_historyRead bottom-up to the failing step
Find stalled runslist_stuck_workflowsTriage by type + stuck_since
Resume a failed runretry_workflowReplays from last committed state
Resume a gated runresolve_workflowAfter fixing the named precondition
Kill an unrecoverable runforce_terminate_workflowDestructive, no compensation
Root-cause a failureLumen logs / error groups / tracesFilter by workflow_id / trace_id

Where to go next

Lumen concepts

How workflow-native observability is structured — logs, error groups, traces, triage memory.

Query API

Filters, mutations, rule JSON.

Use Lumen

App routes mapped to the same endpoints.

Drift detection & self-healing

How Orkestia detects and repairs resources that drift out of declared state.

MCP integration

The run-state tool surface (watch_workflow, retry_workflow, …) and how agents call it.

For per-workflow detail — exact inputs, states, and recovery semantics of any individual workflow type — consult the live catalog at reference.orkestia.dev.