Orkestia
Blog
Operations

Operations

Run Orkestia in production — deployment topologies, monitoring and debugging, and the operational practices that keep workflows reliable

The Operations track is for the people who carry the pager. It assumes you've absorbed the core concepts and have at least one workflow running, and it focuses on the day-two questions: where does Orkestia run, how do I see what it's doing, and what do I do when something goes wrong?

Orkestia's operational model is shaped by one architectural commitment — the Zero Trust / Zero Code Custody posture. The control plane orchestrates and stores workflow state plus Lumen observability data; the actual work executes inside your cloud accounts via runners. That split is what makes operations here different from a classic single-tenant app: you operate two planes that talk over a narrow, audited boundary.

New to the platform? Start with the Introduction and Getting Started, then come back here once you have a runner connected and a workflow or two under Lumen.

In this track

Deployment Models

The topologies Orkestia supports — how the control plane and your runner planes fit together, what lives where, and how to choose a model that matches your trust, latency, and compliance constraints.

Monitoring & Debugging

Using Lumen to watch executions in flight, inspect the event-sourced history of any run, find stuck or failed workflows, and trace a failure back to the exact step and runner.

Best Practices

Operational hygiene that keeps a fleet healthy — least-privilege connections, governed approvals via Staff, idempotent retries, and reconciliation over manual intervention.

The operator's mental model

Three nouns cover almost everything you'll touch operationally:

ConceptWhat it isWhere you operate it
Workflow runA single execution with a unique workflow_id and an event-sourced historyLumen, MCP
RunnerAn execution target (a Kubernetes cluster or cloud account) Orkestia dispatches toRunner management
ConnectionThe least-privilege credential/role that lets the control plane reach your cloudAWS connections, Cloud connections

Everything reconciles toward desired state. Rather than mutating resources imperatively, Orkestia continuously compares desired vs. observed state and either surfaces drift or corrects it — see Drift detection & self-healing.

Observe a run end-to-end

The fastest operational loop is over MCP: discover capabilities, watch live runs, and recover failures without leaving your agent or terminal.

# Discover what's registered for your org
list_workflow_namespaces
list_workflow_types --prefix "k8s."

# Watch a live execution and tail its event history
watch_workflow <workflow_id>
get_workflow_history <workflow_id>

# Find and recover stalled or failed runs
list_stuck_workflows
retry_workflow <workflow_id>
Every run is event-sourced, so get_workflow_history gives you a complete, replayable audit trail — each state transition, its inputs, and its timestamp. This is your primary debugging artifact when a run misbehaves. See the workflow concepts page for the state-machine model behind it.

When something breaks

A quick triage order that maps to the tools above:

  1. Is the run stuck or failed? Check status, then get_workflow_history to find the failing transition. → Monitoring & Debugging
  2. Is the runner reachable and healthy? Drift or a degraded runner shows up in reconciliation. → Runner management, Drift detection
  3. Is it a permissions/connection problem? Re-check the connection's least-privilege role. → AWS connections
  4. Was it gated, not broken? A run may be paused awaiting a Staff approval rather than failed.
Orkestia is generally available. Some operational surfaces (deployment topologies, self-healing automation, and Lumen dashboards) are still evolving, and some actions remain manual today. Treat self-healing as propose-then-apply unless you've explicitly enabled automatic correction, and confirm per-workflow specifics in the workflow types registry and the external catalog at reference.orkestia.dev.

Deployment Models

Where the control plane and runner planes live, and how to pick a topology.

Observability with Lumen

The hands-on guide to monitoring, tracing, and debugging executions.

Runner Management

Connect, size, and keep your Kubernetes and multi-cloud runners healthy.

Security & Compliance

The Zero Trust / Zero Code Custody posture that frames every operational decision.