Operations
The Operations track is for the people who carry the pager. It assumes you've absorbed the core concepts and have at least one workflow running, and it focuses on the day-two questions: where does Orkestia run, how do I see what it's doing, and what do I do when something goes wrong?
Orkestia's operational model is shaped by one architectural commitment — the Zero Trust / Zero Code Custody posture. The control plane orchestrates and stores workflow state plus Lumen observability data; the actual work executes inside your cloud accounts via runners. That split is what makes operations here different from a classic single-tenant app: you operate two planes that talk over a narrow, audited boundary.
- Workflow engineevent-sourced state
- Lumenobservability
- DGIAI design
- RunnersK8s + multi-cloud
- Your resourcescompute, data, APIs
- Workflow engine→ dispatch + state →Runners
- Runners→ telemetry only →Lumen
- Runners→Your resources
In this track
Deployment Models
The topologies Orkestia supports — how the control plane and your runner planes fit together, what lives where, and how to choose a model that matches your trust, latency, and compliance constraints.
Monitoring & Debugging
Using Lumen to watch executions in flight, inspect the event-sourced history of any run, find stuck or failed workflows, and trace a failure back to the exact step and runner.
Best Practices
Operational hygiene that keeps a fleet healthy — least-privilege connections, governed approvals via Staff, idempotent retries, and reconciliation over manual intervention.
The operator's mental model
Three nouns cover almost everything you'll touch operationally:
| Concept | What it is | Where you operate it |
|---|---|---|
| Workflow run | A single execution with a unique workflow_id and an event-sourced history | Lumen, MCP |
| Runner | An execution target (a Kubernetes cluster or cloud account) Orkestia dispatches to | Runner management |
| Connection | The least-privilege credential/role that lets the control plane reach your cloud | AWS connections, Cloud connections |
Everything reconciles toward desired state. Rather than mutating resources imperatively, Orkestia continuously compares desired vs. observed state and either surfaces drift or corrects it — see Drift detection & self-healing.
Observe a run end-to-end
The fastest operational loop is over MCP: discover capabilities, watch live runs, and recover failures without leaving your agent or terminal.
# Discover what's registered for your org
list_workflow_namespaces
list_workflow_types --prefix "k8s."
# Watch a live execution and tail its event history
watch_workflow <workflow_id>
get_workflow_history <workflow_id>
# Find and recover stalled or failed runs
list_stuck_workflows
retry_workflow <workflow_id>
get_workflow_history gives you a complete,
replayable audit trail — each state transition, its inputs, and its timestamp. This
is your primary debugging artifact when a run misbehaves. See the
workflow concepts page for the state-machine model behind it.When something breaks
A quick triage order that maps to the tools above:
- Is the run stuck or failed? Check status, then
get_workflow_historyto find the failing transition. → Monitoring & Debugging - Is the runner reachable and healthy? Drift or a degraded runner shows up in reconciliation. → Runner management, Drift detection
- Is it a permissions/connection problem? Re-check the connection's least-privilege role. → AWS connections
- Was it gated, not broken? A run may be paused awaiting a Staff approval rather than failed.
Related reading
Cost & Performance Optimization
How to keep Orkestia workflows fast and cheap by compiling AI-designed flows to deterministic compositions, right-sizing runners, scaling async over Kafka, and using Lumen to find the expensive steps
Deployment Models
How Orkestia splits a managed control plane from execution that runs entirely inside your own cloud accounts
