Orkestia
Blog
Operations

Best Practices

A checklist-driven field guide for running Orkestia in production — idempotent workflows, least-privilege roles, governed AI actors, and right-sized runners

TL;DR

This page distills the operational habits that keep an Orkestia deployment reliable, governed, and cheap to run. None of it is novel — it is the disciplined application of the mechanisms documented elsewhere in this track. Treat it as a pre-flight checklist: each item links to the deep page that explains why, so you can audit your own setup against it.

The themes, in priority order:

  1. Design for repeatability — idempotent workflows, compositions over bespoke code.
  2. Govern the AI — let DGI design, but compile flows once stable and gate high-impact actors with approvals.
  3. Lock down the blast radius — least-privilege cloud roles, proper tenant isolation.
  4. See everything — Lumen alerts on the signals that matter, right-sized runners with drift detection on.
Orkestia is generally available. Some controls referenced here (per-workflow rate limits, specific cost dashboards, fine-grained alert rules) are still stabilizing. Where a control is not yet GA it is flagged (beta). For per-workflow specifics — exact input schemas, flags, and limits — always defer to the workflow catalog.

The checklist at a glance

Design idempotent workflows

Every step should be safe to run twice. Retries and self-healing depend on it.

Compile AI flows once stable

Explore with DGI, then capture the result as a deterministic composition — no LLM in the hot path.

Scope least-privilege roles

Grant runners and connections only the permissions a workflow actually exercises.

Govern AI actors

Put approval gates in front of high-impact agent actions; bind roles, audit everything.

Set Lumen alerts

Alert on failure rate, stuck workflows, and drift — not just raw logs.

Right-size runners

Match capacity to load, keep drift detection on, let self-healing reconcile.

Isolate tenants properly

Scope identity, connections, and state per organization; never share credentials across tenants.

Prefer compositions over code

Reuse the validated catalog and the virtual engine instead of hand-rolling bespoke logic.

1. Design idempotent workflows

The single most leveraged habit. Orkestia's reliability machinery — retries, drift detection and self-healing, and the event-sourced engine that may replay a transition after a pod restart — all assume that re-running a step produces the same end state, not a duplicate side effect.

A step that creates a resource should first check whether it already exists. A step that charges, emits, or provisions should carry an idempotency key so the provider deduplicates. The atomic workflows in the base library follow a three-state pattern (PENDING → COMPLETED / FAILED); design your compositions so any state can be safely re-entered from PENDING.

PatternDoAvoid
Resource creationCheck-then-create, or create-if-not-existsBlind create that errors on the second run
External mutationsPass an idempotency key the provider honorsRelying on "it only runs once"
Counters / appendsMake the operation set-absolute, not incrementIncrement without a guard
CleanupTolerate "already gone"Fail when the target is missing
If you cannot make a step naturally idempotent, make it detectable: have it record a marker the next run can read. The engine's retry and recovery paths (retry_workflow, self-healing) become safe to use only once every step is replay-safe.
Workflows that are not idempotent are the most common cause of damage during automatic retry. Before enabling self-healing on a runner group, audit the workflows it runs for replay-safety. See Drift Detection & Self-Healing.

2. Compile AI flows once they are stable

DGI's strength is design — figuring out what steps to run for a novel intent. That is open-ended work where AI reasoning earns its keep. But you do not want an LLM in the hot path of a flow you run a thousand times a day: it is slower, costlier, and non-deterministic.

The platform's answer is the hybrid execution model: explore once with DGI, then capture the result as a deterministic virtual workflow (composition). The virtual engine validates the AI-authored plan against the live catalog (does every referenced workflow type exist? do the input mappings type-check?) and compiles it to a plain DAG config the engine runs with no LLM involved.

WhenUseWhy
One-off / exploratoryDGI live designFlexibility beats repeatability
Recurring / productionCompiled compositionByte-stable, cheap, auditable, no LLM latency
Plan changedRe-design with DGI, re-compileKeep the composition as the source of truth
A compiled composition is deterministic: same input plus same installed catalog produces byte-stable output. That makes plans diffable and reviewable — treat a compiled composition like code in review.
Compile-time catalog is a snapshot. If the author environment runs an older library than the runtime, compilation can succeed but dispatch can fail on a missing workflow type. Pin library versions across authoring and runtime, or compile in the same process that dispatches. (beta)

3. Scope least-privilege cloud roles

Orkestia's Zero Trust / Zero Code Custody posture means execution happens in your cloud accounts — Orkestia orchestrates and stores only workflow state and observability data. The corollary: the connections and roles you grant define the entire blast radius. Scope them tightly.

  • Grant only what a workflow exercises. Start from the workflow's declared prerequisites (get_workflow_prerequisites returns the exact platform principal and scope), not a broad managed policy.
  • One role per purpose, per tenant. Avoid a single god-role shared across workflows and organizations. Separate read-only collection from mutating actions.
  • Prefer connection prerequisites over ad-hoc credentials. Connections are the canonical, auditable way to grant access; see AWS Connections and DNS Providers.
  • Rotate and review. Treat every granted role as standing risk; review them on the same cadence as Lumen alerts.
Because Orkestia never holds custody of your code or data, the cloud role is the trust boundary. A least-privilege role that can only do what the workflow needs is the strongest single control you have. Detail: Security & Compliance.

4. Govern your AI actors with approvals

Staff is the governance layer for fleets of AI agents — organizational structure, RBAC, sessions, and oversight. Agents are powerful precisely because they can discover and run capabilities over MCP; that same power is why high-impact actions need a human gate.

  • Bind roles, don't grant ambient power. RBAC is enforced in the workflow engine via declared workflow Capability metadata (RbacGuard). An actor can only run what its role binding allows; denied attempts surface in audit.
  • Put approval gates in front of high-impact actions. Provisioning, spend, destructive operations, and anything customer-facing should require human-in-the-loop sign-off.
  • Scope actors into units. Model staff as an organization (org → unit → actor) so oversight and cost attribution have structure, not a flat pile of agents.
  • Audit everything, and read the audit. staff.list-audit-events and workflow-run history are only useful if someone (or a Lumen alert) actually watches them.
  • Right-size the model and budget per actor. Cost controls and model pricing live in the Staff console; a cheap model on a low-stakes task is the default, not the exception.
Start every new AI actor in the most restrictive mode that still lets it be useful, then graduate its authority based on observed quality — not optimism. Staff & Agents walks through role bindings and approval setup.

5. Set Lumen alerts on the signals that matter

Lumen is the observability plane — it stores workflow state transitions and observability data so you can see what ran, what failed, and where time and money went. Logs alone are noise; alert on signals.

Alert onWhyWhere
Workflow failure rate (per type)Catches a broken integration or a bad deploy earlyObservability with Lumen
Stuck / stalled workflowsA run that never reaches a terminal state ties up capacitylist_stuck_workflows
Drift eventsDesired vs. observed runner state divergedDrift Detection
Denied RBAC attemptsAn actor reaching beyond its role bindingStaff audit
Cost per useful outputCatches runaway agent spend before the invoice doesStaff console
Wire the alerts that have a clear owner and a clear response. An alert nobody acts on is worse than none — it trains people to ignore the dashboard.

6. Right-size your runners

Runners are the execution capacity in your cloud — Kubernetes and multi-cloud — that the engine dispatches work onto. Both under- and over-provisioning cost you: too little capacity stalls workflows; too much burns money idle.

  • Match capacity to observed load, not to peak fear. Use Lumen to see real concurrency and queue depth before scaling.
  • Keep drift detection on. The reconciliation loop continuously compares desired vs. observed runner state and surfaces — or, where you allow it, corrects — divergence. See Drift Detection & Self-Healing.
  • Let self-healing reconcile, but only over idempotent work (see section 1). Enable automatic correction once you trust replay-safety; start with propose-only.
  • Separate runner groups by purpose. Keep agent-session runner groups distinct from general workflow capacity so a noisy agent fleet can't starve production flows. See the runner management guide.
Enabling automatic self-healing over non-idempotent workflows can amplify a problem instead of fixing it — a reconcile loop that re-runs a non-replay-safe step repeats its side effects. Audit first, automate second.

7. Isolate tenants properly

Orkestia is multi-tenant by design, down to "Sign in with Orkestia" for exposing workflows to your own end-users. Isolation is not a single switch — it is consistent scoping at every layer.

  • Scope state per organization. Your organization_uuid is resolved server-side from your token and scopes every run automatically — never pass another org's ID, never share a token across tenants.
  • Never share cloud credentials across tenants. Each organization's connections and roles are its own; a shared credential collapses the isolation boundary.
  • Scope identity for end-users. When exposing workflows via Sign in with Orkestia and App Enablement, keep end-user identity scoped to the tenant that owns the app. (beta)
  • Attribute cost and audit per tenant so spend and access reviews map cleanly to an organization.
Multi-tenancy is only as strong as its weakest shared resource. The most common leak is a credential or connection reused across organizations — keep them strictly per-tenant.

8. Prefer compositions over bespoke code

The platform ships a validated catalog of atomic workflows (one operation each) and business workflows (composed multi-step flows). Before writing custom logic, ask whether a composition of existing catalog entries does the job.

  • Reuse the catalog. Browse the workflow types registry and the full catalog before building anything new.
  • Compose, don't fork. The virtual engine lets you assemble validated steps into a DAG with type-checked input mappings — you get catalog knowledge and structural validation for free.
  • Let the engine own execution semantics. Locking, retries, event sourcing, and audit come from the engine; bespoke code outside it has to reimplement all of that (usually worse).
  • If you must author new workflows, keep them atomic and idempotent, follow the {provider}.{service}.{operation} naming and three-state pattern, and register them so they're discoverable over MCP.
A composition is reviewable, diffable, and runs on the same reliable substrate as every other workflow. Bespoke code outside the engine is a maintenance liability that forfeits the platform's reliability guarantees. When in doubt, compose.

Pre-production checklist

Run through this before promoting a workflow or runner group to production:

  • Every step in the flow is idempotent / replay-safe (§1)
  • Exploratory DGI design has been compiled to a stable composition (§2)
  • Cloud roles grant only what the workflow exercises (§3)
  • High-impact AI actions are gated by approvals; roles are bound, not ambient (§4)
  • Lumen alerts exist for failure rate, stuck runs, drift, and cost — each with an owner (§5)
  • Runner capacity matches observed load; drift detection is on (§6)
  • State, credentials, and identity are scoped per tenant; nothing shared across orgs (§7)
  • Catalog/composition reuse was preferred over bespoke code (§8)

Where to go next

Security & Compliance

The Zero Trust / Zero Code Custody model that underpins least-privilege roles and tenant isolation.

Hybrid Execution Model

Why you design with AI and run with deterministic compositions.

Drift Detection & Self-Healing

The reconciliation loop behind right-sized, self-correcting runners.

Deployment Models

How the control plane and your execution plane are deployed and operated.