Best Practices
TL;DR
This page distills the operational habits that keep an Orkestia deployment reliable, governed, and cheap to run. None of it is novel — it is the disciplined application of the mechanisms documented elsewhere in this track. Treat it as a pre-flight checklist: each item links to the deep page that explains why, so you can audit your own setup against it.
The themes, in priority order:
- Design for repeatability — idempotent workflows, compositions over bespoke code.
- Govern the AI — let DGI design, but compile flows once stable and gate high-impact actors with approvals.
- Lock down the blast radius — least-privilege cloud roles, proper tenant isolation.
- See everything — Lumen alerts on the signals that matter, right-sized runners with drift detection on.
(beta). For per-workflow specifics — exact input schemas,
flags, and limits — always defer to the workflow catalog.The checklist at a glance
Design idempotent workflows
Every step should be safe to run twice. Retries and self-healing depend on it.
Compile AI flows once stable
Explore with DGI, then capture the result as a deterministic composition — no LLM in the hot path.
Scope least-privilege roles
Grant runners and connections only the permissions a workflow actually exercises.
Govern AI actors
Put approval gates in front of high-impact agent actions; bind roles, audit everything.
1. Design idempotent workflows
The single most leveraged habit. Orkestia's reliability machinery — retries, drift detection and self-healing, and the event-sourced engine that may replay a transition after a pod restart — all assume that re-running a step produces the same end state, not a duplicate side effect.
A step that creates a resource should first check whether it already exists. A step that
charges, emits, or provisions should carry an idempotency key so the provider deduplicates.
The atomic workflows in the base library follow a three-state pattern
(PENDING → COMPLETED / FAILED); design your compositions so any state can be safely
re-entered from PENDING.
| Pattern | Do | Avoid |
|---|---|---|
| Resource creation | Check-then-create, or create-if-not-exists | Blind create that errors on the second run |
| External mutations | Pass an idempotency key the provider honors | Relying on "it only runs once" |
| Counters / appends | Make the operation set-absolute, not increment | Increment without a guard |
| Cleanup | Tolerate "already gone" | Fail when the target is missing |
retry_workflow,
self-healing) become safe to use only once every step is replay-safe.2. Compile AI flows once they are stable
DGI's strength is design — figuring out what steps to run for a novel intent. That is open-ended work where AI reasoning earns its keep. But you do not want an LLM in the hot path of a flow you run a thousand times a day: it is slower, costlier, and non-deterministic.
The platform's answer is the hybrid execution model: explore once with DGI, then capture the result as a deterministic virtual workflow (composition). The virtual engine validates the AI-authored plan against the live catalog (does every referenced workflow type exist? do the input mappings type-check?) and compiles it to a plain DAG config the engine runs with no LLM involved.
- Novel intent
- DGI — AI design
- Workflow plan (DAG)
- Virtual enginevalidate + compile
- Compositiondeterministic DAG
- Workflow engine
- Novel intent→DGI — AI design
- DGI — AI design→Workflow plan (DAG)
- Workflow plan (DAG)→Virtual engine
- Virtual engine→Composition
- Composition→ run N times, no LLM →Workflow engine
| When | Use | Why |
|---|---|---|
| One-off / exploratory | DGI live design | Flexibility beats repeatability |
| Recurring / production | Compiled composition | Byte-stable, cheap, auditable, no LLM latency |
| Plan changed | Re-design with DGI, re-compile | Keep the composition as the source of truth |
3. Scope least-privilege cloud roles
Orkestia's Zero Trust / Zero Code Custody posture means execution happens in your cloud accounts — Orkestia orchestrates and stores only workflow state and observability data. The corollary: the connections and roles you grant define the entire blast radius. Scope them tightly.
- Grant only what a workflow exercises. Start from the workflow's declared
prerequisites (
get_workflow_prerequisitesreturns the exact platform principal and scope), not a broad managed policy. - One role per purpose, per tenant. Avoid a single god-role shared across workflows and organizations. Separate read-only collection from mutating actions.
- Prefer connection prerequisites over ad-hoc credentials. Connections are the canonical, auditable way to grant access; see AWS Connections and DNS Providers.
- Rotate and review. Treat every granted role as standing risk; review them on the same cadence as Lumen alerts.
4. Govern your AI actors with approvals
Staff is the governance layer for fleets of AI agents — organizational structure, RBAC, sessions, and oversight. Agents are powerful precisely because they can discover and run capabilities over MCP; that same power is why high-impact actions need a human gate.
- Bind roles, don't grant ambient power. RBAC is enforced in the workflow engine via
declared workflow
Capabilitymetadata (RbacGuard). An actor can only run what its role binding allows; denied attempts surface in audit. - Put approval gates in front of high-impact actions. Provisioning, spend, destructive operations, and anything customer-facing should require human-in-the-loop sign-off.
- Scope actors into units. Model staff as an organization (org → unit → actor) so oversight and cost attribution have structure, not a flat pile of agents.
- Audit everything, and read the audit.
staff.list-audit-eventsand workflow-run history are only useful if someone (or a Lumen alert) actually watches them. - Right-size the model and budget per actor. Cost controls and model pricing live in the Staff console; a cheap model on a low-stakes task is the default, not the exception.
5. Set Lumen alerts on the signals that matter
Lumen is the observability plane — it stores workflow state transitions and observability data so you can see what ran, what failed, and where time and money went. Logs alone are noise; alert on signals.
| Alert on | Why | Where |
|---|---|---|
| Workflow failure rate (per type) | Catches a broken integration or a bad deploy early | Observability with Lumen |
| Stuck / stalled workflows | A run that never reaches a terminal state ties up capacity | list_stuck_workflows |
| Drift events | Desired vs. observed runner state diverged | Drift Detection |
| Denied RBAC attempts | An actor reaching beyond its role binding | Staff audit |
| Cost per useful output | Catches runaway agent spend before the invoice does | Staff console |
6. Right-size your runners
Runners are the execution capacity in your cloud — Kubernetes and multi-cloud — that the engine dispatches work onto. Both under- and over-provisioning cost you: too little capacity stalls workflows; too much burns money idle.
- Match capacity to observed load, not to peak fear. Use Lumen to see real concurrency and queue depth before scaling.
- Keep drift detection on. The reconciliation loop continuously compares desired vs. observed runner state and surfaces — or, where you allow it, corrects — divergence. See Drift Detection & Self-Healing.
- Let self-healing reconcile, but only over idempotent work (see section 1). Enable automatic correction once you trust replay-safety; start with propose-only.
- Separate runner groups by purpose. Keep agent-session runner groups distinct from general workflow capacity so a noisy agent fleet can't starve production flows. See the runner management guide.
7. Isolate tenants properly
Orkestia is multi-tenant by design, down to "Sign in with Orkestia" for exposing workflows to your own end-users. Isolation is not a single switch — it is consistent scoping at every layer.
- Scope state per organization. Your
organization_uuidis resolved server-side from your token and scopes every run automatically — never pass another org's ID, never share a token across tenants. - Never share cloud credentials across tenants. Each organization's connections and roles are its own; a shared credential collapses the isolation boundary.
- Scope identity for end-users. When exposing workflows via Sign in with Orkestia and App Enablement, keep end-user identity scoped to the tenant that owns the app. (beta)
- Attribute cost and audit per tenant so spend and access reviews map cleanly to an organization.
8. Prefer compositions over bespoke code
The platform ships a validated catalog of atomic workflows (one operation each) and business workflows (composed multi-step flows). Before writing custom logic, ask whether a composition of existing catalog entries does the job.
- Reuse the catalog. Browse the workflow types registry and the full catalog before building anything new.
- Compose, don't fork. The virtual engine lets you assemble validated steps into a DAG with type-checked input mappings — you get catalog knowledge and structural validation for free.
- Let the engine own execution semantics. Locking, retries, event sourcing, and audit come from the engine; bespoke code outside it has to reimplement all of that (usually worse).
- If you must author new workflows, keep them atomic and idempotent, follow the
{provider}.{service}.{operation}naming and three-state pattern, and register them so they're discoverable over MCP.
Pre-production checklist
Run through this before promoting a workflow or runner group to production:
- Every step in the flow is idempotent / replay-safe (§1)
- Exploratory DGI design has been compiled to a stable composition (§2)
- Cloud roles grant only what the workflow exercises (§3)
- High-impact AI actions are gated by approvals; roles are bound, not ambient (§4)
- Lumen alerts exist for failure rate, stuck runs, drift, and cost — each with an owner (§5)
- Runner capacity matches observed load; drift detection is on (§6)
- State, credentials, and identity are scoped per tenant; nothing shared across orgs (§7)
- Catalog/composition reuse was preferred over bespoke code (§8)
