Orkestia
Blog
Introduction

Key Benefits

What teams get from Orkestia, each benefit mapped to the mechanism that produces it, privacy by architecture, an event-sourced engine, AI leverage with control, governed agent fleets, self-healing runners, and apps on top

TL;DR

Orkestia's benefits are consequences of architecture, not adjectives. Each one below is tied to a mechanism you can inspect and a page that explains it.

BenefitMechanism
Your code and data never leave your cloudZero Code Custody: scoped, revocable credentials; Orkestia stores state, telemetry, and encrypted connection credentials, never code or data
Reliability in the coreEvent-sourced engine, per-run locks, async transitions, explicit retries
AI leverage without losing controlAI designs, the engine validates and executes deterministically
Governance for agent fleetsStaff: org structure, engine-enforced RBAC, approvals, audit
Multi-cloud runners that heal themselvesReconcile-loop scaling, health checks, drift repair
Apps on top, fastSign in with Orkestia, App Data, exposed compositions
Any AI assistant becomes an operatorOne MCP server: discover, run, watch, recover, with guardrails as resources
Orkestia is generally available. The workflow engine and the AWS, Azure, and Kubernetes runner paths run in production. Other providers are expanding. Roadmap items are flagged in place. Per-workflow detail lives in the workflow reference.

1. Privacy and compliance by architecture

Most platforms ask you to trust a promise. Orkestia removes the need for one: the architecture makes custody impossible.

  • Execution runs in your accounts. Workflows that touch infrastructure execute against your cloud connections. Compute and data stay where they already live.
  • Orkestia is a control plane, not a hosting plane. For runners, the runner binary talks to GitHub, not to Orkestia. Orkestia mints a short-lived registration token and then only observes.
  • What Orkestia stores is workflow state, observability data, and the connection credentials you grant. The event-sourced history of a run, the Lumen telemetry you choose to send, and, encrypted at rest, the scoped role or token each connection needs. Not your source code, not your databases.
For a compliance reviewer, the practical consequence is a much smaller blast radius: a hypothetical Orkestia breach exposes orchestration metadata, not your data plane. See Security & compliance and Deployment models.

2. Reliability from an event-sourced core

Every capability on the platform is a workflow, and every run inherits the same guarantees:

  • Event sourcing. A run is a replayable sequence of transitions. get_workflow_history returns the full record.
  • Concurrency safety. Operations on the same resource are serialized on a per-resource lock. Independent resources run in parallel.
  • Async is first class. Long-running steps advance through the engine's transition bus. Nothing is fire-and-forget.
  • Retries and recovery. Failures are explicit terminal states you can retry_workflow from. Stalled runs surface through list_stuck_workflows.

Compositions sit on top of that runtime. AI may design a flow, but what executes is a validated DAG that is byte-stable for the same inputs.

the loop every run shares, as an assistant sees it
whoami()
list_workflow_types(prefix="connection.")
get_workflow_schema("connection.setup")          → has_prerequisites: true
get_workflow_prerequisites("connection.setup", variant="aws")
start_workflow("connection.setup", {...})        → workflow_id
watch_workflow(workflow_id)                      → COMPLETED | FAILED
get_workflow_history(workflow_id)                → full transition log
retry_workflow(workflow_id)                      → explicit, replayable recovery
The 3-state pattern, middleware pipeline, locks, and async path are covered in Workflows.

3. AI leverage without losing control

The hard part of "AI-native" is not getting a model to suggest an action. It is keeping that suggestion safe, typed, and repeatable. Orkestia's hybrid execution model separates the jobs:

PhaseWho does itProperty
DesignDGI, or an assistant over MCP, authors a DAG of existing capabilitiesConversational, exploratory
ValidateThe virtual engine type-checks structure, references, and input mappings against the live catalogDeterministic gate, typed errors
ExecuteThe workflow engine runs the compiled DAGDeterministic, replayable, observable

The validator refuses a plan that references a capability that does not exist or whose inputs do not type-check. A hallucinated step name comes back as a structured error the model can correct:

{
  "errors": [
    "step 's2': unknown workflow type 'aws.s3.put_bucket_taging'; did you mean 'aws.s3.put_bucket_tagging'?"
  ],
  "warnings": []
}

The MCP server reinforces this at the assistant level with rule://grounding (say only what tool results show) and rule://prerequisites-first (never start a workflow whose setup is missing).

The virtual engine and the dgi.* workflows are evolving. The capability is real and discoverable over MCP; the API shape may still change. See Hybrid execution model and Building with DGI.

4. Governance for fleets of AI agents

Running one agent is a demo. Running many agents against real infrastructure is an organizational problem. Staff turns a collection of agents into an accountable organization:

  • Org, unit, actor structure. Agents are actors inside units, with a full lifecycle (hire, update, pause, resume, archive, invoke).
  • RBAC enforced in the engine. Permissions are Capability metadata on workflows, checked by the engine before the first step runs. A denied action never executes and is recorded.
  • Approvals and oversight. An operator inbox shows what needs a human decision.
  • Cost controls. Per-actor budgets and spend analytics.
See Staff governance for the model and Staff & Agents for the operator path.

5. Multi-cloud runners with drift detection and self-healing

  • Multi-cloud from one plane. Runner groups provision self-hosted GitHub Actions runners across AWS (ECS, EC2), Azure (Container Apps, VMs), Kubernetes, GCP, DigitalOcean, and Magalu Cloud from the same group abstraction. Compute stays in your account.
  • Reconcile-loop scaling. A per-group loop converges the pool to your min and max. Webhooks nudge it; they are never the source of truth.
  • Self-healing. Liveness probes deregister unhealthy runners and reap their compute. Half-failed teardowns are swept on the next health check.
FailureWhat happensRecovery
Provider quota hit at scale-upScale-up fails with a quota error; existing runners unaffectedRetryable once quota is raised
Webhook lostSlightly slower reactionReconcile loop converges on its next pass
Orphaned runner after partial teardownA zombie offline runner appearsReconciler sweeps on the next health check
AWS, Azure, and Kubernetes runner groups run in production today. GCP, DigitalOcean, and Magalu paths are beta. Check the workflow reference for the current per-provider matrix.

6. Fast app delivery on top

The same engine that runs your infrastructure can power features for your users:

  • "Sign in with Orkestia". Hosted OIDC and PKCE login for your app's end-users, with RS256 JWTs and a published JWKS.
  • App Data. Declared tables with injected principals. No DSN in your app. Operators use Query for admitted SQL. Dedicated instances via appdata.instance.*.
  • App Host. Claim a live site, attach Postgres, optionally apply Nostr Buzz. Signing keys are nsec, not org members.
  • Compositions become features. Chain workflows into your logic and expose it to end-users, scoped to each user.

Provisioning an identity app is one workflow call (identity.app.provision). An assistant connected over MCP can do it unattended by following rule://orkestia-auth-setup. See App Enablement.

7. Any AI assistant becomes an operator

This is the benefit that ties the others together. The Orkestia MCP server at https://mcp.orkestia.dev/mcp gives any MCP client the same window onto the engine that the console has:

  • Tools for the whole lifecycle: whoami, discovery, schema, prerequisites, start, watch, history, stuck runs, retry, resolve, and open_app to render the Orkestia console inside the chat.
  • Resources that carry the guardrails as first-class context: rule://getting-started, rule://authenticated-context, rule://prerequisites-first, rule://grounding, rule://orkestia-auth-setup, plus concept:// and knowledge:// orientation documents.
  • Prompts you can invoke from the client: how_to_use_workflow_mcp and diagnose_workflow_run.

Because org scoping, RBAC, approvals, and audit are enforced in the engine, an assistant gets exactly the authority its token carries and nothing more. See Connect an AI assistant and MCP integration.

Ask your AI assistant

prompts
Read concept://product and knowledge://orkestia/capabilities, then tell me which domains my org can operate today.

List the stuck workflows in my organization and explain why each one is stalled.

Use the diagnose_workflow_run prompt on workflow <workflow_id> and propose a recovery.

Show me the runner groups I have and their current min/max policy.

Benefit, mechanism, where to read

BenefitMechanismWhere to read
Code and data stay in your cloudZero Code Custody for cloud work; App Data / App Host are opt-in platform storesSecurity & compliance, App Data, App Host
Reliable executionEvent-sourced engine, per-run locks, async transitions, explicit retriesWorkflows
Deterministic AI-designed flowsVirtual engine compiles validated DAGsHybrid execution model, Virtual workflows
AI leverage, human controlDesign, validate, execute; grounding and prerequisites rulesDGI, MCP integration
Govern agent fleetsStaff structure, engine-enforced RBAC, approvals, auditStaff governance
Self-healing runnersReconcile loop, health checks, drift repairRunners, Drift & self-healing
ObservabilityLumen ingest, error groups, traces, metrics, Lumen MCPLumen, Lumen MCP
Apps on top@orkestia/auth, App Data, App Host, exposed compositionsApp Enablement, App Data, App Host
Assistant as operatorOne MCP server with tools, rule resources, and promptsConnect an AI assistant

Where to go next

Connect an AI assistant

Plug in Claude, ChatGPT, or Cursor and run your first workflow by talking.

Quick Start

Create an org, connect a cloud, run a workflow.

Core Concepts

The engine, types and runs, compositions, identity, and governance.

Browse the catalog

Exact names, inputs, prerequisites, and limits for every workflow.