Key Benefits
TL;DR
Orkestia's benefits are consequences of architecture, not adjectives. Each one below is tied to a mechanism you can inspect and a page that explains it.
| Benefit | Mechanism |
|---|---|
| Your code and data never leave your cloud | Zero Code Custody: scoped, revocable credentials; Orkestia stores state, telemetry, and encrypted connection credentials, never code or data |
| Reliability in the core | Event-sourced engine, per-run locks, async transitions, explicit retries |
| AI leverage without losing control | AI designs, the engine validates and executes deterministically |
| Governance for agent fleets | Staff: org structure, engine-enforced RBAC, approvals, audit |
| Multi-cloud runners that heal themselves | Reconcile-loop scaling, health checks, drift repair |
| Apps on top, fast | Sign in with Orkestia, App Data, exposed compositions |
| Any AI assistant becomes an operator | One MCP server: discover, run, watch, recover, with guardrails as resources |
1. Privacy and compliance by architecture
Most platforms ask you to trust a promise. Orkestia removes the need for one: the architecture makes custody impossible.
- Execution runs in your accounts. Workflows that touch infrastructure execute against your cloud connections. Compute and data stay where they already live.
- Orkestia is a control plane, not a hosting plane. For runners, the runner binary talks to GitHub, not to Orkestia. Orkestia mints a short-lived registration token and then only observes.
- What Orkestia stores is workflow state, observability data, and the connection credentials you grant. The event-sourced history of a run, the Lumen telemetry you choose to send, and, encrypted at rest, the scoped role or token each connection needs. Not your source code, not your databases.
2. Reliability from an event-sourced core
Every capability on the platform is a workflow, and every run inherits the same guarantees:
- Event sourcing. A run is a replayable sequence of transitions.
get_workflow_historyreturns the full record. - Concurrency safety. Operations on the same resource are serialized on a per-resource lock. Independent resources run in parallel.
- Async is first class. Long-running steps advance through the engine's transition bus. Nothing is fire-and-forget.
- Retries and recovery. Failures are explicit terminal states you can
retry_workflowfrom. Stalled runs surface throughlist_stuck_workflows.
Compositions sit on top of that runtime. AI may design a flow, but what executes is a validated DAG that is byte-stable for the same inputs.
whoami()
list_workflow_types(prefix="connection.")
get_workflow_schema("connection.setup") → has_prerequisites: true
get_workflow_prerequisites("connection.setup", variant="aws")
start_workflow("connection.setup", {...}) → workflow_id
watch_workflow(workflow_id) → COMPLETED | FAILED
get_workflow_history(workflow_id) → full transition log
retry_workflow(workflow_id) → explicit, replayable recovery
3. AI leverage without losing control
The hard part of "AI-native" is not getting a model to suggest an action. It is keeping that suggestion safe, typed, and repeatable. Orkestia's hybrid execution model separates the jobs:
| Phase | Who does it | Property |
|---|---|---|
| Design | DGI, or an assistant over MCP, authors a DAG of existing capabilities | Conversational, exploratory |
| Validate | The virtual engine type-checks structure, references, and input mappings against the live catalog | Deterministic gate, typed errors |
| Execute | The workflow engine runs the compiled DAG | Deterministic, replayable, observable |
The validator refuses a plan that references a capability that does not exist or whose inputs do not type-check. A hallucinated step name comes back as a structured error the model can correct:
{
"errors": [
"step 's2': unknown workflow type 'aws.s3.put_bucket_taging'; did you mean 'aws.s3.put_bucket_tagging'?"
],
"warnings": []
}
The MCP server reinforces this at the assistant level with rule://grounding (say only what tool results show) and rule://prerequisites-first (never start a workflow whose setup is missing).
dgi.* workflows are evolving. The capability is real and discoverable over MCP; the API shape may still change. See Hybrid execution model and Building with DGI.4. Governance for fleets of AI agents
Running one agent is a demo. Running many agents against real infrastructure is an organizational problem. Staff turns a collection of agents into an accountable organization:
- Org, unit, actor structure. Agents are actors inside units, with a full lifecycle (hire, update, pause, resume, archive, invoke).
- RBAC enforced in the engine. Permissions are
Capabilitymetadata on workflows, checked by the engine before the first step runs. A denied action never executes and is recorded. - Approvals and oversight. An operator inbox shows what needs a human decision.
- Cost controls. Per-actor budgets and spend analytics.
5. Multi-cloud runners with drift detection and self-healing
- Multi-cloud from one plane. Runner groups provision self-hosted GitHub Actions runners across AWS (ECS, EC2), Azure (Container Apps, VMs), Kubernetes, GCP, DigitalOcean, and Magalu Cloud from the same group abstraction. Compute stays in your account.
- Reconcile-loop scaling. A per-group loop converges the pool to your
minandmax. Webhooks nudge it; they are never the source of truth. - Self-healing. Liveness probes deregister unhealthy runners and reap their compute. Half-failed teardowns are swept on the next health check.
| Failure | What happens | Recovery |
|---|---|---|
| Provider quota hit at scale-up | Scale-up fails with a quota error; existing runners unaffected | Retryable once quota is raised |
| Webhook lost | Slightly slower reaction | Reconcile loop converges on its next pass |
| Orphaned runner after partial teardown | A zombie offline runner appears | Reconciler sweeps on the next health check |
6. Fast app delivery on top
The same engine that runs your infrastructure can power features for your users:
- "Sign in with Orkestia". Hosted OIDC and PKCE login for your app's end-users, with RS256 JWTs and a published JWKS.
- App Data. Declared tables with injected principals. No DSN in your app. Operators use Query for admitted SQL. Dedicated instances via
appdata.instance.*. - App Host. Claim a live site, attach Postgres, optionally apply Nostr Buzz. Signing keys are nsec, not org members.
- Compositions become features. Chain workflows into your logic and expose it to end-users, scoped to each user.
Provisioning an identity app is one workflow call (identity.app.provision). An assistant connected over MCP can do it unattended by following rule://orkestia-auth-setup. See App Enablement.
7. Any AI assistant becomes an operator
This is the benefit that ties the others together. The Orkestia MCP server at https://mcp.orkestia.dev/mcp gives any MCP client the same window onto the engine that the console has:
- Tools for the whole lifecycle:
whoami, discovery, schema, prerequisites, start, watch, history, stuck runs, retry, resolve, andopen_appto render the Orkestia console inside the chat. - Resources that carry the guardrails as first-class context:
rule://getting-started,rule://authenticated-context,rule://prerequisites-first,rule://grounding,rule://orkestia-auth-setup, plusconcept://andknowledge://orientation documents. - Prompts you can invoke from the client:
how_to_use_workflow_mcpanddiagnose_workflow_run.
Because org scoping, RBAC, approvals, and audit are enforced in the engine, an assistant gets exactly the authority its token carries and nothing more. See Connect an AI assistant and MCP integration.
Ask your AI assistant
Read concept://product and knowledge://orkestia/capabilities, then tell me which domains my org can operate today.
List the stuck workflows in my organization and explain why each one is stalled.
Use the diagnose_workflow_run prompt on workflow <workflow_id> and propose a recovery.
Show me the runner groups I have and their current min/max policy.
Benefit, mechanism, where to read
| Benefit | Mechanism | Where to read |
|---|---|---|
| Code and data stay in your cloud | Zero Code Custody for cloud work; App Data / App Host are opt-in platform stores | Security & compliance, App Data, App Host |
| Reliable execution | Event-sourced engine, per-run locks, async transitions, explicit retries | Workflows |
| Deterministic AI-designed flows | Virtual engine compiles validated DAGs | Hybrid execution model, Virtual workflows |
| AI leverage, human control | Design, validate, execute; grounding and prerequisites rules | DGI, MCP integration |
| Govern agent fleets | Staff structure, engine-enforced RBAC, approvals, audit | Staff governance |
| Self-healing runners | Reconcile loop, health checks, drift repair | Runners, Drift & self-healing |
| Observability | Lumen ingest, error groups, traces, metrics, Lumen MCP | Lumen, Lumen MCP |
| Apps on top | @orkestia/auth, App Data, App Host, exposed compositions | App Enablement, App Data, App Host |
| Assistant as operator | One MCP server with tools, rule resources, and prompts | Connect an AI assistant |
Where to go next
Core Philosophy
The three pillars behind Orkestia, Zero Code Custody, AI-designed deterministic execution, and customer-owned compute, and what they mean for humans and for AI assistants
Quick Start
From zero to a first workflow run in your own cloud. Create an org, connect a cloud, then run a workflow from an AI assistant over MCP, the console, or the SDKs
