Orkestia
Blog
Advanced Topics

Advanced Topics

Deep dives into the mechanisms that make Orkestia reliable, governed, and efficient at scale

The advanced track is for platform engineers and technical leaders who have already absorbed the core concepts and want to understand how Orkestia behaves under load, under failure, and under organizational pressure. These pages trade introductory framing for mechanism: state machines, reconciliation loops, approval gates, and the cost/performance levers you can pull.

If you are new to the platform, start with the Introduction and Getting Started first, then come back here once you are comfortable with workflows, DGI, and runners.

What you'll find here

Hybrid AI + Deterministic Execution

How DGI's AI reasoning designs a workflow once, then compiles it into a deterministic virtual workflow (composition) that runs the same way every time — the boundary between probabilistic planning and reproducible execution.

Drift Detection & Self-Healing

The reconciliation loop that continuously compares desired vs. observed state across your Kubernetes and multi-cloud runners, surfaces drift, and either proposes or applies a correction.

Governance & Approvals

Staff as a control plane for fleets of AI agents — organizational structure, approval gates, and human-in-the-loop oversight for high-impact actions.

Cost & Performance Optimization

Where time and money actually go in an orchestration run, and the levers — compositions, async via the workflow engine, runner sizing — that move them.

The two execution planes

Almost everything in the advanced track comes back to one architectural fact: Orkestia keeps a hard line between orchestration (what Orkestia runs and stores) and execution (what runs in your cloud). This is the Zero Trust / Zero Code Custody posture described in Security & Compliance.

The control plane is event-sourced: the workflow engine records every state transition rather than mutating a single mutable row, which is what makes the engine's behavior auditable, replayable, and safe to retry. Execution happens on runners inside the customer's accounts; Orkestia retains only workflow state plus the observability data that Lumen collects — never customer code or data.

Keep this diagram in mind as you read the advanced pages. "Drift" is about the Customer box diverging from what the Orkestia box recorded as desired; "hybrid execution" is about the path from DGI to ENG; "governance" is the gate AI agents pass through before a transition reaches a runner.

How the pieces connect

TopicThe mechanism underneathDeep-dive
Hybrid executionDGI designs → compiles to a deterministic virtual workflow → engine executesHybrid Execution Model
Drift & self-healingReconciliation loop over runner stateDrift Detection & Self-Healing
GovernanceStaff approval gates on agent-initiated actionsGovernance & Approvals
Cost & performanceCompositions, Kafka-backed async, runner sizingCost & Performance

Hybrid AI + deterministic execution

AI is excellent at deciding what should happen and poor at doing the same thing twice. Orkestia resolves that tension by letting DGI reason over your intent and the available capabilities (discovered over MCP), then compiling that plan into a deterministic composition. Once compiled, the virtual workflow runs through the engine with no model in the hot path — same inputs, same DAG, same outputs. See Building with DGI for the authoring side.

Drift detection & self-healing

A runner that was healthy at deploy time drifts: a node is recycled, a manifest is edited out-of-band, a cloud resource is deleted. Orkestia's reconciliation loop compares the desired state (recorded in the engine) against what it observes on the runner and flags the delta. Depending on policy, it surfaces the drift for review or self-heals by re-applying the desired state. The full loop is covered in Drift Detection & Self-Healing and operationally in Runner Management.

Governance & approvals

When fleets of AI agents can start workflows, "who approved this?" becomes a first-class question. Staff is the governance layer: agents live inside an organizational structure with explicit authority, and high-impact transitions pass through approval gates before they reach a runner. This is how oversight scales with the number of agents rather than collapsing under it.

Approval-policy surfaces and the breadth of self-healing actions are still expanding. Treat the governance and drift features documented here as evolving, and confirm per-action behavior in the reference catalog for production change control.

Cost & performance optimization

The biggest performance wins come from not using AI where you don't need it (compile to a composition) and from moving long-running work off the request path. All long-running work runs through the workflow engine's Kafka backend — a dedicated consumer advances workflow.transition messages, and there is no second async system to reason about. Lumen is where you see where the time and spend actually go, so you can size runners and shape DAGs against real data rather than guesses.

Where to go next

Operations

Deployment models, environments, and the operational runbooks for running Orkestia.

Reference

Workflow type registry, MCP integration, and the per-workflow catalog at reference.orkestia.dev.