Orkestia
Blog
Advanced Topics

Hybrid AI + Deterministic Execution

How Orkestia lets AI design workflows at runtime, then compiles those plans into deterministic virtual workflows that run repeatably, cheaply, and auditably

TL;DR

Orkestia's flagship execution model separates two jobs that most "AI workflow" tools conflate:

  1. Design — figuring out what steps to run, in what order, with what inputs. This is open-ended, benefits from reasoning, and is where AI (DGI and agents) earns its keep.
  2. Execution — actually running those steps reliably, every time, with locking, retries, observability, and an audit trail. This wants determinism, not creativity.

The hybrid model uses AI for design, then compiles the resulting plan into a deterministic virtual workflow (a composition) that the engine runs with no LLM in the hot path. You explore once with AI, capture the result as a reusable composition, and from then on it runs like any other DAG: byte-stable, cheap, and fully auditable.

"Virtual workflow" and "composition" are the same thing in this document: an AI-authored (or tool-authored) DAG that has been validated and compiled into engine-ready config. See Concepts → Workflows and the Virtual Workflows guide.

Why split design from execution

An LLM in the execution path is a liability: it is non-deterministic, slow, expensive per invocation, and hard to audit ("why did it do that this time?"). But an LLM is also the best tool we have for turning a vague human goal into a concrete plan against a large, changing catalog of capabilities.

Orkestia resolves the tension by letting the model operate only where its strengths apply:

ConcernAI design phaseDeterministic execution phase
Decides step order / inputsYes (reasoning)No — fixed by the compiled DAG
Cost per runHigh (token cost)Near-zero (no model call)
ReproducibilityLow (re-plans each time)Byte-stable given same inputs
LatencyReasoning + tool roundtripsEngine dispatch only
AuditabilityPlan is the artifactFull event-sourced history
Best forNovel / one-off / ambiguous goalsRepeated, known, high-volume flows

The key insight from the platform's internal substrate: the compiler emits a dict that matches the real engine's DAGWorkflow format exactly — there is no runtime marker. The engine does not know or care that a config was authored by an AI. Once compiled, a virtual workflow is indistinguishable from a hand-written one, which is precisely what makes it trustworthy.

The lifecycle: explore → capture → run

The three phases:

1. Explore with AI

A human expresses a goal. DGI (the Dialog Generative Interface) or an autonomous agent reasons about it, discovering available capabilities over MCP rather than against a static, hand-maintained list. The catalog is whatever is installed and importable — so the model plans against ground truth, not a stale snapshot.

The model emits a plan in the VirtualWorkflow wire format: layers of steps, where every step references a {provider}.{service}.{operation} workflow type and every input is an explicit mapping with a known source.

{
  "name": "provision_and_tag_bucket",
  "layers": [
    {
      "name": "create",
      "steps": [
        {
          "name": "s1",
          "workflow_type": "aws.s3.create_bucket",
          "input_mapping": {
            "name": { "source": "input", "field_name": "bucket_name" }
          }
        }
      ]
    },
    {
      "name": "tag",
      "steps": [
        {
          "name": "s2",
          "workflow_type": "aws.s3.put_bucket_tagging",
          "input_mapping": {
            "bucket":  { "source": "step", "step": "s1", "field_name": "bucket_name" },
            "tagging": { "source": "static", "value": { "TagSet": [{ "Key": "env", "Value": "demo" }] } }
          }
        }
      ]
    }
  ]
}
Workflow type names like aws.s3.create_bucket are illustrative. The authoritative, per-workflow catalog lives at reference.orkestia.dev and the Workflow Types Registry.

Input mappings come from exactly three sources, with no implicit global scope and no runtime string templating:

SourceMeaning
inputA field from the workflow's own input
stepAn output field of a prior step
staticA hardcoded literal

This explicitness is what makes the plan compilable — and later, auditable.

2. Capture as a composition

Before anything runs, the plan goes through three-phase validation, each phase short-circuiting with a typed error so the AI can self-correct in a loop:

PhaseChecksCatches
StructureDAG acyclicity, unique step IDsmalformed plans
Referencesevery {provider}.{service}.{operation} exists in the live cataloghallucinated / typo'd workflow types
Compatibilityevery InputMapping resolves to a compatible source field & typeshape drift, wrong wiring

A failure is structured data, not a stack trace in a string — designed to be fed straight back into the model:

{
  "phase": "references",
  "step_id": "s2",
  "workflow_type": "aws.s3.put_bucket_taging",
  "reason": "unknown workflow type; did you mean 'aws.s3.put_bucket_tagging'?"
}

Once validated, compile_workflow emits engine-ready JSON:

{
  "id": "provision_and_tag_bucket",
  "layers": [
    { "steps": [ { "id": "s1", "workflow": "aws.s3.create_bucket",
                   "inputs": { "name": {"from": "input.bucket_name"} } } ] },
    { "steps": [ { "id": "s2", "workflow": "aws.s3.put_bucket_tagging",
                   "inputs": {
                     "bucket":  {"from": "step.s1.bucket_name"},
                     "tagging": {"value": { "TagSet": [{ "Key": "env", "Value": "demo" }] }}
                   } } ] }
  ]
}

This JSON is the composition. It's the durable artifact you version, diff, review, and reuse.

Because compilation is deterministic — same plan + same installed catalog → byte-stable JSON — you can diff two AI-authored plans the way you'd diff source code. That makes AI output reviewable in a pull request rather than trusted on faith.

3. Run deterministically

From here the AI is gone. The compiled DAG is dispatched to the workflow engine, which runs it like any other workflow: PG advisory locks for concurrency, the middleware pipeline, retries, and an event-sourced history that records every transition. The workflow.transition events flow through Kafka to the consumer (the canonical async path), and Lumen captures observability and the audit trail.

No tokens are spent. No model is queried. The run at 3am behaves exactly like the run at noon.

This is also where Orkestia's Zero Code Custody posture holds: the deterministic run executes against the customer's own cloud accounts via runners. Orkestia orchestrates state and observability; it never takes custody of customer code or data. See Security & Compliance and Deployment Models.

When to stay AI-driven vs. when to compile

Compiling is not always the right move. The decision turns on how stable and repeated the task is.

Stay AI-driven

The goal is novel, ambiguous, or one-off. Inputs vary structurally each time. A human is in the loop and wants conversational iteration. The catalog or environment is still changing under you. Cost-per-run is acceptable because volume is low.

Compile to a composition

The same shape of task recurs. You need reproducibility and a stable audit trail. Volume makes per-run LLM cost or latency unacceptable. You want the flow reviewable in version control and runnable by non-AI callers (schedules, webhooks, end-users).

A useful mental model: AI authors the first run; the composition owns every run after. You let the model explore until a plan is good, then "freeze" it into a deterministic workflow. If requirements change later, you go back to the AI, regenerate, re-validate, and ship a new compiled version — the same way you'd edit and re-merge code.

Catalog drift is the sharp edge. The catalog is a snapshot at compile time. If a workflow type is added after you compile, that composition won't see it; if a type is removed after compile, the JSON still references it and the run fails at engine dispatch — a runtime "unknown workflow" error, not a validation error. Mitigation: compile in (or against) the same environment that will dispatch, and pin library versions across authoring and runtime. See the Virtual Workflows guide for the recommended setup.

How AI reaches the catalog

The design phase doesn't guess at capabilities — it discovers them. Agents and DGI talk to the engine over MCP (MCP integration): listing workflow namespaces and types, fetching schemas, checking prerequisites, then starting and watching runs. The same surface that an AI uses to discover capabilities is the one it uses to run a compiled composition, so there's a single, consistent contract from exploration through execution.

The dgi.* workflow family makes the design phase itself a first-class, discoverable surface — construction, conversion, template resolution, and execution are workflow types the engine can run, not hidden internal calls. See Concepts → DGI.

Governance fits naturally

Because compilation produces a reviewable artifact and execution produces an event-sourced history, the hybrid model slots cleanly into Staff governance: a fleet of AI agents can propose compositions, while approval gates decide which ones are allowed to compile and run. The deterministic phase is where oversight has teeth — you approve a frozen plan, not a probability distribution over future model behavior.

Notes & limitations

A few things to keep in mind when building on the hybrid model:
  • The VirtualWorkflow input format may still evolve; the wire JSON is the more stable contract.
  • There is no cross-version catalog arbitration today — if two library versions coexist, import resolution order wins.
  • The compiler is not a runtime: it cannot step, pause, resume, or observe a run. Use the engine for that.
  • Per-workflow specifics (exact names, inputs, limits) come from reference.orkestia.dev, not inferred from examples here.

Where to go next

Virtual Workflows guide

Hands-on: author, validate, and compile a composition end to end.

Building with DGI

Drive the design phase conversationally with the Dialog Generative Interface.

Workflows concept

The engine, the 3-state pattern, and DAG execution that compositions run on.

MCP integration

How agents discover and run capabilities over MCP.

Drift detection & self-healing

The deterministic counterpart for infrastructure state.

Lumen observability

The audit trail and observability behind every deterministic run.