Hybrid AI + Deterministic Execution
TL;DR
Orkestia's flagship execution model separates two jobs that most "AI workflow" tools conflate:
- Design — figuring out what steps to run, in what order, with what inputs. This is open-ended, benefits from reasoning, and is where AI (DGI and agents) earns its keep.
- Execution — actually running those steps reliably, every time, with locking, retries, observability, and an audit trail. This wants determinism, not creativity.
The hybrid model uses AI for design, then compiles the resulting plan into a deterministic virtual workflow (a composition) that the engine runs with no LLM in the hot path. You explore once with AI, capture the result as a reusable composition, and from then on it runs like any other DAG: byte-stable, cheap, and fully auditable.
Why split design from execution
An LLM in the execution path is a liability: it is non-deterministic, slow, expensive per invocation, and hard to audit ("why did it do that this time?"). But an LLM is also the best tool we have for turning a vague human goal into a concrete plan against a large, changing catalog of capabilities.
Orkestia resolves the tension by letting the model operate only where its strengths apply:
| Concern | AI design phase | Deterministic execution phase |
|---|---|---|
| Decides step order / inputs | Yes (reasoning) | No — fixed by the compiled DAG |
| Cost per run | High (token cost) | Near-zero (no model call) |
| Reproducibility | Low (re-plans each time) | Byte-stable given same inputs |
| Latency | Reasoning + tool roundtrips | Engine dispatch only |
| Auditability | Plan is the artifact | Full event-sourced history |
| Best for | Novel / one-off / ambiguous goals | Repeated, known, high-volume flows |
The key insight from the platform's internal substrate: the compiler emits a dict that matches the real engine's DAGWorkflow format exactly — there is no runtime marker. The engine does not know or care that a config was authored by an AI. Once compiled, a virtual workflow is indistinguishable from a hand-written one, which is precisely what makes it trustworthy.
The lifecycle: explore → capture → run
- Human goal / intent
- DGI / agent reasoning
- Live workflow catalog
- VirtualWorkflow planlayers - steps - input mappings
- Validate
- compile_workflow
- DAGWorkflow JSONengine-ready, no LLM marker
- Workflow enginedeterministic run
- Event-sourced history + Lumen observability
- Human goal / intent→DGI / agent reasoning
- DGI / agent reasoning→ discovers capabilities over MCP →Live workflow catalog
- DGI / agent reasoning→VirtualWorkflow plan
- VirtualWorkflow plan→Validate
- Validate→ structure / references / compatibility →compile_workflow
- Validate→ typed error →DGI / agent reasoning
- compile_workflow→DAGWorkflow JSON
- DAGWorkflow JSON→Workflow engine
- Workflow engine→Event-sourced history + Lumen observability
- DAGWorkflow JSON→ reusable composition →Workflow engine
The three phases:
1. Explore with AI
A human expresses a goal. DGI (the Dialog Generative Interface) or an autonomous agent reasons about it, discovering available capabilities over MCP rather than against a static, hand-maintained list. The catalog is whatever is installed and importable — so the model plans against ground truth, not a stale snapshot.
The model emits a plan in the VirtualWorkflow wire format: layers of steps, where every step references a {provider}.{service}.{operation} workflow type and every input is an explicit mapping with a known source.
{
"name": "provision_and_tag_bucket",
"layers": [
{
"name": "create",
"steps": [
{
"name": "s1",
"workflow_type": "aws.s3.create_bucket",
"input_mapping": {
"name": { "source": "input", "field_name": "bucket_name" }
}
}
]
},
{
"name": "tag",
"steps": [
{
"name": "s2",
"workflow_type": "aws.s3.put_bucket_tagging",
"input_mapping": {
"bucket": { "source": "step", "step": "s1", "field_name": "bucket_name" },
"tagging": { "source": "static", "value": { "TagSet": [{ "Key": "env", "Value": "demo" }] } }
}
}
]
}
]
}
aws.s3.create_bucket are illustrative. The authoritative, per-workflow catalog lives at reference.orkestia.dev and the Workflow Types Registry.Input mappings come from exactly three sources, with no implicit global scope and no runtime string templating:
| Source | Meaning |
|---|---|
input | A field from the workflow's own input |
step | An output field of a prior step |
static | A hardcoded literal |
This explicitness is what makes the plan compilable — and later, auditable.
2. Capture as a composition
Before anything runs, the plan goes through three-phase validation, each phase short-circuiting with a typed error so the AI can self-correct in a loop:
| Phase | Checks | Catches |
|---|---|---|
| Structure | DAG acyclicity, unique step IDs | malformed plans |
| References | every {provider}.{service}.{operation} exists in the live catalog | hallucinated / typo'd workflow types |
| Compatibility | every InputMapping resolves to a compatible source field & type | shape drift, wrong wiring |
A failure is structured data, not a stack trace in a string — designed to be fed straight back into the model:
{
"phase": "references",
"step_id": "s2",
"workflow_type": "aws.s3.put_bucket_taging",
"reason": "unknown workflow type; did you mean 'aws.s3.put_bucket_tagging'?"
}
Once validated, compile_workflow emits engine-ready JSON:
{
"id": "provision_and_tag_bucket",
"layers": [
{ "steps": [ { "id": "s1", "workflow": "aws.s3.create_bucket",
"inputs": { "name": {"from": "input.bucket_name"} } } ] },
{ "steps": [ { "id": "s2", "workflow": "aws.s3.put_bucket_tagging",
"inputs": {
"bucket": {"from": "step.s1.bucket_name"},
"tagging": {"value": { "TagSet": [{ "Key": "env", "Value": "demo" }] }}
} } ] }
]
}
This JSON is the composition. It's the durable artifact you version, diff, review, and reuse.
3. Run deterministically
From here the AI is gone. The compiled DAG is dispatched to the workflow engine, which runs it like any other workflow: PG advisory locks for concurrency, the middleware pipeline, retries, and an event-sourced history that records every transition. The workflow.transition events flow through Kafka to the consumer (the canonical async path), and Lumen captures observability and the audit trail.
No tokens are spent. No model is queried. The run at 3am behaves exactly like the run at noon.
When to stay AI-driven vs. when to compile
Compiling is not always the right move. The decision turns on how stable and repeated the task is.
Stay AI-driven
The goal is novel, ambiguous, or one-off. Inputs vary structurally each time. A human is in the loop and wants conversational iteration. The catalog or environment is still changing under you. Cost-per-run is acceptable because volume is low.
Compile to a composition
The same shape of task recurs. You need reproducibility and a stable audit trail. Volume makes per-run LLM cost or latency unacceptable. You want the flow reviewable in version control and runnable by non-AI callers (schedules, webhooks, end-users).
A useful mental model: AI authors the first run; the composition owns every run after. You let the model explore until a plan is good, then "freeze" it into a deterministic workflow. If requirements change later, you go back to the AI, regenerate, re-validate, and ship a new compiled version — the same way you'd edit and re-merge code.
How AI reaches the catalog
The design phase doesn't guess at capabilities — it discovers them. Agents and DGI talk to the engine over MCP (MCP integration): listing workflow namespaces and types, fetching schemas, checking prerequisites, then starting and watching runs. The same surface that an AI uses to discover capabilities is the one it uses to run a compiled composition, so there's a single, consistent contract from exploration through execution.
The dgi.* workflow family makes the design phase itself a first-class, discoverable surface — construction, conversion, template resolution, and execution are workflow types the engine can run, not hidden internal calls. See Concepts → DGI.
Governance fits naturally
Because compilation produces a reviewable artifact and execution produces an event-sourced history, the hybrid model slots cleanly into Staff governance: a fleet of AI agents can propose compositions, while approval gates decide which ones are allowed to compile and run. The deterministic phase is where oversight has teeth — you approve a frozen plan, not a probability distribution over future model behavior.
Notes & limitations
- The VirtualWorkflow input format may still evolve; the wire JSON is the more stable contract.
- There is no cross-version catalog arbitration today — if two library versions coexist, import resolution order wins.
- The compiler is not a runtime: it cannot step, pause, resume, or observe a run. Use the engine for that.
- Per-workflow specifics (exact names, inputs, limits) come from reference.orkestia.dev, not inferred from examples here.
Where to go next
Advanced Topics
Deep dives into the mechanisms that make Orkestia reliable, governed, and efficient at scale
Drift Detection & Self-Healing
How Orkestia reconciles desired and actual state for runners and network infrastructure in your cloud, computes readiness verdicts, and decides what to heal automatically versus surface for approval
