# Workflows

> The execution model in depth. Workflow types versus runs, the event-sourced state machine, DAGs, how an assistant drives it over MCP, and compositions that compile into deterministic engine config

## TL;DR

- **A type is a capability; a run is one execution.** Types have dotted names and typed schemas. Runs have a `workflow_id`, a state, and a history.
- **Three kinds of type:** `state_machine` (atomic, `PENDING → COMPLETED | FAILED`), `data` (read-only lookups, safe to start), and `dag` (multi-step, layers of steps).
- **The engine is event-sourced.** A run is an append-only log of transitions, so history is complete and recovery is deterministic.
- **Compositions are your logic.** Chain existing types with explicit input mapping, validate against the live catalog, compile to a normal DAG. No code.
- **The same loop everywhere.** Discover, schema, prerequisites, start, watch, recover, whether from the console, an SDK, or an assistant over MCP.

<note>

Orkestia never holds your source code or your data. Workflows execute against your cloud accounts through [runners](/concepts/runners), using connection credentials you granted and can revoke. The engine stores workflow **state** and, if you enable it, [Lumen](/concepts/lumen) telemetry.

</note>

## Types vs runs

<table>
<thead>
  <tr>
    <th>
      
    </th>
    
    <th>
      Workflow <strong>
        type
      </strong>
    </th>
    
    <th>
      Workflow <strong>
        run
      </strong>
    </th>
  </tr>
</thead>

<tbody>
  <tr>
    <td>
      What it is
    </td>
    
    <td>
      A registered <em>
        capability
      </em>
    </td>
    
    <td>
      A single <em>
        execution
      </em>
      
       of a type
    </td>
  </tr>
  
  <tr>
    <td>
      Identity
    </td>
    
    <td>
      A dotted name, such as <code>
        aws.s3.create_bucket
      </code>
    </td>
    
    <td>
      A unique <code>
        workflow_id
      </code>
    </td>
  </tr>
  
  <tr>
    <td>
      Lifetime
    </td>
    
    <td>
      Installed in the catalog
    </td>
    
    <td>
      From start to a terminal state
    </td>
  </tr>
  
  <tr>
    <td>
      Carries
    </td>
    
    <td>
      Typed input and output schema, prerequisites, <code>
        workflow_kind
      </code>
      
      , <code>
        featured
      </code>
      
      , <code>
        scope
      </code>
    </td>
    
    <td>
      Inputs, outputs, full transition history
    </td>
  </tr>
  
  <tr>
    <td>
      Discover with
    </td>
    
    <td>
      <code>
        list_workflow_types
      </code>
      
      , <code>
        get_workflow_schema
      </code>
      
      , <code>
        get_workflow_definition
      </code>
      
      , <code>
        get_workflow_dag
      </code>
    </td>
    
    <td>
      <code>
        get_workflow_status
      </code>
      
      , <code>
        watch_workflow
      </code>
      
      , <code>
        get_workflow_history
      </code>
      
      , <code>
        list_workflows
      </code>
    </td>
  </tr>
</tbody>
</table>

### Types: registered capabilities

A type is a named, typed capability the platform knows how to run. Atomic types are named `{provider}.{service}.{operation}` and grouped into namespaces you can browse. Each type declares:

- **A typed input and output schema**, validated before a run starts and before a result is accepted.
- **Prerequisites**: what must exist first, most often a cloud [connection](/aws-connections). When a schema reports `has_prerequisites: true`, fetch the setup guide before starting.
- **Metadata** an assistant uses to choose well: `workflow_kind` (`state_machine`, `data`, or `dag`), `featured` (a human-runnable entry point; internal `*.prepare` and `*.finalize` sub-steps are not), and `scope` (`organization`, `library`, or `none`).

The catalog is the source of truth. Never invent a type name. Quote the exact name the engine returns.

```text [discover a capability, as an assistant does it]
list_workflow_namespaces()
list_workflow_types(prefix="aws.s3.")
get_workflow_schema("aws.s3.create_bucket")          → fields, read_only, has_prerequisites
get_workflow_prerequisites("aws.s3.create_bucket", variant="aws")
```

<tip>

Every per-type detail (inputs, outputs, flags, prerequisites) lives in the live catalog at [reference.orkestia.dev](https://reference.orkestia.dev) and through `get_workflow_schema`. This page describes capabilities at the model level.

</tip>

### Runs: executions with a `workflow_id`

A run is one execution of a type. You start it with inputs, get back a `workflow_id`, then **watch** it to a terminal state, inspect its **history**, or **retry** it if it failed.

```text
start_workflow("aws.s3.create_bucket", { "bucket": "my-bucket", "connection_uuid": "…" })
  → { "workflow_id": "wf_8f3c…", "status": "PENDING" }

watch_workflow("wf_8f3c…")          → blocks until COMPLETED or FAILED
get_workflow_history("wf_8f3c…")    → full append-only transition log
```

<callout icon="i-lucide-info">

Your organization scopes every run automatically. It is resolved server-side from your token and never passed by hand. End-user runs through [App Enablement](/app-enablement) are scoped the same way, to the signed-in user.

</callout>

## The event-sourced state machine

Every run, atomic or composed, executes on one **event-sourced state machine**. A run is not a row mutated in place. It is the **replay of an append-only sequence of transitions**. Current state is derived from history. Two properties follow:

<card-group>
<card icon="i-lucide-history" title="Full history">

Every state change, input, and output is preserved. You can reconstruct exactly how a run reached its outcome.

</card>

<card icon="i-lucide-rotate-ccw" title="Deterministic recovery">

State is rebuilt from events, so a run resumes or retries from a known point instead of restarting blindly.

</card>
</card-group>

### The 3-state pattern

Atomic workflows follow one predictable shape. A run enters `PENDING`, auto-advances to execute its single operation, and lands in exactly one terminal state.

<table>
<thead>
  <tr>
    <th>
      State
    </th>
    
    <th>
      Role
    </th>
    
    <th>
      Notes
    </th>
  </tr>
</thead>

<tbody>
  <tr>
    <td>
      <code>
        PENDING
      </code>
    </td>
    
    <td>
      Initial
    </td>
    
    <td>
      Runs the work through <code>
        auto_advance
      </code>
      
       to <code>
        execute
      </code>
    </td>
  </tr>
  
  <tr>
    <td>
      <code>
        COMPLETED
      </code>
    </td>
    
    <td>
      Terminal
    </td>
    
    <td>
      Success, carries the output
    </td>
  </tr>
  
  <tr>
    <td>
      <code>
        FAILED
      </code>
    </td>
    
    <td>
      Terminal
    </td>
    
    <td>
      Failure, carries a structured reason
    </td>
  </tr>
</tbody>
</table>

<dag-diagram :edges="[{"from":"init","to":"PENDING"},{"from":"PENDING","to":"COMPLETED","label":"execute succeeds"},{"from":"PENDING","to":"FAILED","label":"execute raises"},{"from":"COMPLETED","to":"done"},{"from":"FAILED","to":"done"}]" :nodes="[{"id":"init","label":"[*]","kind":"start"},{"id":"PENDING","label":"PENDING"},{"id":"COMPLETED","label":"COMPLETED","kind":"terminal"},{"id":"FAILED","label":"FAILED","kind":"failed"},{"id":"done","label":"[*]","kind":"terminal"}]">



</dag-diagram>

This uniformity is the point. Because every primitive obeys the same contract, the console, SDKs, and AI assistants treat every capability identically: discover it, start it, watch it to one of two terminal states.

### DAGs: multi-step workflows

Real processes are many steps with data threaded between them. A **DAG workflow** arranges atomic steps in **layers**: layers run in order; steps within a layer can run together. Each step still obeys the 3-state contract, so a DAG is a composition of primitives, not a different runtime. Ask `get_workflow_dag(type)` for the layer and step structure.

<dag-diagram :edges="[{"from":"A","to":"B"},{"from":"A","to":"C"}]" :nodes="[{"id":"A","label":"aws.s3.create_bucket","kind":"cloud"},{"id":"B","label":"aws.s3.put_bucket_versioning","kind":"cloud"},{"id":"C","label":"aws.s3.put_bucket_policy","kind":"cloud"}]" direction="LR">



</dag-diagram>

A DAG step that fails on a *fixable* precondition can park in `remediation_pending` instead of compensating. The run's `state_data.remediation` envelope names the fix. Apply it, then call `resolve_workflow(workflow_id, "remediated")` and the engine re-runs only the failed step. Resolve `"denied"` to compensate and fail.

### Middleware, plugins, and RBAC

Every transition runs through a **middleware pipeline** with a **plugin system**, so validation, authorization, and observability live in one place. **RBAC is enforced in the engine** through declared capability metadata. Permissions are checked the same way for every run, whether it started from the console, an SDK, the API, or an assistant over MCP. There is no back door.

### Concurrency

To keep a single run's transitions serialized, the engine takes a **per-run lock**. No two workers can advance the same run into conflicting states. Concurrency safety is by construction.

### Async

Work that takes time never blocks your call. The engine publishes a `workflow.transition` message; a dedicated consumer picks it up and advances the run out of band. You start a run and watch it. There is no queue to provision.

### Retries, recovery, and stuck runs

<table>
<thead>
  <tr>
    <th>
      Operation
    </th>
    
    <th>
      Tool
    </th>
    
    <th>
      Purpose
    </th>
  </tr>
</thead>

<tbody>
  <tr>
    <td>
      Inspect
    </td>
    
    <td>
      <code>
        get_workflow_status
      </code>
    </td>
    
    <td>
      Current state of a run
    </td>
  </tr>
  
  <tr>
    <td>
      Replay
    </td>
    
    <td>
      <code>
        get_workflow_history
      </code>
    </td>
    
    <td>
      Full append-only transition log (<code>
        include_state_data=True
      </code>
      
       for payloads)
    </td>
  </tr>
  
  <tr>
    <td>
      Recover
    </td>
    
    <td>
      <code>
        retry_workflow
      </code>
    </td>
    
    <td>
      Re-advance a <code>
        FAILED
      </code>
      
       run from its last good point
    </td>
  </tr>
  
  <tr>
    <td>
      Unpark
    </td>
    
    <td>
      <code>
        resolve_workflow
      </code>
    </td>
    
    <td>
      Answer a remediation gate with <code>
        "remediated"
      </code>
      
       or <code>
        "denied"
      </code>
    </td>
  </tr>
  
  <tr>
    <td>
      Triage
    </td>
    
    <td>
      <code>
        list_stuck_workflows
      </code>
    </td>
    
    <td>
      Find runs that stalled before reaching a terminal state
    </td>
  </tr>
  
  <tr>
    <td>
      Abandon
    </td>
    
    <td>
      <code>
        force_terminate_workflow
      </code>
    </td>
    
    <td>
      Append a failed terminal state to a confirmed-stale run, with a reason and guards
    </td>
  </tr>
</tbody>
</table>

The MCP prompt template `diagnose_workflow_run` walks exactly this sequence for a given `workflow_id`.

## Compositions: virtual workflows

The platform ships thousands of capabilities, but *your* business logic is the **order** you run them in and the **data** you thread between them. **Compositions** (virtual workflows) express that by chaining existing workflows, **with no code**.

You declare *which* workflows run, *in what order*, *with which inputs*. The platform validates that against the live catalog and **compiles it into a regular engine DAG**. From then on a composition runs, is watched, and is retried like any other run.

### The three building blocks

<table>
<thead>
  <tr>
    <th>
      Piece
    </th>
    
    <th>
      Role
    </th>
  </tr>
</thead>

<tbody>
  <tr>
    <td>
      <strong>
        Layer
      </strong>
    </td>
    
    <td>
      A stage. Layers run in order; steps within a layer can run together.
    </td>
  </tr>
  
  <tr>
    <td>
      <strong>
        Step
      </strong>
    </td>
    
    <td>
      One existing workflow type to invoke, placed in a layer.
    </td>
  </tr>
  
  <tr>
    <td>
      <strong>
        Input mapping
      </strong>
    </td>
    
    <td>
      Where each step argument's value comes from.
    </td>
  </tr>
</tbody>
</table>

### Input mapping

Every step input is explicit about its source. There is no implicit global scope and no runtime string templating. Three sources:

<card-group>
<card icon="i-lucide-log-in" title="input">

The composition's own input.

</card>

<card icon="i-lucide-arrow-right" title="step">

A prior step's output field.

</card>

<card icon="i-lucide-lock" title="static">

A literal fixed at authoring time.

</card>
</card-group>

Because mappings are explicit, the whole composition is **type-checked before it runs**. Errors are precise (which step, which reason) rather than runtime surprises.

### Validate, compile, stamp, run

Authoring is a deterministic pipeline driven by the `composition.save` workflow: validate the definition, compile it to a DAG, stamp the compiled cache, and the type becomes runnable as `virtual.<uuid>@<version>`.

<dag-diagram :edges="[{"from":"I","to":"V"},{"from":"V","to":"V2"},{"from":"V2","to":"V3"},{"from":"V3","to":"C"},{"from":"C","to":"S"},{"from":"S","to":"R"},{"from":"V","to":"I","label":"errors[]","dashed":true},{"from":"V2","to":"I","label":"errors[]","dashed":true}]" :nodes="[{"id":"I","label":"Definition: layers + steps + mappings","kind":"start"},{"id":"V","label":"Validate: structure","kind":"engine"},{"id":"V2","label":"Validate: references","kind":"engine"},{"id":"V3","label":"Validate: compatibility","kind":"engine"},{"id":"C","label":"Compile to DAG","kind":"engine"},{"id":"S","label":"Stamp: content_hash + format version","kind":"data"},{"id":"R","label":"Run on the engine","kind":"engine"}]" direction="LR">



</dag-diagram>

1. **Structure**: layers are acyclic, step names are unique.
2. **References**: every referenced type exists in the installed catalog.
3. **Compatibility**: every input mapping resolves to a compatible source field and type.

#### Authoring shape

You author the **definition**. Each step points at a real workflow type and declares an `input_mapping` from each parameter to its source (`input`, `step`, or `static`):

```json
{
  "name": "vw_demo",
  "version": "1.0",
  "layers": [
    { "name": "layer_1", "steps": [
        { "name": "s1", "workflow_type": "aws.s3.create_bucket",
          "input_mapping": { "bucket": { "source": "input", "field": "bucket_name" } } }
    ] },
    { "name": "layer_2", "steps": [
        { "name": "s2", "workflow_type": "aws.s3.put_bucket_tagging",
          "input_mapping": {
            "bucket":  { "source": "step",   "step": "s1", "field": "bucket_name" },
            "tagging": { "source": "static", "value": { "TagSet": [{ "Key": "env", "Value": "demo" }] } }
          } }
    ] }
  ]
}
```

The compiled form is a server-owned cache: the per-step `input_mapping` becomes `input`, a `step` source becomes `state_data`, the workflow is marked `type: "dag"`, and a `metadata` block with `content_hash` and `compiled_format_version` is stamped on. The definition stays the durable source of truth; the engine can always recompile from it.

Validation returns two lists of plain strings, precise enough to feed straight back into an authoring loop:

```json
{
  "errors": [
    "step 's2': unknown workflow type 'aws.s3.put_bucket_taging'; did you mean 'aws.s3.put_bucket_tagging'?"
  ],
  "warnings": []
}
```

<warning>

The compilation substrate is **alpha**. The stable contract is the **wire JSON**, not the internal builder API. The catalog is read at compile time: if author and runtime environments differ, a composition can compile cleanly yet fail at dispatch with an "unknown workflow type". Pin library versions across both.

</warning>

### Why this is how apps express logic

A composition is just another workflow type, which means it can be **exposed to your app's end-users** through [App Enablement](/app-enablement). A signed-in user invokes it; Orkestia runs it scoped to that user. Your app's business logic lives on the platform with no backend of your own and no database credentials in your code.

## Ask your AI assistant

```text [prompts]
List the workflow types under "aws.s3." with their workflow_kind, and tell me which ones are featured entry points.

Show me the DAG structure of <dag_workflow_type> as layers and steps.

Read concept://dag, then design a composition that creates a bucket, enables versioning, and applies a tag. Validate it but do not save it.

Workflow <workflow_id> failed. Read its history with state data and tell me the failing step and the reason.
```

## For AI agents

<table>
<thead>
  <tr>
    <th>
      Rule
    </th>
    
    <th>
      Detail
    </th>
  </tr>
</thead>

<tbody>
  <tr>
    <td>
      Confirm the type
    </td>
    
    <td>
      Never assume a type exists. Use <code>
        list_workflow_types
      </code>
      
       (with <code>
        prefix
      </code>
      
       or <code>
        q
      </code>
      
      ) or <code>
        get_workflow_schema
      </code>
      
      .
    </td>
  </tr>
  
  <tr>
    <td>
      Prefer featured
    </td>
    
    <td>
      <code>
        featured
      </code>
      
       rows are human-runnable entry points. Ignore <code>
        *.prepare
      </code>
      
       and <code>
        *.finalize
      </code>
      
       sub-steps.
    </td>
  </tr>
  
  <tr>
    <td>
      Reads are safe
    </td>
    
    <td>
      <code>
        data
      </code>
      
       kinds and names containing <code>
        list
      </code>
      
      , <code>
        get
      </code>
      
      , <code>
        query
      </code>
      
      , <code>
        describe
      </code>
      
      , <code>
        status
      </code>
      
       can be started directly. Confirm creates and mutations.
    </td>
  </tr>
  
  <tr>
    <td>
      Prerequisites first
    </td>
    
    <td>
      If <code>
        has_prerequisites
      </code>
      
       is true, call <code>
        get_workflow_prerequisites(type, variant)
      </code>
      
       before <code>
        start_workflow
      </code>
      
      .
    </td>
  </tr>
  
  <tr>
    <td>
      Virtual types are per org
    </td>
    
    <td>
      Saved compositions run as <code>
        virtual.<uuid>@<version>
      </code>
      
       and do <strong>
        not
      </strong>
      
       appear in <code>
        list_workflow_types
      </code>
      
      . An empty <code>
        prefix="virtual."
      </code>
      
       browse does not mean the user has none. Use <code>
        audit.workflow-run.query
      </code>
      
       to see what ran.
    </td>
  </tr>
  
  <tr>
    <td>
      Report properly
    </td>
    
    <td>
      Workflow type, workflow ID, final state, and the output fields that matter. See <code>
        knowledge://mcp/execution
      </code>
      
      .
    </td>
  </tr>
</tbody>
</table>

## Where this fits

<card-group>
<card icon="i-lucide-sparkles" title="DGI" to="/concepts/dgi">

AI reasoning that designs workflows and compiles them into compositions.

</card>

<card icon="i-lucide-server" title="Runners" to="/concepts/runners">

Where steps actually execute, in your own cloud accounts.

</card>

<card icon="i-lucide-activity" title="Lumen" to="/concepts/lumen">

Observability over every run's history.

</card>

<card icon="i-lucide-users" title="Identity & multi-tenancy" to="/concepts/identity-multi-tenancy">

How runs are scoped per organization and per end-user.

</card>
</card-group>

## Next

- Build a composition step by step: [Virtual workflows guide](/guides/virtual-workflows).
- Let AI design one: [Building with DGI](/guides/building-with-dgi).
- Drive capabilities from an assistant: [Connect an AI assistant](/getting-started/connect-an-ai-assistant) and [MCP integration](/reference/mcp-integration).
- The complete catalog: [reference.orkestia.dev](https://reference.orkestia.dev) and the [workflow types registry](/reference/workflow-types-registry).
