Cost & Performance Optimization
Orkestia gives you two execution styles that sit on the same engine: AI reasoning (DGI designs and adapts flows on the fly) and deterministic compositions (a compiled DAG runs the same way every time). Cost and latency live almost entirely in the choice between them, in how much runner capacity you keep warm, and in how you fan work out asynchronously. This page is the practical playbook: where the money and the milliseconds go, the levers you have, and the tradeoffs each lever carries.
The three cost centers
Most Orkestia spend traces back to three places. Optimize them in roughly this order — the first usually has the biggest leverage.
LLM reasoning
Every time DGI reasons over a request — interpreting intent, choosing capabilities, sequencing steps — you pay for model tokens and wall-clock latency. Re-reasoning the same flow on every run is the most common avoidable cost.
Runner / compute capacity
Self-hosted runners and execution environments cost money whenever they are warm, whether or not a job is in flight. Idle warm pools trade cost for pickup latency.
Step execution + waiting
Individual workflow steps call provider APIs, wait on long-running cloud operations, and pass data between layers. Slow or redundant steps inflate both latency and the compute time you pay for.
Lever 1 — Compile AI-designed flows to deterministic compositions
This is the highest-leverage move in the platform. DGI is excellent at designing a workflow from intent, but you should not pay an LLM to re-derive that design on every execution. The Virtual Workflow Engine exists precisely to capture a design once and turn it into config the deterministic engine can run forever — with no model in the hot path.
How the compile step removes LLM cost
When DGI (or any producer — a UI, an API client, an agent) settles on a plan, it expresses it as a VirtualWorkflow: typed layers of steps with explicit input mappings. The virtual engine then:
- Discovers the live catalog of installed workflow types (no static list — it reads what the installed libraries actually register).
- Validates structure (acyclic DAG, unique step IDs), references (every
{provider}.{service}.{operation}exists), and compatibility (every input mapping type-checks). - Compiles to a plain JSON DAG matching the engine's
DAGWorkflowshape.
The real workflow engine consumes that JSON exactly as it would any hand-authored DAG. There is no runtime marker and no model call — at execution time it is pure deterministic orchestration.
- User intent
- DGI reasoning
- VirtualWorkflow plan
- Validated DAG JSON
- Workflow engine
- run 1
- run 2
- run ...
- User intent→ tokens $$ →DGI reasoning
- DGI reasoning→VirtualWorkflow plan
- VirtualWorkflow plan→ compile once →Validated DAG JSON
- Validated DAG JSON→ run N times, $0 LLM →Workflow engine
- Workflow engine→run 1
- Workflow engine→run 2
- Workflow engine→run ...
The economic shape is simple: reasoning cost is amortized across every subsequent run. A flow designed once and run ten thousand times pays the LLM bill once.
When to keep DGI in the loop anyway
Compilation is the right default for anything repeated, but not everything should be frozen. Keep reasoning live when the shape of the work genuinely changes per request.
| Pattern | Recommended style | Why |
|---|---|---|
| Repeated, well-understood operation (nightly provision, standard deploy) | Compiled composition | Stable shape; pay LLM once |
| One-off or exploratory request | DGI live | Not worth compiling something you run once |
| Flow whose branching depends on unpredictable runtime data | DGI live, or compiled DAG with conditional steps | Decide whether the variation is structural or just data |
| High-volume, latency-sensitive automation | Compiled composition | Removes model latency from the hot path entirely |
The deeper treatment of when each style fires — and how they interoperate — is in Hybrid Execution Model.
Lever 2 — Right-size runner warm pools and capacity
Runners (and the broader execution environments they model) are where idle cost accumulates. A runner that is warm but unused costs you provider minutes for zero work; a cold start costs you pickup latency. Tuning the warm pool is the act of choosing where on that curve you want to sit.
The core tradeoff
cost ◄────────────────────────────────► latency
high warm pool balanced scale-to-zero
(instant pickup, (some warm, (cheapest,
pay for idle) spill cold) cold-start delay)
Orkestia scales runner groups with a reconciliation loop that continuously converges the pool toward the group's configured min/max — your primary cost dial. GitHub workflow_job events are a best-effort nudge that triggers an immediate reconcile (queued may scale up, completed may scale down); because the loop, not the webhook, is authoritative, scaling stays correct and bounded even if an event is missed.
{
"scale": { "min": 0, "max": 10, "idle_timeout_s": 300 }
}
| Setting | Effect on cost | Effect on latency |
|---|---|---|
min: 0 (scale to zero) | Lowest — no idle runners | Cold start on first queued job |
min: N (warm floor) | Pay for N idle runners always | First N jobs picked up instantly |
max (ceiling) | Hard cap on spend — Orkestia never exceeds it even under queue pressure | Jobs queue once the cap is hit |
idle_timeout_s | Higher = fewer scale cycles, more idle time | Higher = warmer runners between bursts |
min to your steady-state concurrency floor (the number of jobs you almost always have in flight) and let the reconcile loop scale up to max for bursts. Reserve a non-zero min only for the groups whose pickup latency users actually feel — batch and nightly groups are usually fine at min: 0.max. If you would rather burst than wait, raise max — but that is an explicit decision to spend more. Self-hosted runners also scale in seconds-to-minutes, not milliseconds, so they are not the right tool when you need extreme instantaneous burst.Reap dead capacity automatically
Idle isn't the only waste — orphaned capacity is worse, because it costs money and serves nothing. Runner groups run periodic health checks: unhealthy runners are deregistered from GitHub and the compute reaped so dead runners don't sit on the bill or starve the queue. The reconciler also sweeps zombie registrations left behind by a failed teardown. You get this for free, but it's worth watching the orphaned-runner rate (it should trend to zero) and reacting if it doesn't. See Runner Management and drift detection & self-healing.
Lever 3 — Use Kafka async for throughput, not just durability
Long-running steps — cloud provisioning, drains, AWS waiters, anything that blocks on an external system — must not occupy a synchronous request. Orkestia's canonical async path puts these on the Kafka workflow.transition bus, consumed by the workflow Kafka consumer. This is both a reliability property (work survives pod restarts) and a performance property: it decouples submission rate from execution rate, so a burst of work fans out across consumers instead of serializing behind a request thread.
- API / MCP submit
- Engine: enqueue transition
- Kafka
- consumer
- consumer
- consumer
- step exec
- step exec
- step exec
- API / MCP submit→Engine: enqueue transition
- Engine: enqueue transition→ workflow.transition →Kafka
- Kafka→consumer
- Kafka→consumer
- Kafka→consumer
- consumer→step exec
- consumer→step exec
- consumer→step exec
Practical consequences for cost and speed:
- Parallelism is the default. Independent workflow groups execute fully in parallel; only operations on the same serialized entity (e.g. the same runner group, guarded by a PG advisory lock) are ordered. Structure work so unrelated items don't share a lock key.
- No long-poll waiting threads. Waiting on a cloud operation happens via transitions, not a held connection, so you don't pay for idle compute holding a socket.
- Backpressure is visible. Consumer lag is your throughput signal — if it grows, scale consumers rather than runners.
Lever 4 — Use Lumen to find the slow and expensive steps
You can't optimize what you can't see. Lumen is Orkestia's workflow-native observability layer, and it attributes every signal to structural context — workflow, actor, state, trace, span, project, environment, release. That attribution is exactly what lets you answer "which step is costing me" instead of guessing.
The signals that matter for cost and performance work:
| Lumen surface | What it tells you | Optimization it drives |
|---|---|---|
| Traces & spans | Per-step duration, slow spans, latency p95/p99 | Find the step to compile, parallelize, or cache |
| Metrics | Aggregate time-series by metric name | Track cost/latency trends across releases |
| Error groups | Repeated failures grouped by fingerprint | Kill retry storms that burn compute |
| Dashboard | Project summary, latency and cost signals, top error groups | First place to look for a regression |
A typical optimization loop:
- Open the Lumen dashboard, sort by latency or error volume.
- Drill into the slow trace; identify the dominant span.
- Decide the lever — is it an LLM call that should be compiled away, a step that should run in a parallel layer, or a flaky step generating retries?
- Apply the change; watch the same span's p95 across the next releases to confirm.
trace_id / span_id and workflow/actor context attached. Missing trace IDs collapse the trace UI into coarse pseudo-traces grouped by workflow, which makes per-step attribution much harder. Instrument emitting context early — it's the difference between "this workflow is slow" and "step 3 of 7 is slow." See Observability with Lumen.Beware ingest rules that hide the signal
Lumen's ingest rules can drop or sample logs to control observability cost. That's a legitimate lever — but over-aggressive sampling can discard the very spans you need to find a slow step, and bad fingerprint rules can mask a retry storm by mis-grouping it. Tune ingest rules with the same care as the workflows they observe.
Putting it together — an optimization checklist
Compile what repeats
Freeze any stable, repeated DGI-designed flow into a deterministic composition. Amortize the LLM cost to near zero.
Right-size warm pools
Set min to your real concurrency floor, cap max for cost safety, let the reconcile loop handle bursts.
Tradeoffs at a glance
| Lever | You gain | You give up |
|---|---|---|
| Compile to composition | Lower LLM cost, lower latency, reproducibility | Runtime adaptability; must recompile when the design changes |
Non-zero warm pool min | Instant pickup | Pay for idle runners |
Aggressive min: 0 / scale-to-zero | Lowest idle cost | Cold-start latency on first job |
High max ceiling | Burst capacity | Higher peak spend |
| Kafka async fan-out | Throughput, durability | More moving parts to monitor (consumer lag) |
| Lumen ingest sampling | Lower observability cost | Risk of dropping the signal you need |
Related reading
- Hybrid Execution Model — how AI reasoning and deterministic compositions interoperate
- Drift Detection & Self-Healing — keeping runner capacity correct and reaped
- Concepts: Workflows, Runners, Lumen, DGI
- Operations: Deployment Models and Security & Compliance
- Reference: Workflow Types Registry and the external Orkestia capability catalog
Governance & Approvals
How Orkestia enforces role-based authority, human-in-the-loop approval gates, and an immutable evidence trail over fleets of autonomous AI agents
Operations
Run Orkestia in production — deployment topologies, monitoring and debugging, and the operational practices that keep workflows reliable
