Orkestia
Blog
Advanced Topics

Cost & Performance Optimization

How to keep Orkestia workflows fast and cheap by compiling AI-designed flows to deterministic compositions, right-sizing runners, scaling async over Kafka, and using Lumen to find the expensive steps

Orkestia gives you two execution styles that sit on the same engine: AI reasoning (DGI designs and adapts flows on the fly) and deterministic compositions (a compiled DAG runs the same way every time). Cost and latency live almost entirely in the choice between them, in how much runner capacity you keep warm, and in how you fan work out asynchronously. This page is the practical playbook: where the money and the milliseconds go, the levers you have, and the tradeoffs each lever carries.

Because Orkestia runs under Zero Code Custody, the compute cost of your jobs (runner minutes, container time, egress) is billed to you by your own cloud provider — Orkestia never holds it. That makes cost optimization here largely about reducing work and right-sizing your own infrastructure, not about shaving an Orkestia hosting bill. Orkestia's own footprint is workflow STATE plus Lumen observability data. For how the Orkestia side itself is billed (a prepaid subscription with seats and add-ons, plus execution and request meters above the included volume), see Billing, Pricing & Seats.

The three cost centers

Most Orkestia spend traces back to three places. Optimize them in roughly this order — the first usually has the biggest leverage.

LLM reasoning

Every time DGI reasons over a request — interpreting intent, choosing capabilities, sequencing steps — you pay for model tokens and wall-clock latency. Re-reasoning the same flow on every run is the most common avoidable cost.

Runner / compute capacity

Self-hosted runners and execution environments cost money whenever they are warm, whether or not a job is in flight. Idle warm pools trade cost for pickup latency.

Step execution + waiting

Individual workflow steps call provider APIs, wait on long-running cloud operations, and pass data between layers. Slow or redundant steps inflate both latency and the compute time you pay for.

Lever 1 — Compile AI-designed flows to deterministic compositions

This is the highest-leverage move in the platform. DGI is excellent at designing a workflow from intent, but you should not pay an LLM to re-derive that design on every execution. The Virtual Workflow Engine exists precisely to capture a design once and turn it into config the deterministic engine can run forever — with no model in the hot path.

How the compile step removes LLM cost

When DGI (or any producer — a UI, an API client, an agent) settles on a plan, it expresses it as a VirtualWorkflow: typed layers of steps with explicit input mappings. The virtual engine then:

  1. Discovers the live catalog of installed workflow types (no static list — it reads what the installed libraries actually register).
  2. Validates structure (acyclic DAG, unique step IDs), references (every {provider}.{service}.{operation} exists), and compatibility (every input mapping type-checks).
  3. Compiles to a plain JSON DAG matching the engine's DAGWorkflow shape.

The real workflow engine consumes that JSON exactly as it would any hand-authored DAG. There is no runtime marker and no model call — at execution time it is pure deterministic orchestration.

The economic shape is simple: reasoning cost is amortized across every subsequent run. A flow designed once and run ten thousand times pays the LLM bill once.

Compiled output is byte-stable for the same input plan and the same installed catalog. That means you can diff an LLM-authored plan in review, check it into version control, and treat it as a reproducible artifact — not a fresh generation each time. See Virtual Workflows for the authoring loop.

When to keep DGI in the loop anyway

Compilation is the right default for anything repeated, but not everything should be frozen. Keep reasoning live when the shape of the work genuinely changes per request.

PatternRecommended styleWhy
Repeated, well-understood operation (nightly provision, standard deploy)Compiled compositionStable shape; pay LLM once
One-off or exploratory requestDGI liveNot worth compiling something you run once
Flow whose branching depends on unpredictable runtime dataDGI live, or compiled DAG with conditional stepsDecide whether the variation is structural or just data
High-volume, latency-sensitive automationCompiled compositionRemoves model latency from the hot path entirely

The deeper treatment of when each style fires — and how they interoperate — is in Hybrid Execution Model.

The virtual engine is alpha (versioned in the 0.x range) and its catalog is a snapshot at compile time. If the producer process has an older library than the runtime, compilation can succeed while engine dispatch later fails with an "unknown workflow type" error. Pin library versions across your authoring and runtime environments, or compile in the same process that dispatches. Treat the JSON wire format as the stable contract, not the Python builder API.

Lever 2 — Right-size runner warm pools and capacity

Runners (and the broader execution environments they model) are where idle cost accumulates. A runner that is warm but unused costs you provider minutes for zero work; a cold start costs you pickup latency. Tuning the warm pool is the act of choosing where on that curve you want to sit.

The core tradeoff

 cost  ◄────────────────────────────────► latency
 high warm pool          balanced          scale-to-zero
 (instant pickup,        (some warm,        (cheapest,
  pay for idle)          spill cold)        cold-start delay)

Orkestia scales runner groups with a reconciliation loop that continuously converges the pool toward the group's configured min/max — your primary cost dial. GitHub workflow_job events are a best-effort nudge that triggers an immediate reconcile (queued may scale up, completed may scale down); because the loop, not the webhook, is authoritative, scaling stays correct and bounded even if an event is missed.

{
  "scale": { "min": 0, "max": 10, "idle_timeout_s": 300 }
}
SettingEffect on costEffect on latency
min: 0 (scale to zero)Lowest — no idle runnersCold start on first queued job
min: N (warm floor)Pay for N idle runners alwaysFirst N jobs picked up instantly
max (ceiling)Hard cap on spend — Orkestia never exceeds it even under queue pressureJobs queue once the cap is hit
idle_timeout_sHigher = fewer scale cycles, more idle timeHigher = warmer runners between bursts
Set min to your steady-state concurrency floor (the number of jobs you almost always have in flight) and let the reconcile loop scale up to max for bursts. Reserve a non-zero min only for the groups whose pickup latency users actually feel — batch and nightly groups are usually fine at min: 0.
Bounded scaling is a deliberate cost guardrail, not best-effort: Orkestia will let the queue grow rather than exceed your max. If you would rather burst than wait, raise max — but that is an explicit decision to spend more. Self-hosted runners also scale in seconds-to-minutes, not milliseconds, so they are not the right tool when you need extreme instantaneous burst.

Reap dead capacity automatically

Idle isn't the only waste — orphaned capacity is worse, because it costs money and serves nothing. Runner groups run periodic health checks: unhealthy runners are deregistered from GitHub and the compute reaped so dead runners don't sit on the bill or starve the queue. The reconciler also sweeps zombie registrations left behind by a failed teardown. You get this for free, but it's worth watching the orphaned-runner rate (it should trend to zero) and reacting if it doesn't. See Runner Management and drift detection & self-healing.

Lever 3 — Use Kafka async for throughput, not just durability

Long-running steps — cloud provisioning, drains, AWS waiters, anything that blocks on an external system — must not occupy a synchronous request. Orkestia's canonical async path puts these on the Kafka workflow.transition bus, consumed by the workflow Kafka consumer. This is both a reliability property (work survives pod restarts) and a performance property: it decouples submission rate from execution rate, so a burst of work fans out across consumers instead of serializing behind a request thread.

Practical consequences for cost and speed:

  • Parallelism is the default. Independent workflow groups execute fully in parallel; only operations on the same serialized entity (e.g. the same runner group, guarded by a PG advisory lock) are ordered. Structure work so unrelated items don't share a lock key.
  • No long-poll waiting threads. Waiting on a cloud operation happens via transitions, not a held connection, so you don't pay for idle compute holding a socket.
  • Backpressure is visible. Consumer lag is your throughput signal — if it grows, scale consumers rather than runners.
Long-running work always rides the engine's Kafka backend — there is no second queueing system to provision, monitor, or pay for (an early Celery-based path was fully retired). Consumer lag is the one backpressure signal to watch.

Lever 4 — Use Lumen to find the slow and expensive steps

You can't optimize what you can't see. Lumen is Orkestia's workflow-native observability layer, and it attributes every signal to structural context — workflow, actor, state, trace, span, project, environment, release. That attribution is exactly what lets you answer "which step is costing me" instead of guessing.

The signals that matter for cost and performance work:

Lumen surfaceWhat it tells youOptimization it drives
Traces & spansPer-step duration, slow spans, latency p95/p99Find the step to compile, parallelize, or cache
MetricsAggregate time-series by metric nameTrack cost/latency trends across releases
Error groupsRepeated failures grouped by fingerprintKill retry storms that burn compute
DashboardProject summary, latency and cost signals, top error groupsFirst place to look for a regression

A typical optimization loop:

  1. Open the Lumen dashboard, sort by latency or error volume.
  2. Drill into the slow trace; identify the dominant span.
  3. Decide the lever — is it an LLM call that should be compiled away, a step that should run in a parallel layer, or a flaky step generating retries?
  4. Apply the change; watch the same span's p95 across the next releases to confirm.
Lumen's strength depends on logs arriving with trace_id / span_id and workflow/actor context attached. Missing trace IDs collapse the trace UI into coarse pseudo-traces grouped by workflow, which makes per-step attribution much harder. Instrument emitting context early — it's the difference between "this workflow is slow" and "step 3 of 7 is slow." See Observability with Lumen.
Lumen is alpha. Its self-serve onboarding and control plane ship today, but standalone packaging is still maturing. The latency/cost signals are real and usable for the optimization loop above; treat any per-step cost number as directional, and use the external catalog for authoritative per-workflow detail.

Beware ingest rules that hide the signal

Lumen's ingest rules can drop or sample logs to control observability cost. That's a legitimate lever — but over-aggressive sampling can discard the very spans you need to find a slow step, and bad fingerprint rules can mask a retry storm by mis-grouping it. Tune ingest rules with the same care as the workflows they observe.

Putting it together — an optimization checklist

Compile what repeats

Freeze any stable, repeated DGI-designed flow into a deterministic composition. Amortize the LLM cost to near zero.

Right-size warm pools

Set min to your real concurrency floor, cap max for cost safety, let the reconcile loop handle bursts.

Fan out over Kafka

Push long-running work async; parallelize unrelated groups; scale consumers on lag.

Measure with Lumen

Find the dominant span, confirm every change against p95 over subsequent releases.

Tradeoffs at a glance

LeverYou gainYou give up
Compile to compositionLower LLM cost, lower latency, reproducibilityRuntime adaptability; must recompile when the design changes
Non-zero warm pool minInstant pickupPay for idle runners
Aggressive min: 0 / scale-to-zeroLowest idle costCold-start latency on first job
High max ceilingBurst capacityHigher peak spend
Kafka async fan-outThroughput, durabilityMore moving parts to monitor (consumer lag)
Lumen ingest samplingLower observability costRisk of dropping the signal you need