Orkestia
Blog
Core Concepts

Runners & Execution Environments

Managed execution environments provisioned inside your own cloud where workflow steps, CI jobs, and agent sessions run, runner groups, warm pools, reconcile-loop scaling, and the control-plane / execution split

TL;DR

  • A runner is where work physically runs, and it is always in your cloud. Orkestia provisions, registers, scales, and heals it. Compute, data, and egress stay with you.
  • Two entities: a runner group (the long-lived, configured pool) and runner executions (launched runners that come and go with demand).
  • Scaling is a reconcile loop, bounded by min and max. Webhooks nudge it; they are never the source of truth.
  • AWS, Azure, and Kubernetes run production fleets today. GCP, DigitalOcean, and Magalu Cloud are beta.
  • Everything is a runner.* workflow, so an assistant over MCP can provision, inspect, and scale groups like any other capability.
Compute lives in the customer cloud. The runner binary talks to its job source (GitHub), not to Orkestia. Orkestia observes state and emits signals. See Security & compliance.

Control plane vs execution plane

PlaneOwnsLives inHolds your code or data?
Control planeProvisioning DAGs, scaling decisions, registration, drift repair, run stateOrkestiaNo. Only workflow state and observability
Execution planeThe compute, the runner binary, the job workloadYour cloud accountYes, and it never leaves

The control plane drives runners through runner.* workflows on the engine. Because those are ordinary workflows, provisioning a fleet is just another DAG with full history, per-group locking, and async transitions.

The two core entities

Runner group

The provisioned compute substrate: a binding of job-source org, target cloud environment, scaling policy, and network placement. The long-lived thing you configure.

Runner execution

A single launched runner serving jobs on a group. Executions come and go with demand; the group persists.

A group binds to exactly one job-source org. Multi-org users create multiple groups. Cross-org runner sharing is not modeled.

Kinds (backend × purpose × integration)

A group is not “an AWS runner.” It stores three enums:

AxisFieldExamples
Backend kindbackend_typefargate, ec2_vm, kubernetes, azure_container_apps_job, devkit, …
Purposepurposegithub_actions, gitlab_runner, agent, generic
Integrationintegration_typegithub, gitlab, none

The create wizard’s Runner type step is the backend-kind list, filtered by the cloud connection you picked. Full catalog, required config keys, and per-kind pages: Runner groups.

ProviderKindsStatus
AWSfargate, ec2_auto_scaling, ec2_vmProduction
Kuberneteskubernetes (use this even on EKS; eks is a legacy alias)Production, carries live fleets today
Azureazure_container_apps_job, azure_vmss, azure_vmProduction, carries live fleets today
GCPcloud_run, gceBeta, partial coverage
DigitalOceando_app_job, do_dropletBeta, partial coverage
Magalu Cloudmgc_vmBeta, partial coverage
None (laptop / hosted broker)devkitProduction for coding agents
AWS, Azure, and Kubernetes are production paths. Azure and Kubernetes groups carry Orkestia's own agent fleet. Treat GCP, DigitalOcean, and Magalu as beta and check reference.orkestia.dev for the current per-kind capability. purpose=agent is a separate gate — see Purposes.

Connect a cloud first (AWS connections, Cloud connections) and pick placement through network profiles. A broken cloud connection means a broken group. There is deliberately no Orkestia compute fallback.

Warm pools and capacity

A cold runner has to provision compute, pull an image, and register before it picks up a job. Warm pools keep a floor of ready runners so jobs start immediately. Capacity is governed per group:

{
  "scale": { "min": 0, "max": 10, "idle_timeout_s": 300 },
  "labels": ["self-hosted", "orkestia", "prod-builds"]
}
  • min: the warm floor. 0 means scale to zero when idle.
  • max: the hard ceiling. Orkestia never exceeds it, even if the queue grows.
  • idle_timeout_s: how long a runner sits idle before it is reaped.
Tune min against your pickup-latency target and max against your cost ceiling. Scale to zero suits bursty, cost-sensitive work. A warm floor suits latency-sensitive pipelines.

Reconcile-loop scaling

A per-group controller continuously compares the live pool against your policy and converges it toward min and max. Job-source workflow_job events are a best-effort nudge that makes the loop react immediately. They are never the source of truth.

sequenceDiagram
  participant GH as Job source
  participant CP as Runner control plane
  participant Cloud as Your cloud
  Note over CP: reconcile loop converges pool → min/max (continuous)
  GH->>CP: workflow_job.queued (webhook nudge)
  CP->>CP: reconcile (queued vs busy, bounded by max)
  CP->>Cloud: launch runner execution (async transition)
  Cloud->>GH: register + poll for job
  GH->>CP: workflow_job.completed (nudge)
  CP->>CP: reconcile: maybe scale down after idle_timeout

Registration is pull, not push: Orkestia hands the runner a short-lived registration token and the runner registers itself with the job source. No long-lived secrets sit on disk.

If a webhook is lost, the loop still converges on its next pass. Provisioning, scaling, and teardown of the same group are serialized on a per-group lock. Different groups scale in parallel.

Group lifecycle

StateWhat it means
draftConfig exists; no cloud resources yet
provisioningThe provision DAG is running
activeRunners registered; executions picking up jobs
scalingA scale DAG is in flight; the group still serves jobs
archivedDrain and teardown complete; registrations revoked, resources reaped

Drift detection and self-healing

Runners are long-lived infrastructure in an account you also touch directly. Reality drifts from intent: a runner dies, a token expires, a teardown half-completes. Periodic health checks detect drift and self-heal: unhealthy runners are deregistered and reaped, orphaned registrations are swept, failed scale-ups are retried.

Common cases the healer handles: expired registration tokens (new runners cannot join; existing ones keep working), provider quota at scale-up (queue backs up; existing runners unaffected), orphaned runners after a partial teardown, unreachable runner images.

Runner group kinds

Every backend_type, required config, and purpose/integration pairing.

Drift detection & self-healing

The reconciler sweep and the failure-mode taxonomy.

Runner management guide

Create a group, set capacity, operate a fleet.

Observability

Every runner action emits structured signals on the same transition bus the rest of the platform uses. Group provisioned, scaled up, scaled down, unhealthy, archived, and per-job pickup and completion all flow into Lumen.

SignalFeeds
runner.group.scaled_up / scaled_downCost and SLO dashboards
runner.group.unhealthyAlerting
runner.job.picked_up / completedPickup latency and utilization

Runners vs agent runner groups

Staff also has "runner groups": cloud pools that host AI agent sessions (the Agents substrate). Same RunnerGroup row, different purpose=agent. Related infrastructure, different job. CI kinds are this page; agent eligibility is Staff & Agents; every kind is listed under Runner groups.

Ask your AI assistant

prompts
List my runner groups with provider, state, and scaling policy.

Show me the schema for provisioning a Kubernetes runner group. If it has prerequisites, fetch them.

Which of my runner groups are unhealthy, and what does the last health check say?

Scale group <group_id> to min 2, max 8. Show me the plan and wait for my confirmation.

For AI agents

RuleDetail
Everything is runner.*Discover with list_workflow_types(prefix="runner."). Provision, scale, and teardown are mutations: confirm first.
Connections firstRunner provisioning needs a cloud connection. Honor has_prerequisites.
Stalled provisioninglist_stuck_workflows(workflow_type="runner.…") surfaces stalled DAGs. Read history before retrying.
Never exceed maxThe platform will not, and neither should a plan you propose.

Status

AWS, Azure, and Kubernetes run production fleets; GCP, DigitalOcean, and Magalu Cloud are beta. Job sources: GitHub Actions (GA) and GitLab (purpose=gitlab_runner). Kind catalog: Runner groups.

Runner group kinds

Fargate, EC2, Kubernetes, Azure, GCP, DigitalOcean, Magalu, DevKit.

Runner management

Provision and operate runner fleets step by step.

Hybrid execution model

How AI-designed workflows compile to deterministic steps that run on your runners.

Full workflow catalog

Per-workflow inputs, flags, and limits for every runner.* capability.