Runners & Execution Environments
TL;DR
- A runner is where work physically runs, and it is always in your cloud. Orkestia provisions, registers, scales, and heals it. Compute, data, and egress stay with you.
- Two entities: a runner group (the long-lived, configured pool) and runner executions (launched runners that come and go with demand).
- Scaling is a reconcile loop, bounded by
minandmax. Webhooks nudge it; they are never the source of truth. - AWS, Azure, and Kubernetes run production fleets today. GCP, DigitalOcean, and Magalu Cloud are beta.
- Everything is a
runner.*workflow, so an assistant over MCP can provision, inspect, and scale groups like any other capability.
Control plane vs execution plane
- Runner control planeprovision · register · scale · heal
- Workflow state + Lumenstate and observability only
- Runner groupECS / AKS / K8s / GCE / Droplet / MGC VM
- runner exec
- runner exec
- warm pool
- Job source (GitHub Actions)
- Runner control plane→ short-lived token + DAG →Runner group
- Runner group→runner exec
- Runner group→runner exec
- Runner group→warm pool
- runner exec→ register + poll jobs →Job source (GitHub Actions)
- runner exec→ transition signals →Workflow state + Lumen
| Plane | Owns | Lives in | Holds your code or data? |
|---|---|---|---|
| Control plane | Provisioning DAGs, scaling decisions, registration, drift repair, run state | Orkestia | No. Only workflow state and observability |
| Execution plane | The compute, the runner binary, the job workload | Your cloud account | Yes, and it never leaves |
The control plane drives runners through runner.* workflows on the engine. Because those are ordinary workflows, provisioning a fleet is just another DAG with full history, per-group locking, and async transitions.
The two core entities
Runner group
The provisioned compute substrate: a binding of job-source org, target cloud environment, scaling policy, and network placement. The long-lived thing you configure.
Runner execution
A single launched runner serving jobs on a group. Executions come and go with demand; the group persists.
Kinds (backend × purpose × integration)
A group is not “an AWS runner.” It stores three enums:
| Axis | Field | Examples |
|---|---|---|
| Backend kind | backend_type | fargate, ec2_vm, kubernetes, azure_container_apps_job, devkit, … |
| Purpose | purpose | github_actions, gitlab_runner, agent, generic |
| Integration | integration_type | github, gitlab, none |
The create wizard’s Runner type step is the backend-kind list, filtered by the cloud connection you picked. Full catalog, required config keys, and per-kind pages: Runner groups.
| Provider | Kinds | Status |
|---|---|---|
| AWS | fargate, ec2_auto_scaling, ec2_vm | Production |
| Kubernetes | kubernetes (use this even on EKS; eks is a legacy alias) | Production, carries live fleets today |
| Azure | azure_container_apps_job, azure_vmss, azure_vm | Production, carries live fleets today |
| GCP | cloud_run, gce | Beta, partial coverage |
| DigitalOcean | do_app_job, do_droplet | Beta, partial coverage |
| Magalu Cloud | mgc_vm | Beta, partial coverage |
| None (laptop / hosted broker) | devkit | Production for coding agents |
purpose=agent is a separate gate — see Purposes.Connect a cloud first (AWS connections, Cloud connections) and pick placement through network profiles. A broken cloud connection means a broken group. There is deliberately no Orkestia compute fallback.
Warm pools and capacity
A cold runner has to provision compute, pull an image, and register before it picks up a job. Warm pools keep a floor of ready runners so jobs start immediately. Capacity is governed per group:
{
"scale": { "min": 0, "max": 10, "idle_timeout_s": 300 },
"labels": ["self-hosted", "orkestia", "prod-builds"]
}
min: the warm floor.0means scale to zero when idle.max: the hard ceiling. Orkestia never exceeds it, even if the queue grows.idle_timeout_s: how long a runner sits idle before it is reaped.
min against your pickup-latency target and max against your cost ceiling. Scale to zero suits bursty, cost-sensitive work. A warm floor suits latency-sensitive pipelines.Reconcile-loop scaling
A per-group controller continuously compares the live pool against your policy and converges it toward min and max. Job-source workflow_job events are a best-effort nudge that makes the loop react immediately. They are never the source of truth.
sequenceDiagram participant GH as Job source participant CP as Runner control plane participant Cloud as Your cloud Note over CP: reconcile loop converges pool → min/max (continuous) GH->>CP: workflow_job.queued (webhook nudge) CP->>CP: reconcile (queued vs busy, bounded by max) CP->>Cloud: launch runner execution (async transition) Cloud->>GH: register + poll for job GH->>CP: workflow_job.completed (nudge) CP->>CP: reconcile: maybe scale down after idle_timeout
Registration is pull, not push: Orkestia hands the runner a short-lived registration token and the runner registers itself with the job source. No long-lived secrets sit on disk.
Group lifecycle
| State | What it means |
|---|---|
| draft | Config exists; no cloud resources yet |
| provisioning | The provision DAG is running |
| active | Runners registered; executions picking up jobs |
| scaling | A scale DAG is in flight; the group still serves jobs |
| archived | Drain and teardown complete; registrations revoked, resources reaped |
Drift detection and self-healing
Runners are long-lived infrastructure in an account you also touch directly. Reality drifts from intent: a runner dies, a token expires, a teardown half-completes. Periodic health checks detect drift and self-heal: unhealthy runners are deregistered and reaped, orphaned registrations are swept, failed scale-ups are retried.
Common cases the healer handles: expired registration tokens (new runners cannot join; existing ones keep working), provider quota at scale-up (queue backs up; existing runners unaffected), orphaned runners after a partial teardown, unreachable runner images.
Observability
Every runner action emits structured signals on the same transition bus the rest of the platform uses. Group provisioned, scaled up, scaled down, unhealthy, archived, and per-job pickup and completion all flow into Lumen.
| Signal | Feeds |
|---|---|
runner.group.scaled_up / scaled_down | Cost and SLO dashboards |
runner.group.unhealthy | Alerting |
runner.job.picked_up / completed | Pickup latency and utilization |
Runners vs agent runner groups
RunnerGroup row, different purpose=agent. Related infrastructure, different job. CI kinds are this page; agent eligibility is Staff & Agents; every kind is listed under Runner groups.Ask your AI assistant
List my runner groups with provider, state, and scaling policy.
Show me the schema for provisioning a Kubernetes runner group. If it has prerequisites, fetch them.
Which of my runner groups are unhealthy, and what does the last health check say?
Scale group <group_id> to min 2, max 8. Show me the plan and wait for my confirmation.
For AI agents
| Rule | Detail |
|---|---|
Everything is runner.* | Discover with list_workflow_types(prefix="runner."). Provision, scale, and teardown are mutations: confirm first. |
| Connections first | Runner provisioning needs a cloud connection. Honor has_prerequisites. |
| Stalled provisioning | list_stuck_workflows(workflow_type="runner.…") surfaces stalled DAGs. Read history before retrying. |
Never exceed max | The platform will not, and neither should a plan you propose. |
Status
AWS, Azure, and Kubernetes run production fleets; GCP, DigitalOcean, and Magalu Cloud are beta. Job sources: GitHub Actions (GA) and GitLab (purpose=gitlab_runner). Kind catalog: Runner groups.
Agents — the execution substrate
The primitives behind every AI worker on Orkestia, agent configs, workflow-backed skills, per-agent MCP servers, sessions on your runners, memory, budgets, and end-user agents
Lumen Observability
Orkestia's telemetry store and triage engine, JSON HTTP ingest, SHA-256 error groups, traces, metrics, the query API, and the Lumen MCP server for assistant-driven triage
