Runner Management & Provisioning
This guide walks through the full operational lifecycle of an Orkestia runner group — from picking a kind and provisioning compute, through warm-pool capacity tuning and health monitoring, to clean decommissioning. For the conceptual overview, start with Runners concepts. For the catalog of every backend_type / purpose / integration_type, see Runner groups.
The defining property to keep in mind throughout: the runner compute always lives in your cloud account. Orkestia is the control plane — it provisions, registers, scales, and reaps runners — but it never hosts the compute and never sees your job contents. The runner binary phones home to GitHub directly; Orkestia mints short-lived registration tokens and observes. This is the Zero Code Custody posture applied to CI: your builds run on infrastructure you own and pay for.
Mental model: two object families
Everything in Runner Management reduces to two first-class entities:
| Entity | What it is | Lifecycle |
|---|---|---|
| Runner group (environment) | The provisioned compute substrate — a Kubernetes namespace, an ECS/Fargate or EC2 group, etc. — bound to one GitHub org, one target cloud, a scaling policy, and a network placement | draft → provisioning → active → scaling → archived |
| Runner execution | A single launched runner process serving jobs on a group | launched → serving → drained / reaped |
The control loop you operate is: provision environment → register runners with GitHub → launch executions → serve jobs → scale → drift-repair → deregister. Each arrow is a runner.* workflow on the Orkestia engine; you rarely call them by hand (the Runners UI and SDK do), but understanding the chain is what makes monitoring and recovery legible.
- draftconfig only
- provisioning
- active
- scaling
- archived
- draft→ provision DAG →provisioning
- provisioning→ runners register with GitHub →active
- active→ workflow_job events →scaling
- scaling→active
- active→ drain + teardown →archived
Prerequisites
Before you can provision a group you need:
A cloud connection
Runner compute lands in your cloud, so Orkestia needs a scoped connection to it. AWS is the GA path — see AWS Connections. Other providers connect through the same connection model.
A GitHub connection (App)
Orkestia uses a GitHub App to mint short-lived runner registration tokens and to receive workflow_job events. One connection per GitHub org.
has_prerequisites: true, fetch the setup guide first — it returns Orkestia's own cloud principal pre-filled so you can grant exactly the trust needed and nothing more. See Security & compliance.Step 1 — Choose a kind
The create wizard’s Runner type step is not “pick a cloud.” It is backend_type, filtered by the connection you selected. Purpose (github_actions / gitlab_runner / agent / generic) and integration (github / gitlab / none) are a separate choice — Purposes & integrations.
| Cloud | Kinds (backend_type) | Status | When to choose |
|---|---|---|---|
| AWS | fargate, ec2_auto_scaling, ec2_vm | GA | Production CI; tightest drift repair on Fargate / EC2 VM |
| Azure | azure_container_apps_job, azure_vmss, azure_vm | Production | Azure-resident orgs |
| Kubernetes | kubernetes | Production | Any conformant cluster (including EKS). Do not create eks groups |
| GCP | cloud_run, gce | Beta | Workloads already on GCP |
| DigitalOcean | do_app_job, do_droplet | Beta | Lightweight / cost-sensitive setups |
| Magalu Cloud | mgc_vm | Beta | Brazil-region data-residency |
| None | devkit | Production (coding) | Laptop / hosted provider-blind coding broker |
A group binds to one kind and one connection. There is no cross-cloud pool — a job cannot transparently spill from AWS to GCP within a single group. Multi-cloud orgs run multiple groups from the one control plane. Required config keys per kind come from data.runner.list-provider-config-specs.
Step 2 — Create a runner group
The fastest path is the Runners command center UI (the dedicated console at runners.orkestia.dev), which walks you through GitHub org → cloud target → scaling policy → network placement and submits the provisioning DAG for you. The same provisioning is available programmatically — useful for IaC-driven group creation.
A provisioning request describes the GitHub binding, the cloud target, the scaling policy, and the labels GitHub will route jobs by:
{
"workflow": "runner.environment-provision-kubernetes",
"input": {
"organization_id": "org_01HXZ...",
"github_connection_id": "cn_gh_main",
"github_org": "myorg",
"target": {
"provider": "kubernetes",
"env": "in-cluster",
"connection_id": "cn_k8s_prod",
"namespace": "ci-runners",
"network_profile_id": "np_01HXZ..."
},
"scale": { "min": 0, "max": 10, "idle_timeout_s": 300 },
"labels": ["self-hosted", "orkestia", "prod-builds"]
}
}
env values are owned by the live registry, not this guide. Discover the current runner.* catalog and its schemas via the MCP (list_workflow_types(prefix="runner.") → get_workflow_schema(...)) or the workflow types registry, and treat the JSON above as a shape, not a contract.What happens after you submit:
- The group enters provisioning — the provider's environment DAG runs (
runner.group-creation/runner.environment-provision-*), standing up the ECS service / EC2 group / namespace inside your cloud. - Orkestia mints a short-lived registration token via the GitHub App.
- Runner executions launch and register themselves with GitHub, labelled per the group config.
- The group goes active and starts picking up jobs that match its labels.
Registration is pull, not push: Orkestia never pushes runners onto GitHub. It hands each runner binary a short-lived token and the runner calls GitHub itself — so no long-lived secret sits on disk.
Step 3 — Configure warm pools & capacity
Capacity is governed by the group's scaling policy. The two knobs that matter most:
| Setting | Effect | Trade-off |
|---|---|---|
min | Warm pool — runners kept ready even when idle | Higher min = faster pickup, higher idle cost |
max | Hard ceiling on concurrent runners | Protects your cloud bill; queue waits rather than overspending |
idle_timeout_s | How long an idle runner lives before being reaped | Lower = cheaper; higher = fewer cold starts under bursty load |
Set min: 0 for spiky, cost-sensitive pipelines (every runner is on-demand, accepting a cold-start delay). Set min to your steady-state concurrency for latency-sensitive CI where developers wait on builds.
max even when the queue grows. This is deliberate: protecting your cloud spend is an explicit guarantee, not best-effort. If jobs are queuing, raise max rather than expecting auto-burst.How scaling actually fires
Scaling is driven by a reconciliation loop, not by webhooks. A per-group controller continuously converges the live pool toward your min/max — the same reconcile model Orkestia uses for drift & self-healing. GitHub workflow_job webhooks are a best-effort nudge that triggers an immediate reconcile, never the source of truth:
workflow_job.queued→ nudge: reconcile may scale up (bounded bymax)workflow_job.completed/ idle timeout → nudge: reconcile may scale down (bounded bymin)
Because the loop — not the webhook — is authoritative, a missed or delayed workflow_job event never strands a group: the next reconcile pass converges it anyway. All long-running provisioning, scale, and drain steps run asynchronously over the platform's Kafka bus (workflow.transition), so a scale event never blocks an API call.
Step 4 — Register the build/runner app
Two registrations connect a group to your CI:
- The GitHub App (done once per org, at connection time) authorizes Orkestia to mint registration tokens and subscribe to
workflow_job/workflow_runevents. This is set up when you create the GitHub connection (see Cloud connections). - The job labels. Orkestia registers each runner with the labels you set on the group (e.g.
self-hosted,orkestia,prod-builds). To route a job to the group, target those labels in your workflow YAML:
jobs:
build:
runs-on: [self-hosted, orkestia, prod-builds]
steps:
- uses: actions/checkout@v4
- run: make build
When a matching job is queued, GitHub emits workflow_job.queued, Orkestia (via the dispatch/scaling workflows) decides whether a new runner is needed, and an execution picks the job up. No runner endpoint is exposed to you — the binding is entirely through GitHub labels.
workflow_job events reach your machine through the DevKit CLI's hook redirect (see hooks in Platform Services) rather than a public webhook URL, so you can exercise scaling against a real org without a public endpoint.Step 5 — Monitor health
Runner activity is surfaced as workflow.transition events on the same Kafka bus as the rest of the platform — there is no runners-specific log silo. The signals you watch:
| Signal | Kind | Fires when |
|---|---|---|
runner.group.provisioned | event | Group reaches active |
runner.group.scaled_up / scaled_down | event + metric | A scale DAG completes |
runner.group.unhealthy | event | A health check fails for the group |
runner.job.picked_up | event | A group runner starts a job (pickup-latency metric) |
runner.job.completed | event | A job finishes (utilization metric) |
runner.group.archived | event | Teardown completes |
Health checks are periodic liveness probes. When a runner fails one, Orkestia deregisters it from GitHub and reaps the compute, so a dead runner never sits in the pool starving the queue.
These signals are first-class in Lumen observability — pickup latency, utilization (% of time runners are busy), scale-event counts, and failure rate by environment all roll up there. See the observability guide for building runner dashboards and alerts.
Drift detection & self-healing
Orkestia continuously reconciles the desired group state against what actually exists in your cloud. Orphaned runners (compute gone but GitHub still shows them, or vice versa) are swept on the next health check. Read the deep dive for how reconciliation and self-healing work.
Common failure modes & what to do
| Symptom | Likely cause | Action |
|---|---|---|
| New runners can't join, existing ones keep working | GitHub App token expired/revoked | Re-check the GitHub connection; re-authorize the App |
| Scale-up DAG fails with a quota error | Provider quota hit (ECS tasks, EC2 instances, Cloud Run max-instances) | Raise the cloud quota; existing runners are unaffected |
| GitHub shows zombie offline runners | Teardown failed after destroying compute but before deregistration | Wait for the reconciler sweep on next health check, or re-run teardown |
| A scale decision didn't happen | workflow_job webhook lost (network blip) | The reconcile loop converges the group on its next pass — a lost webhook self-heals, not a stuck group |
| Provisioning fails at placement | Network profile references a deleted VPC/subnet | Fix the network profile; active groups are unaffected |
| Scale-up fails at image pull | Provider can't reach the runner image | Check Aggregated Registry reachability; retryable |
Step 6 — Decommission a group
When you no longer need a group, archive it. Teardown drains in-flight jobs, revokes the GitHub registrations, and reaps the cloud resources — leaving no orphaned compute and no zombie runners on GitHub.
- active
- draining
- teardown
- archived
- active→ stop accepting new jobs →draining
- draining→ in-flight jobs complete →teardown
- teardown→ GitHub regs revoked, compute reaped →archived
Because teardown is serialized per group by the advisory lock, you can safely archive one group while others keep serving. If teardown is interrupted mid-flight, the self-healing reconciler detects the partial state and completes the cleanup on the next sweep.
Limits & constraints
- GitHub Actions is the GA job source. GitLab is a first-class
integration_typein the create wizard (purpose=gitlab_runner); Bitbucket and Jenkins are not in scope. See Purposes. - One GitHub org per group — cross-org runner sharing isn't modeled; multi-org users create multiple groups.
- Compute stays in your cloud — a broken cloud connection means a broken group, not an Orkestia fallback. You pay the provider directly.
- No cross-cloud pool — a group targets exactly one environment.
- Scaling is bounded by group policy — Orkestia never exceeds
max, even under queue pressure. - Not a millisecond-burst replacement — self-hosted runners scale in seconds-to-minutes, not the instant elasticity of GitHub-hosted ephemeral runners.
Where to go next
Creating & Exposing Virtual Workflows
Compose existing workflows into a validated, versioned virtual workflow and expose it to your app's end-users through App Enablement
Security & Compliance
How Orkestia keeps customer code and data out of its custody, isolates every tenant by construction, and gives evaluators an audit trail they can hand to a reviewer
