Orkestia
Blog
Guides

Runner Management & Provisioning

Provision, scale, monitor, and decommission self-hosted GitHub Actions runners that live entirely in your own cloud

This guide walks through the full operational lifecycle of an Orkestia runner group — from picking a kind and provisioning compute, through warm-pool capacity tuning and health monitoring, to clean decommissioning. For the conceptual overview, start with Runners concepts. For the catalog of every backend_type / purpose / integration_type, see Runner groups.

The defining property to keep in mind throughout: the runner compute always lives in your cloud account. Orkestia is the control plane — it provisions, registers, scales, and reaps runners — but it never hosts the compute and never sees your job contents. The runner binary phones home to GitHub directly; Orkestia mints short-lived registration tokens and observes. This is the Zero Code Custody posture applied to CI: your builds run on infrastructure you own and pay for.

Runners is GA end-to-end for AWS, and Azure and Kubernetes groups run production fleets today (they carry Orkestia's own agents). GCP, DigitalOcean, and Magalu Cloud are progressively maturing through their provider-specific workflow libraries — treat those as beta. Per-provider coverage is tracked in the workflow reference.

Mental model: two object families

Everything in Runner Management reduces to two first-class entities:

EntityWhat it isLifecycle
Runner group (environment)The provisioned compute substrate — a Kubernetes namespace, an ECS/Fargate or EC2 group, etc. — bound to one GitHub org, one target cloud, a scaling policy, and a network placementdraft → provisioning → active → scaling → archived
Runner executionA single launched runner process serving jobs on a grouplaunched → serving → drained / reaped

The control loop you operate is: provision environment → register runners with GitHub → launch executions → serve jobs → scale → drift-repair → deregister. Each arrow is a runner.* workflow on the Orkestia engine; you rarely call them by hand (the Runners UI and SDK do), but understanding the chain is what makes monitoring and recovery legible.

Prerequisites

Before you can provision a group you need:

A cloud connection

Runner compute lands in your cloud, so Orkestia needs a scoped connection to it. AWS is the GA path — see AWS Connections. Other providers connect through the same connection model.

A GitHub connection (App)

Orkestia uses a GitHub App to mint short-lived runner registration tokens and to receive workflow_job events. One connection per GitHub org.

A network profile

A profile selects the VPC, subnets, and security groups your runners live in. Required for VPC/region-constrained placement.

A runner image source

Runner containers are pulled from your Aggregated Registry. The provider must be able to pull the image at scale-up time.

Connections are the canonical "prerequisite" in Orkestia. If you provision through the MCP or SDK and a workflow reports has_prerequisites: true, fetch the setup guide first — it returns Orkestia's own cloud principal pre-filled so you can grant exactly the trust needed and nothing more. See Security & compliance.

Step 1 — Choose a kind

The create wizard’s Runner type step is not “pick a cloud.” It is backend_type, filtered by the connection you selected. Purpose (github_actions / gitlab_runner / agent / generic) and integration (github / gitlab / none) are a separate choice — Purposes & integrations.

CloudKinds (backend_type)StatusWhen to choose
AWSfargate, ec2_auto_scaling, ec2_vmGAProduction CI; tightest drift repair on Fargate / EC2 VM
Azureazure_container_apps_job, azure_vmss, azure_vmProductionAzure-resident orgs
KuberneteskubernetesProductionAny conformant cluster (including EKS). Do not create eks groups
GCPcloud_run, gceBetaWorkloads already on GCP
DigitalOceando_app_job, do_dropletBetaLightweight / cost-sensitive setups
Magalu Cloudmgc_vmBetaBrazil-region data-residency
NonedevkitProduction (coding)Laptop / hosted provider-blind coding broker
For GCP, DigitalOcean, and Magalu, end-to-end provisioning, reconcile-loop scaling, multi-region, and health-check auto-reap have partial coverage in the provider libraries. Verify current per-kind capability in the workflow reference and the kind pages before committing a production pipeline to a beta kind.

A group binds to one kind and one connection. There is no cross-cloud pool — a job cannot transparently spill from AWS to GCP within a single group. Multi-cloud orgs run multiple groups from the one control plane. Required config keys per kind come from data.runner.list-provider-config-specs.

Step 2 — Create a runner group

The fastest path is the Runners command center UI (the dedicated console at runners.orkestia.dev), which walks you through GitHub org → cloud target → scaling policy → network placement and submits the provisioning DAG for you. The same provisioning is available programmatically — useful for IaC-driven group creation.

A provisioning request describes the GitHub binding, the cloud target, the scaling policy, and the labels GitHub will route jobs by:

{
  "workflow": "runner.environment-provision-kubernetes",
  "input": {
    "organization_id": "org_01HXZ...",
    "github_connection_id": "cn_gh_main",
    "github_org": "myorg",
    "target": {
      "provider": "kubernetes",
      "env": "in-cluster",
      "connection_id": "cn_k8s_prod",
      "namespace": "ci-runners",
      "network_profile_id": "np_01HXZ..."
    },
    "scale": { "min": 0, "max": 10, "idle_timeout_s": 300 },
    "labels": ["self-hosted", "orkestia", "prod-builds"]
  }
}
Exact workflow names, input field names, and provider env values are owned by the live registry, not this guide. Discover the current runner.* catalog and its schemas via the MCP (list_workflow_types(prefix="runner.") → get_workflow_schema(...)) or the workflow types registry, and treat the JSON above as a shape, not a contract.

What happens after you submit:

  1. The group enters provisioning — the provider's environment DAG runs (runner.group-creation / runner.environment-provision-*), standing up the ECS service / EC2 group / namespace inside your cloud.
  2. Orkestia mints a short-lived registration token via the GitHub App.
  3. Runner executions launch and register themselves with GitHub, labelled per the group config.
  4. The group goes active and starts picking up jobs that match its labels.

Registration is pull, not push: Orkestia never pushes runners onto GitHub. It hands each runner binary a short-lived token and the runner calls GitHub itself — so no long-lived secret sits on disk.

Step 3 — Configure warm pools & capacity

Capacity is governed by the group's scaling policy. The two knobs that matter most:

SettingEffectTrade-off
minWarm pool — runners kept ready even when idleHigher min = faster pickup, higher idle cost
maxHard ceiling on concurrent runnersProtects your cloud bill; queue waits rather than overspending
idle_timeout_sHow long an idle runner lives before being reapedLower = cheaper; higher = fewer cold starts under bursty load

Set min: 0 for spiky, cost-sensitive pipelines (every runner is on-demand, accepting a cold-start delay). Set min to your steady-state concurrency for latency-sensitive CI where developers wait on builds.

Scaling is bounded by group policy — always. Orkestia will not exceed max even when the queue grows. This is deliberate: protecting your cloud spend is an explicit guarantee, not best-effort. If jobs are queuing, raise max rather than expecting auto-burst.

How scaling actually fires

Scaling is driven by a reconciliation loop, not by webhooks. A per-group controller continuously converges the live pool toward your min/max — the same reconcile model Orkestia uses for drift & self-healing. GitHub workflow_job webhooks are a best-effort nudge that triggers an immediate reconcile, never the source of truth:

  • workflow_job.queued → nudge: reconcile may scale up (bounded by max)
  • workflow_job.completed / idle timeout → nudge: reconcile may scale down (bounded by min)

Because the loop — not the webhook — is authoritative, a missed or delayed workflow_job event never strands a group: the next reconcile pass converges it anyway. All long-running provisioning, scale, and drain steps run asynchronously over the platform's Kafka bus (workflow.transition), so a scale event never blocks an API call.

Operations on the same group (provision, scale, teardown) are serialized by a Postgres advisory lock keyed on the group ID, so they can't race. Different groups scale fully in parallel.

Step 4 — Register the build/runner app

Two registrations connect a group to your CI:

  1. The GitHub App (done once per org, at connection time) authorizes Orkestia to mint registration tokens and subscribe to workflow_job / workflow_run events. This is set up when you create the GitHub connection (see Cloud connections).
  2. The job labels. Orkestia registers each runner with the labels you set on the group (e.g. self-hosted, orkestia, prod-builds). To route a job to the group, target those labels in your workflow YAML:
jobs:
  build:
    runs-on: [self-hosted, orkestia, prod-builds]
    steps:
      - uses: actions/checkout@v4
      - run: make build

When a matching job is queued, GitHub emits workflow_job.queued, Orkestia (via the dispatch/scaling workflows) decides whether a new runner is needed, and an execution picks the job up. No runner endpoint is exposed to you — the binding is entirely through GitHub labels.

In local development, GitHub workflow_job events reach your machine through the DevKit CLI's hook redirect (see hooks in Platform Services) rather than a public webhook URL, so you can exercise scaling against a real org without a public endpoint.

Step 5 — Monitor health

Runner activity is surfaced as workflow.transition events on the same Kafka bus as the rest of the platform — there is no runners-specific log silo. The signals you watch:

SignalKindFires when
runner.group.provisionedeventGroup reaches active
runner.group.scaled_up / scaled_downevent + metricA scale DAG completes
runner.group.unhealthyeventA health check fails for the group
runner.job.picked_upeventA group runner starts a job (pickup-latency metric)
runner.job.completedeventA job finishes (utilization metric)
runner.group.archivedeventTeardown completes

Health checks are periodic liveness probes. When a runner fails one, Orkestia deregisters it from GitHub and reaps the compute, so a dead runner never sits in the pool starving the queue.

These signals are first-class in Lumen observability — pickup latency, utilization (% of time runners are busy), scale-event counts, and failure rate by environment all roll up there. See the observability guide for building runner dashboards and alerts.

Drift detection & self-healing

Orkestia continuously reconciles the desired group state against what actually exists in your cloud. Orphaned runners (compute gone but GitHub still shows them, or vice versa) are swept on the next health check. Read the deep dive for how reconciliation and self-healing work.

Hybrid execution model

Runner provisioning and scaling are deterministic compiled workflows. Understand how Orkestia compiles DAGs and runs them efficiently.

Common failure modes & what to do

SymptomLikely causeAction
New runners can't join, existing ones keep workingGitHub App token expired/revokedRe-check the GitHub connection; re-authorize the App
Scale-up DAG fails with a quota errorProvider quota hit (ECS tasks, EC2 instances, Cloud Run max-instances)Raise the cloud quota; existing runners are unaffected
GitHub shows zombie offline runnersTeardown failed after destroying compute but before deregistrationWait for the reconciler sweep on next health check, or re-run teardown
A scale decision didn't happenworkflow_job webhook lost (network blip)The reconcile loop converges the group on its next pass — a lost webhook self-heals, not a stuck group
Provisioning fails at placementNetwork profile references a deleted VPC/subnetFix the network profile; active groups are unaffected
Scale-up fails at image pullProvider can't reach the runner imageCheck Aggregated Registry reachability; retryable

Step 6 — Decommission a group

When you no longer need a group, archive it. Teardown drains in-flight jobs, revokes the GitHub registrations, and reaps the cloud resources — leaving no orphaned compute and no zombie runners on GitHub.

Because teardown is serialized per group by the advisory lock, you can safely archive one group while others keep serving. If teardown is interrupted mid-flight, the self-healing reconciler detects the partial state and completes the cleanup on the next sweep.

Limits & constraints

  • GitHub Actions is the GA job source. GitLab is a first-class integration_type in the create wizard (purpose=gitlab_runner); Bitbucket and Jenkins are not in scope. See Purposes.
  • One GitHub org per group — cross-org runner sharing isn't modeled; multi-org users create multiple groups.
  • Compute stays in your cloud — a broken cloud connection means a broken group, not an Orkestia fallback. You pay the provider directly.
  • No cross-cloud pool — a group targets exactly one environment.
  • Scaling is bounded by group policy — Orkestia never exceeds max, even under queue pressure.
  • Not a millisecond-burst replacement — self-hosted runners scale in seconds-to-minutes, not the instant elasticity of GitHub-hosted ephemeral runners.

Where to go next

Runners concepts

The conceptual model: groups, executions, and the control-plane posture.

Runner group kinds

Every backend_type, required config, and purpose pairing.

Runner reference

Browse the live runner.* workflow catalog and per-provider capability detail.

Observability with Lumen

Build dashboards and alerts on runner pickup latency and utilization.

Security & compliance

How Zero Code Custody and scoped connections apply to runner compute.