Orkestia
Blog
Advanced Topics

Drift Detection & Self-Healing

How Orkestia reconciles desired and actual state for runners and network infrastructure in your cloud, computes readiness verdicts, and decides what to heal automatically versus surface for approval

Orkestia provisions runners and reads network topology inside your cloud account — but the cloud is a living system. Quotas change, consoles get edited by hand, runners die, webhooks get dropped, and a VPC referenced by a profile gets decommissioned. Drift is the gap between the state Orkestia intends and the state the provider actually reports. This page explains how that gap is detected, how readiness verdicts are computed from it, and the deliberate line between what Orkestia heals on its own and what it surfaces for a human to approve.

This is a control-plane reconciliation model, not an agent loop running inside your account. Per the Zero Code Custody posture, Orkestia stores only workflow state and observability data — the compute being reconciled always lives in your cloud, and Orkestia observes and orchestrates it rather than holding it. See Deployment models.

What "drift" means on Orkestia

There are two distinct surfaces where desired and actual state can diverge, and they reconcile differently:

SurfaceDesired stateActual stateReconciliation style
RunnersGroup config: min/max, labels, network placement, registered runner setLive cloud compute (ECS tasks / EC2 / K8s pods) + GitHub's view of registered runnersActive — reconcile-loop scaling, health-reap, teardown repair
Network ManagementSaved network profiles (VPC + subnets + SGs, versioned)Provider network inventory mirrored by periodic syncDetect-and-surface — Orkestia never mutates customer networks

The asymmetry is intentional. Orkestia owns the lifecycle of the runners it provisions, so it can act on runner drift directly. It does not own your networks — Network Management is pull-only inventory, so network drift is detected and flagged, never silently corrected.

Drift detection and self-healing are GA end-to-end for AWS runner groups, and Azure and Kubernetes groups run production fleets today. Coverage for GCP, DigitalOcean, and Magalu ships in the provider workflow libraries with partial end-to-end paths — treat auto-reap and placement repair on those providers as (beta). See the live catalog at reference.orkestia.dev.

How desired state is expressed

Both surfaces declare desired state as versioned, org-local metadata — never as mutable in-place rows. This is what makes drift computable: there is always a stable reference to diff actual state against.

  • A runner group holds (GitHub org + target cloud env + scaling policy + network placement + labels). It moves through draft → provisioning → active → scaling → archived.
  • A network profile holds (VPC, subnets, security groups, optional public/private intent) and is versioned: every edit bumps vN → vN+1, and an in-flight launch keeps the version it resolved. A profile edit can never retroactively break a running deploy.
{
  "id": "np_01HXZ...",
  "name": "prod-usa-private",
  "version": 3,
  "intent": "private",
  "targets": [
    {
      "provider": "aws",
      "connection_id": "cn_aws_prod",
      "region": "us-east-1",
      "vpc_id": "vpc-0abc123",
      "subnets": ["subnet-0a...", "subnet-0b..."],
      "security_groups": ["sg-0ff..."]
    }
  ],
  "readiness": "green",
  "updated_at": "2026-04-15T18:22:04Z"
}
Because state is event-sourced in the workflow engine, every reconciliation step is an append-only transition with full history. You can replay exactly how a group reached its current shape rather than guessing from logs. See The workflow engine.

Detecting network drift (pull, detect, surface)

Network Management mirrors cloud topology locally so apps don't re-query each provider on every deploy. Drift is the gap between that mirror and reality, and it is detected by re-running the sync.

Key properties of network drift detection:

  • Pull-only, eventually consistent. A network created directly in a cloud console appears only after the next sync — not instantly. Downstream products must tolerate "not yet known."
  • Per-region freshness, not all-or-nothing. A provider error in one region/account does not invalidate prior data elsewhere; the failed slice is marked stale with its last successful discovered_at, surfaced as the network.sync.age_seconds gauge.
  • Silent SG drift is the hard case. Security-group rules edited in the console are only picked up on the next sync. Between syncs the policy-violation view can under-report; on-demand re-sync (UI or MCP) is the mitigation.
  • No mutation, ever. Orkestia is not an IPAM. It does not allocate CIDRs, reserve IPs, or touch route tables. Detected network drift produces a verdict and a notification, not a corrective write.

When a profile's underlying connection or network vanishes, the profile is marked orphaned: existing resolutions remain readable, but new launches are blocked until it is rebound — a deliberate fail-closed choice rather than letting a deploy land somewhere unknown.

Readiness verdicts

Every network and profile carries a deploy-ready verdict — green / yellow / red — so users see "don't pick this" before they pick it. The verdict is a rollup of several signals:

Input signalPushes toward red
Sync freshness (network.sync.age_seconds)Inventory stale beyond threshold
Quota headroomSubnet effectively full / no IP space
Required tagsMandatory governance tags missing
Security-group sanityRules violating baseline (e.g. 0.0.0.0/0 on sensitive ports)
Connection healthSource connection orphaned / decommissioned

A verdict flip emits network.readiness.changed, which drives UI badges and SLO probes through Lumen.

A readiness verdict is advisory at resolution time, not a reservation. Profile selection does not lock IPs or capacity. A subnet that reads green can fill between the readiness check and the actual launch — that surfaces as a provider error at launch, not a profile error. Readiness reduces the odds of a bad pick; it does not guarantee the placement.

Detecting and healing runner drift

Runners is where Orkestia acts, because it owns the runner lifecycle. Drift here shows up in three forms, each with its own reconciliation path.

1. Scale drift — desired concurrency vs. queue reality

Scaling is reconcile-driven, not webhook-driven. A per-group reconcile loop continuously converges the pool toward the group's min/max; GitHub workflow_job webhooks (handled by runner.dispatch-from-job-queued) are a best-effort nudge that triggers an immediate reconcile, not the source of truth.

{
  "workflow": "runner.dispatch-from-job-queued",
  "input": {
    "group_id": "rg_01HXZ...",
    "reason": "workflow_job.queued",
    "queued_jobs": 3,
    "current_runners": 1
  }
}

If a webhook is lost (network blip, redirect down in local dev), nothing is stranded — the webhook was only a nudge. The reconcile loop converges the group toward its desired min/max on its next pass regardless. This is self-healing by convergence: each reconcile drives toward the desired set rather than mutating a fragile counter, so correctness never depends on any single event arriving.

2. Health drift — runners that are registered but dead

Runner groups run periodic liveness probes. An unhealthy runner is deregistered from GitHub and its compute reaped automatically, so dead runners don't sit registered and starve the queue. This is fully automated — no approval — because the action is unambiguously corrective and bounded to compute Orkestia provisioned.

sequenceDiagram
  participant HC as Health check
  participant GH as GitHub
  participant Cloud as Customer cloud
  HC->>Cloud: probe runner liveness
  Cloud-->>HC: unhealthy / unreachable
  HC->>GH: deregister runner
  HC->>Cloud: reap compute
  HC->>HC: emit runner.group.unhealthy

3. State drift — orphaned runners after a partial teardown

The hardest case: a teardown DAG fails after some cloud resources are destroyed but before GitHub deregistration. GitHub is left with zombie offline runners. The reconciler sweeps on the next group health check, reconciling GitHub's view back to actual compute. Outside the production providers (AWS, Azure, Kubernetes), treat orphan reconciliation as (beta).

Runner registration is pull, not push. Orkestia hands the runner binary a short-lived registration token and the binary calls GitHub itself; Orkestia only observes the result. So "healing" a registration means re-driving the desired set and letting runners (de)register themselves — Orkestia never force-writes GitHub's runner list out of band.

What is automated vs. surfaced for approval

This is the governance boundary. The rule of thumb: Orkestia auto-heals only state it owns and only actions that are unambiguously corrective and bounded by your declared policy. Everything else is surfaced.

Drift conditionReconciliationApproval needed?
Job queued / completed → scale within min–maxAuto: runner.scaling*No — bounded by group policy
Unhealthy runner detectedAuto: deregister + reapNo
Orphaned/zombie GitHub runner after partial teardownAuto: reconciler sweep (AWS GA; others beta)No
Lost workflow_job webhookAuto: reconcile loop converges on its next passNo
Queue exceeds max runnersNot auto-scaled past maxSurfaced — raising max is a human decision (protects your cost)
Provider quota hit at scale-upScale-up DAG fails, surfacedSurfaced — raising quota is a customer cloud action
Network profile orphaned (connection/VPC gone)Detect + block new launchesSurfaced — rebind is a human action
Security-group drift / policy violationDetect + flag verdictSurfaced — Orkestia never edits your networks
Stale network inventoryVerdict → red + readiness.changedSurfaced (on-demand re-sync available)
Bounds are hard, not best-effort. Orkestia will not exceed a group's configured max even if the queue keeps growing — protecting your spend is explicit. Likewise it never mutates customer networks. When self-healing would cross either line, it stops and surfaces rather than acting. For fleets of AI agents driving these workflows, that boundary is itself governed — see Staff governance.

Observing reconciliation

Every detection and healing action is emitted as a workflow.transition payload on the platform Kafka bus — the same bus as every other Orkestia signal — so Lumen consumes them with no runner- or network-specific log stream.

SignalSurfaceMeaning
runner.group.unhealthyRunnersHealth check failed for a group → reap path
runner.group.scaled_up / scaled_downRunnersScale reconciliation completed
runner.group.archivedRunnersTeardown DAG completed
network.readiness.changedNetworkVerdict flipped (green↔yellow↔red)
network.sync.failedNetworkProvider call failed non-retryably
network.profile.orphanedNetworkTarget/connection disappeared on sync
network.sync.age_secondsNetworkContinuous staleness gauge per connection/region
network.policy.violationsNetworkContinuous count of SG rules failing baseline

You can also poll reconciliation state directly over MCP — list_stuck_workflows surfaces runs (including provisioning/scaling/teardown DAGs) that have stalled and may need a retry_workflow. See Observability with Lumen.

Concurrency: why reconciliation is safe

Provisioning, scaling, and teardown of the same runner group are serialized by a PostgreSQL advisory lock keyed on group ID — two reconcilers can never drive the same group into conflicting states. Different groups reconcile fully in parallel. All long-running provisioning, drain, and cloud-wait steps go through workflow.transition on Kafka, consumed by the workflow consumer; there is no Celery in this path. This is the same concurrency model the workflow engine applies everywhere.

Failure modes to expect

These are known, bounded failure modes — the reconciler is designed so each degrades gracefully rather than wedging a group.
  • Stale inventory used for a deploy. A profile resolves to a network deleted in-cloud between syncs → launch fails at the provider. The verdict degrades to red on the next sync; it does not retroactively stop the in-flight launch.
  • GitHub App token expired/revoked. New runners can't register; existing runners keep working until their own registration expires. Surfaced at the token-mint step.
  • Provider quota at scale-up. ECS task / EC2 instance / Cloud Run max-instances cap hit → scale-up DAG fails with a quota error; existing runners unaffected; queue backs up until quota is raised.
  • Runner image unreachable. Aggregated-registry metadata exists but the provider can't pull → scale-up fails at image pull; retryable.
  • Network profile points to a decommissioned connection. Profile marked orphaned; reads OK, new launches blocked until rebound.

Runners

The runner control plane: groups, warm pools, executions, and the lifecycle that drift reconciliation operates on.

Hybrid execution model

How AI-designed workflows compile into the deterministic compositions that drive provisioning and scaling.

Runner management

Day-2 operations: creating groups, setting scaling bounds, and reading health.

Observability with Lumen

Consume the reconciliation signals and readiness verdicts emitted on the workflow bus.

Deployment models

Where compute lives and why Orkestia reconciles rather than hosts.

Workflow types registry

Per-workflow detail for runner.* and network.* — also at reference.orkestia.dev.