Drift Detection & Self-Healing
Orkestia provisions runners and reads network topology inside your cloud account — but the cloud is a living system. Quotas change, consoles get edited by hand, runners die, webhooks get dropped, and a VPC referenced by a profile gets decommissioned. Drift is the gap between the state Orkestia intends and the state the provider actually reports. This page explains how that gap is detected, how readiness verdicts are computed from it, and the deliberate line between what Orkestia heals on its own and what it surfaces for a human to approve.
What "drift" means on Orkestia
There are two distinct surfaces where desired and actual state can diverge, and they reconcile differently:
| Surface | Desired state | Actual state | Reconciliation style |
|---|---|---|---|
| Runners | Group config: min/max, labels, network placement, registered runner set | Live cloud compute (ECS tasks / EC2 / K8s pods) + GitHub's view of registered runners | Active — reconcile-loop scaling, health-reap, teardown repair |
| Network Management | Saved network profiles (VPC + subnets + SGs, versioned) | Provider network inventory mirrored by periodic sync | Detect-and-surface — Orkestia never mutates customer networks |
The asymmetry is intentional. Orkestia owns the lifecycle of the runners it provisions, so it can act on runner drift directly. It does not own your networks — Network Management is pull-only inventory, so network drift is detected and flagged, never silently corrected.
How desired state is expressed
Both surfaces declare desired state as versioned, org-local metadata — never as mutable in-place rows. This is what makes drift computable: there is always a stable reference to diff actual state against.
- A runner group holds
(GitHub org + target cloud env + scaling policy + network placement + labels). It moves throughdraft → provisioning → active → scaling → archived. - A network profile holds
(VPC, subnets, security groups, optional public/private intent)and is versioned: every edit bumpsvN → vN+1, and an in-flight launch keeps the version it resolved. A profile edit can never retroactively break a running deploy.
{
"id": "np_01HXZ...",
"name": "prod-usa-private",
"version": 3,
"intent": "private",
"targets": [
{
"provider": "aws",
"connection_id": "cn_aws_prod",
"region": "us-east-1",
"vpc_id": "vpc-0abc123",
"subnets": ["subnet-0a...", "subnet-0b..."],
"security_groups": ["sg-0ff..."]
}
],
"readiness": "green",
"updated_at": "2026-04-15T18:22:04Z"
}
Detecting network drift (pull, detect, surface)
Network Management mirrors cloud topology locally so apps don't re-query each provider on every deploy. Drift is the gap between that mirror and reality, and it is detected by re-running the sync.
- Scheduled sync per connection
- Provider call
- Update inventory + discovered_at
- Mark slice stalekeep last discovered_at
- Recompute readiness verdict
- Verdict changed?
- Emit network.readiness.changed
- No-op
- Scheduled sync per connection→Provider call
- Provider call→ success →Update inventory + discovered_at
- Provider call→ partial fail →Mark slice stale
- Update inventory + discovered_at→Recompute readiness verdict
- Mark slice stale→Recompute readiness verdict
- Recompute readiness verdict→Verdict changed?
- Verdict changed?→ yes →Emit network.readiness.changed
- Verdict changed?→ no →No-op
Key properties of network drift detection:
- Pull-only, eventually consistent. A network created directly in a cloud console appears only after the next sync — not instantly. Downstream products must tolerate "not yet known."
- Per-region freshness, not all-or-nothing. A provider error in one region/account does not invalidate prior data elsewhere; the failed slice is marked stale with its last successful
discovered_at, surfaced as thenetwork.sync.age_secondsgauge. - Silent SG drift is the hard case. Security-group rules edited in the console are only picked up on the next sync. Between syncs the policy-violation view can under-report; on-demand re-sync (UI or MCP) is the mitigation.
- No mutation, ever. Orkestia is not an IPAM. It does not allocate CIDRs, reserve IPs, or touch route tables. Detected network drift produces a verdict and a notification, not a corrective write.
When a profile's underlying connection or network vanishes, the profile is marked orphaned: existing resolutions remain readable, but new launches are blocked until it is rebound — a deliberate fail-closed choice rather than letting a deploy land somewhere unknown.
Readiness verdicts
Every network and profile carries a deploy-ready verdict — green / yellow / red — so users see "don't pick this" before they pick it. The verdict is a rollup of several signals:
| Input signal | Pushes toward red |
|---|---|
Sync freshness (network.sync.age_seconds) | Inventory stale beyond threshold |
| Quota headroom | Subnet effectively full / no IP space |
| Required tags | Mandatory governance tags missing |
| Security-group sanity | Rules violating baseline (e.g. 0.0.0.0/0 on sensitive ports) |
| Connection health | Source connection orphaned / decommissioned |
A verdict flip emits network.readiness.changed, which drives UI badges and SLO probes through Lumen.
Detecting and healing runner drift
Runners is where Orkestia acts, because it owns the runner lifecycle. Drift here shows up in three forms, each with its own reconciliation path.
1. Scale drift — desired concurrency vs. queue reality
Scaling is reconcile-driven, not webhook-driven. A per-group reconcile loop continuously converges the pool toward the group's min/max; GitHub workflow_job webhooks (handled by runner.dispatch-from-job-queued) are a best-effort nudge that triggers an immediate reconcile, not the source of truth.
{
"workflow": "runner.dispatch-from-job-queued",
"input": {
"group_id": "rg_01HXZ...",
"reason": "workflow_job.queued",
"queued_jobs": 3,
"current_runners": 1
}
}
If a webhook is lost (network blip, redirect down in local dev), nothing is stranded — the webhook was only a nudge. The reconcile loop converges the group toward its desired min/max on its next pass regardless. This is self-healing by convergence: each reconcile drives toward the desired set rather than mutating a fragile counter, so correctness never depends on any single event arriving.
2. Health drift — runners that are registered but dead
Runner groups run periodic liveness probes. An unhealthy runner is deregistered from GitHub and its compute reaped automatically, so dead runners don't sit registered and starve the queue. This is fully automated — no approval — because the action is unambiguously corrective and bounded to compute Orkestia provisioned.
sequenceDiagram participant HC as Health check participant GH as GitHub participant Cloud as Customer cloud HC->>Cloud: probe runner liveness Cloud-->>HC: unhealthy / unreachable HC->>GH: deregister runner HC->>Cloud: reap compute HC->>HC: emit runner.group.unhealthy
3. State drift — orphaned runners after a partial teardown
The hardest case: a teardown DAG fails after some cloud resources are destroyed but before GitHub deregistration. GitHub is left with zombie offline runners. The reconciler sweeps on the next group health check, reconciling GitHub's view back to actual compute. Outside the production providers (AWS, Azure, Kubernetes), treat orphan reconciliation as (beta).
What is automated vs. surfaced for approval
This is the governance boundary. The rule of thumb: Orkestia auto-heals only state it owns and only actions that are unambiguously corrective and bounded by your declared policy. Everything else is surfaced.
| Drift condition | Reconciliation | Approval needed? |
|---|---|---|
Job queued / completed → scale within min–max | Auto: runner.scaling* | No — bounded by group policy |
| Unhealthy runner detected | Auto: deregister + reap | No |
| Orphaned/zombie GitHub runner after partial teardown | Auto: reconciler sweep (AWS GA; others beta) | No |
Lost workflow_job webhook | Auto: reconcile loop converges on its next pass | No |
Queue exceeds max runners | Not auto-scaled past max | Surfaced — raising max is a human decision (protects your cost) |
| Provider quota hit at scale-up | Scale-up DAG fails, surfaced | Surfaced — raising quota is a customer cloud action |
| Network profile orphaned (connection/VPC gone) | Detect + block new launches | Surfaced — rebind is a human action |
| Security-group drift / policy violation | Detect + flag verdict | Surfaced — Orkestia never edits your networks |
| Stale network inventory | Verdict → red + readiness.changed | Surfaced (on-demand re-sync available) |
max even if the queue keeps growing — protecting your spend is explicit. Likewise it never mutates customer networks. When self-healing would cross either line, it stops and surfaces rather than acting. For fleets of AI agents driving these workflows, that boundary is itself governed — see Staff governance.Observing reconciliation
Every detection and healing action is emitted as a workflow.transition payload on the platform Kafka bus — the same bus as every other Orkestia signal — so Lumen consumes them with no runner- or network-specific log stream.
| Signal | Surface | Meaning |
|---|---|---|
runner.group.unhealthy | Runners | Health check failed for a group → reap path |
runner.group.scaled_up / scaled_down | Runners | Scale reconciliation completed |
runner.group.archived | Runners | Teardown DAG completed |
network.readiness.changed | Network | Verdict flipped (green↔yellow↔red) |
network.sync.failed | Network | Provider call failed non-retryably |
network.profile.orphaned | Network | Target/connection disappeared on sync |
network.sync.age_seconds | Network | Continuous staleness gauge per connection/region |
network.policy.violations | Network | Continuous count of SG rules failing baseline |
You can also poll reconciliation state directly over MCP — list_stuck_workflows surfaces runs (including provisioning/scaling/teardown DAGs) that have stalled and may need a retry_workflow. See Observability with Lumen.
Concurrency: why reconciliation is safe
Provisioning, scaling, and teardown of the same runner group are serialized by a PostgreSQL advisory lock keyed on group ID — two reconcilers can never drive the same group into conflicting states. Different groups reconcile fully in parallel. All long-running provisioning, drain, and cloud-wait steps go through workflow.transition on Kafka, consumed by the workflow consumer; there is no Celery in this path. This is the same concurrency model the workflow engine applies everywhere.
Failure modes to expect
- Stale inventory used for a deploy. A profile resolves to a network deleted in-cloud between syncs → launch fails at the provider. The verdict degrades to red on the next sync; it does not retroactively stop the in-flight launch.
- GitHub App token expired/revoked. New runners can't register; existing runners keep working until their own registration expires. Surfaced at the token-mint step.
- Provider quota at scale-up. ECS task / EC2 instance / Cloud Run max-instances cap hit → scale-up DAG fails with a quota error; existing runners unaffected; queue backs up until quota is raised.
- Runner image unreachable. Aggregated-registry metadata exists but the provider can't pull → scale-up fails at image pull; retryable.
- Network profile points to a decommissioned connection. Profile marked orphaned; reads OK, new launches blocked until rebound.
Related reading
Runners
The runner control plane: groups, warm pools, executions, and the lifecycle that drift reconciliation operates on.
Hybrid execution model
How AI-designed workflows compile into the deterministic compositions that drive provisioning and scaling.
Hybrid AI + Deterministic Execution
How Orkestia lets AI design workflows at runtime, then compiles those plans into deterministic virtual workflows that run repeatably, cheaply, and auditably
Governance & Approvals
How Orkestia enforces role-based authority, human-in-the-loop approval gates, and an immutable evidence trail over fleets of autonomous AI agents
