Governance & Approvals
Autonomy is the point of an AI workforce — and the risk. An agent that can discover capabilities over MCP, read connections, and start workflows in your cloud has real reach. Governance is the layer that turns that reach into something you can trust in production: every action is constrained by role-based authority, sensitive actions stop at a human-in-the-loop gate, and every authorize/deny/run is recorded in an evidence trail you can query and export.
This page is the deep dive on how those controls actually work. For the conceptual overview, start with Staff & AI Workforce Governance; to operate the controls day-to-day, see Staff & Agents.
RbacGuard) and the engine transition log are live; the audit.* query library is active; the approval-gate UX and the security.* assessment library are still maturing. Items below tagged (beta) or (roadmap) are expected to change.The governance invariant: one enforcement point
Most authorization bugs come from drift — the same role check re-implemented across REST routes, MCP tools, async consumers, and SDKs, where one of them inevitably forgets a guard. Orkestia eliminates that class of bug by design: every staff-side operation is a workflow, and authority is enforced at exactly one place — inside the workflow engine.
- Staff console
- AI agent / MCP
- REST / SDK
- Kafka consumer
- engine.start_workflow
- Auth middleware
- RbacGuard middleware
- Workflow first transition
- Stop + audit_event
- Transition log + audit_event
- Staff console→engine.start_workflow
- AI agent / MCP→engine.start_workflow
- REST / SDK→engine.start_workflow
- Kafka consumer→engine.start_workflow
- engine.start_workflow→Auth middleware
- Auth middleware→RbacGuard middleware
- RbacGuard middleware→ authorize →Workflow first transition
- RbacGuard middleware→ deny →Stop + audit_event
- RbacGuard middleware→ emits →Transition log + audit_event
- Workflow first transition→ emits →Transition log + audit_event
Because the Staff console, an MCP agent, a REST/SDK caller, and the async consumer all invoke work through engine.start_workflow() / engine.transition(), they all pass through the same RbacGuard. There are no @requires_role decorators sprinkled across route handlers to forget. Subverting authority requires editing the private workflow-engine core — not a single endpoint.
Role-based authority
Authority is modeled as a real organization. The core entities — OrgUnit, Actor, RoleBinding, Capability — are first-class platform models, managed exclusively through staff.* workflows.
| Entity | What it is |
|---|---|
| OrgUnit | A team/department node. Units form a tree and are the natural scope for permissions and oversight. |
| Actor | A principal — a human operator or an AI worker — with a lifecycle (invite/hire, suspend, resume, remove) and its own inbox/outbox/journal. |
| RoleBinding | Binds a role to an actor at a unit. Effective authority is resolved by walking the unit tree. |
| Capability | The unit of authority. Each workflow class declares Capability metadata; RBAC checks the caller's effective role against it. |
How a decision is made
Each staff.* workflow declares its authority requirement in its definition — for example, the staff.grant-role-binding workflow carries metadata of this shape:
{
"workflow_type": "staff.grant-role-binding",
"capability": {
"min_role": "ADMIN",
"scope": "unit"
}
}
At run time, RbacGuard runs after auth and before the workflow's first transition. It:
- Reads
workflow_cls.capability. - Resolves the caller's effective role on the targeted unit via
resolve_effective_role(actor, unit), which walks theOrgUnittree (a binding higher in the tree inherits downward). - Compares the effective role against
Capability.min_role. - Emits an
audit_eventrow for both authorize and deny outcomes, then either proceeds or stops.
sequenceDiagram
participant C as Caller (UI / agent / SDK)
participant E as Workflow engine
participant G as RbacGuard
participant D as OrgUnit tree
participant L as audit_event
C->>E: start_workflow("staff.grant-role-binding", {unit_id, ...})
E->>G: before first transition
G->>D: resolve_effective_role(actor, unit)
D-->>G: OrgRole.ADMIN
G->>G: ADMIN >= capability.min_role? yes
G->>L: write audit_event (outcome=authorize)
G-->>E: proceed
Note over C,L: a denied call writes audit_event (outcome=deny) and stops
The roles matrix
Five roles compose additively; an effective role on a unit grants the union of capabilities at that unit's scope and inherits down the tree.
| Surface | OWNER | ADMIN | OPERATOR | VIEWER | AUDITOR |
|---|---|---|---|---|---|
| Org-wide settings (billing, SSO, plan) | RW | R | — | — | R |
| Unit tree (create / move / delete) | RW | RW | — | R | R |
| Actor lifecycle (invite / suspend / remove) | RW | RW | — | R | R |
| Role bindings (grant / revoke) | RW | RW | — | — | R |
Staff workflow execution (run staff.* on bound unit) | RW | RW | RW | R | R |
| Audit log read | R | R | R | — | RW |
| Audit log export | RW | RW | — | — | RW |
Scope-bound delegation is the headline capability this unlocks: grant a contractor OPERATOR on a single unit for a fixed window without handing over the rest of the org. See the operator procedure in Staff & Agents and the per-workflow authority requirements in the Workflow Types Registry.
Human-in-the-loop approval gates
RBAC answers "is this actor allowed to attempt this?" Approval gates answer the next question — "should this specific attempt proceed right now?" — by interposing a human decision before a sensitive action takes effect.
Because every action is a workflow and the engine is event-sourced, a gate is a natural workflow state: the run reaches a pending-approval transition and parks there until an authorized human approves or declines. Nothing downstream executes while the run waits, and the wait itself is durable — it survives pod restarts and HPA scale events because the engine persists state, not in-memory sessions.
- Start
- Proposed
- Pending Approval
- Approved
- Declined
- Executing
- Completed
- Rejected
- End
- Start→ agent proposes action →Proposed
- Proposed→ sensitive capability gated →Pending Approval
- Pending Approval→ human approves →Approved
- Pending Approval→ human declines →Declined
- Approved→Executing
- Executing→Completed
- Declined→Rejected
- Completed→End
- Rejected→End
When an AI agent proposes an action that is declined at the gate, that outcome is first-class: the agent invocation is sealed with status='rejected' (distinct from a runtime failed), so "the agent wanted to do X and a human said no" is a queryable fact, not a log line. See agent actuation for how proposals flow from the agent runtime into gated workflows.
rejected outcome — is live. Treat the exact gate-configuration surface as subject to change and confirm per-workflow gating in the registry.Where gates fit relative to guardrails
Think of the controls as a defense-in-depth stack — a request must clear every layer:
1. Tenant isolation
Org scoping is resolved server-side from the caller's token; an actor can never see or act on another org's state.
2. RBAC authority
RbacGuard denies any capability the actor's effective role doesn't grant — before the first transition runs.
Guardrails on autonomous AI actions
The same machinery that governs human operators constrains AI agents — but agents get extra guardrails because they act at machine speed and scale.
- Capability-scoped reach. An agent acts as an
Actorwith role bindings. A "finance agent" bound only to the Finance unit withfinance-operatorcannot touch infrastructure workflows;RbacGuarddenies the attempt and records it. This is how you keep authority from sprawling as you add agents. - Discovery is gated too. Agents discover capabilities over MCP, but discovery and invocation both resolve org and authority server-side from the token — an agent cannot widen its own scope by passing a different
organization_uuid. - Hashed invocation audit. Every agent run writes one
agent_invocationrow (running→success/failed/rejected). Tool arguments, results, and prompts are stored as SHA-256 hashes of canonical JSON, never literals — so the trail is safe to expose to operators and auditors while still supporting correlation ("all runs that calledstart_workflowwith the same arg shape"). The agent's own final message and error reasons are kept verbatim because those are the artifacts humans review. - Crash detection. A run that vanishes mid-flight (pod killed, OOM) leaves its row stuck in
running— a code-detectable signal that an agent crashed without a clean exit, surfaced rather than silently dropped. - Deterministic compilation. When AI (DGI) designs a workflow, the result compiles into a deterministic virtual workflow. Review and govern the compiled composition, not an opaque model decision at execution time — see the hybrid execution model.
Capability lives on the class, and an actor's reach is the union of its RoleBinding rows. You can answer "what can this agent do?" by reading data, not by reasoning about prompts.The audit & evidence trail
Accountability rests on the engine's event-sourced transition log: every workflow run is an append-only sequence of transitions. Nothing mutates that history; the governance value is that "what ran, when, on whose behalf, and what came of it" is already recorded as a side effect of how the engine works.
The Audit library (audit.*) exposes that log as a typed, read-only, org-scoped query surface — so callers ask questions without writing raw SQL against engine internals or risking a mutation.
| Workflow | Kind | What it answers |
|---|---|---|
audit.workflow-run.query | data (read-only) | Paginated runs for the org, filterable by type prefix, state, terminal status, actor, time range |
audit.workflow-run.get-history | data (read-only) | Full transition log for one run, after verifying it belongs to the caller's org |
audit.workflow-run.aggregate | data (read-only) | Per-type counts and last_started_at over a time range |
audit.workflow-health.scan | workflow | Surfaces unhealthy / stuck runs |
An evidence pack is composed from these read-only queries — query for the run list, get-history per run, aggregate for totals — over the time range and prefixes under review.
Grouping is prefix-composed: callers pass their own workflow_type_prefixes (e.g. ['kubernetes.', 'deploy.k8s.', 'runner.']) so "all Kubernetes activity" works without hard-coding filters into the engine. Confirm exact inputs and outputs against the external catalog.
{
"workflow_type": "audit.workflow-run.query",
"initial_data": {
"workflow_type_prefixes": ["staff.", "kubernetes."],
"started_after": "2026-06-01T00:00:00Z",
"started_before": "2026-06-30T23:59:59Z"
}
}
audit.*) is the system of record for workflow runs and RbacGuard authorize/deny events. The agent_invocation table is the per-run record for AI agent executions (hashed args/results, rejected outcomes). Together they answer both "what did the platform do for this org?" and "what did each agent attempt?".Tie to compliance
The governance model maps cleanly onto common compliance controls:
| Control objective | How Orkestia satisfies it |
|---|---|
| Least privilege | Role bindings scoped to units; RbacGuard denies anything not explicitly granted |
| Separation of duties | Distinct roles (e.g. OPERATOR runs, AUDITOR reviews); approval gates require a different human to approve |
| Change approval | Human-in-the-loop gates on sensitive workflows; declines recorded as rejected |
| Audit logging | Append-only transition log; authorize and deny both written to audit_event |
| Evidence on demand | Evidence packs composed from the read-only audit.* queries — time-range / prefix-scoped artifacts for reviewers |
| Data minimization | Agent audit stores hashes, not literals; Zero Code Custody means customer code/data never leaves your cloud |
| Tenant isolation | Org scope resolved server-side from the token; every audit query re-verifies org ownership |
security.* library now carries org-level workflow policy controls (security.org-workflow-policy.*) — org-wide rules over which workflow surfaces may run — and the policy layer adds feature flags and entitlements (policy.*) enforced at the invoke gate. The broader security-assessment surface (assessment campaigns, posture collection, findings triage) remains roadmap; see the security and compliance guide.Orkestia provides the mechanisms (enforcement, gates, immutable trail) that make a control environment auditable; it does not by itself constitute a certification. Map these controls to your own framework with your auditor, and see deployment models for how the posture differs across hosting topologies.
Where to go next
Drift Detection & Self-Healing
How Orkestia reconciles desired and actual state for runners and network infrastructure in your cloud, computes readiness verdicts, and decides what to heal automatically versus surface for approval
Cost & Performance Optimization
How to keep Orkestia workflows fast and cheap by compiling AI-designed flows to deterministic compositions, right-sizing runners, scaling async over Kafka, and using Lumen to find the expensive steps
