Lumen Observability
TL;DR
- Lumen is a separate host:
https://lumen-api.orkestia.dev. The workflow API keeps run state; Lumen keeps the telemetry you send plus what grouping derives. - Off until provisioned. Every write is
403 LUMEN_NOT_PROVISIONEDuntil an org admin enables a plan. - JSON over HTTP, not OTLP. Logs, metrics, and Pulse product events. Optional Kafka topic for logs only.
- Error groups are the unit of work, keyed by a SHA-256 fingerprint. Alerts fire on new, regressed, or threshold groups.
- Assistants triage through the Lumen MCP at
https://mcp-lumen.orkestia.dev/mcp, a second server next to the workflow MCP.
Why it exists
Orkestia's default posture is Zero Code Custody: runners execute in your cloud; the control plane retains workflow state plus orchestration metadata. Observability is the scoped exception. To answer "why did this fail, and has it failed before?" you need the line, the similar group, the span tree, and the last resolution, together.
Telemetry is not App Data and not source. A secret in message, traceback, or context is stored. See Security & compliance.
Pipeline
producer --POST /api/logs|metrics|product/ingest[/batch]--> lumen-api.orkestia.dev
| optional: Kafka topic lumen.logs (same JSON as HTTP logs)
v
normalize / redact / ingest rules (drop | sample)
v
persist → 201 { id, received_at } | 200 { dropped: true } | 429 quota
v
SHA-256 fingerprint (sync) + queue
v
processor: error groups · alert rules · lumen.alert.fired (ticket)
v
GET /api/logs · /error-groups · /traces/{id} · /metrics · Lumen MCP
| Signal | Write | Read |
|---|---|---|
| Logs | POST /api/logs/ingest[/batch] | GET /api/logs (default last 24h, limit ≤ 500), /logs/stats, /logs/pulse |
| Error groups | Derived, not a write API | GET /api/error-groups; PATCH …/resolve, ignore, reopen, assign, severity |
| Traces / spans | UUID trace_id and span_id on logs, or POST /api/traces (org session) | GET /api/traces/{id}, /spans/stats, /slow, /errors |
| Metrics | POST /api/metrics/ingest[/batch] | GET /api/metrics, /aggregate, /names |
| Pulse | POST /api/product/ingest[/batch] with lump_ | GET /api/product/events, /stats |
Batch body is { "items": [ … ] } (Pulse also accepts "events"). Max 1000 items; browser Pulse keys cap at 50. Kafka lumen.logs is logs only. channel=ticket alerts publish lumen.alert.fired.
Keys
| Prefix | Scope | Header | Allowed |
|---|---|---|---|
lumk_… | ingest | X-Api-Key | Log and metric writes |
lumk_… | read | X-Api-Key | GET and HEAD only |
lump_… | product | X-Api-Key or Authorization: Bearer | Pulse only |
The key binds the org. Do not send organization_uuid with customer keys. Never put a lumk_ ingest key in a browser.
Ingest status codes: 201 stored · 200 dropped · 401 bad or missing key · 403 not provisioned · 422 schema · 429 rate or monthly quota (Retry-After, X-RateLimit-*).
Fingerprint and error groups
A group is the unit of work, not a line. On write:
- Ingest rules (
drop,sample). Dropped lines never group. - A 64-char lowercase hex SHA-256 is computed synchronously.
- A background processor attaches the log to an error group and evaluates alert rules.
Levels that group by default: ERROR, WARNING, CRITICAL. Other levels are stored, not grouped.
Hash order: matching fingerprint rule → sha256(fingerprint_key); else error_type|message_norm or location|message_norm; else weak project|level|message_norm. Dynamic tokens in message are collapsed before hashing.
The line is searchable immediately. The group and any alert appear after the processor runs. Full field list: Send data.
Structural attribution
| Field | Effect |
|---|---|
workflow_id, workflow_type, workflow_state_id, actor_id | Filter a run without reconstructing context |
trace_id, span_id, parent_span_id | UUID required for a real span tree; anything else becomes external_trace_id |
project, environment, release | Scope and compare deploys. The first stored project string creates the project |
exception_class, error_code, traceback | Stronger grouping |
source | Default "workflow"; collector sends kubernetes / kubernetes-event; Python SDK sends python-sdk |
Rules
Applied along the pipeline: ingest (before persist) → fingerprint (hash override) → alert (after group write).
| Type | Actions | Permanent? |
|---|---|---|
| Ingest | drop, sample | Dropped data is gone |
| Fingerprint | regex → sha256(fingerprint_key) | Reshapes grouping going forward |
| Alert | new_group, regressed, threshold | Channels: webhook, slack, email, ticket |
JSON shapes: Query API, Rules.
Surfaces
| Surface | URL | Auth |
|---|---|---|
| Provision and plan | app.orkestia.dev/governance/lumen | Org admin |
| App | lumen.orkestia.dev | Same session; routes map 1:1 to the Query API |
| API | https://lumen-api.orkestia.dev | lumk_, lump_, or org Bearer. Query windows capped at 90 days |
| MCP | https://mcp-lumen.orkestia.dev/mcp | Bearer forwarded. Not mcp.orkestia.dev |
| Health | GET /health | none → {"status":"healthy"} |
The Lumen MCP
The Lumen MCP is a second server you connect next to the workflow MCP. Its tools are the same reads and triage mutations the app exposes: list_logs, get_trace, error-group listing and resolve_error_group, metrics queries, and more. A typical triage turn:
list_error_groups(since="24h", status="open") → ranked groups
get_trace(trace_id) → span tree for the top one
list_logs(fingerprint=…, limit=50) → the raw lines
resolve_error_group(group_id, note="fixed in v1.2") → triage mutation
Inventory and setup: Lumen MCP. Connect it the same way as the workflow server, described in Connect an AI assistant.
Privacy
| Stored | Not stored |
|---|---|
| Lines, fingerprints, spans, series, triage notes you sent or grouping derived | Source, App Data rows, long-lived cloud credentials |
Scrubbing reduces incidental PII; it does not remove a secret you logged. Pair structured fields with ingest drop rules.
SLA Health (preview)
SLA Health is a separate preview: policies over uptime, latency, and error rate, not a Lumen query. Today it is demo-driven. Production health is Lumen pulse, slow spans, and error-group counts.
Ask your AI assistant
List the open error groups from the last 24 hours, rank them by occurrence, and show the trace of the top one.
Find logs for workflow <workflow_id> and tell me the first ERROR line and its fingerprint.
Which projects had a regressed error group since the last release?
Resolve error group <group_id> with the note "fixed by PR #42". Confirm before you do it.
For AI agents
| Rule | Detail |
|---|---|
| Two servers | Workflow operations go to mcp.orkestia.dev. Telemetry goes to mcp-lumen.orkestia.dev. Do not look for log tools on the workflow server. |
| Provisioned or not | A 403 LUMEN_NOT_PROVISIONED means the org has not enabled Lumen. Say so; do not retry. |
| Groups, not lines | Reason about error groups. Use the fingerprint to pull the lines behind one. |
| Triage writes are mutations | resolve, ignore, assign, and severity changes should be confirmed with the user. |
| Correlate with runs | workflow_id on a log line links straight back to get_workflow_history on the workflow MCP. |
Where to go next
Runners & Execution Environments
Managed execution environments provisioned inside your own cloud where workflow steps, CI jobs, and agent sessions run, runner groups, warm pools, reconcile-loop scaling, and the control-plane / execution split
Identity & Multi-Tenancy
Two identity planes, org members who operate the platform and end-users who sign in to apps you build, with automatic org scoping, per-user isolation, and a one-call setup an assistant can run for you
