Orkestia
Blog
Core Concepts

Lumen Observability

Orkestia's telemetry store and triage engine, JSON HTTP ingest, SHA-256 error groups, traces, metrics, the query API, and the Lumen MCP server for assistant-driven triage

TL;DR

  • Lumen is a separate host: https://lumen-api.orkestia.dev. The workflow API keeps run state; Lumen keeps the telemetry you send plus what grouping derives.
  • Off until provisioned. Every write is 403 LUMEN_NOT_PROVISIONED until an org admin enables a plan.
  • JSON over HTTP, not OTLP. Logs, metrics, and Pulse product events. Optional Kafka topic for logs only.
  • Error groups are the unit of work, keyed by a SHA-256 fingerprint. Alerts fire on new, regressed, or threshold groups.
  • Assistants triage through the Lumen MCP at https://mcp-lumen.orkestia.dev/mcp, a second server next to the workflow MCP.
The customer contract (provision, keys, ingest schema, query params, collector, app, MCP) lives in the Lumen section. This page is the mechanism.

Why it exists

Orkestia's default posture is Zero Code Custody: runners execute in your cloud; the control plane retains workflow state plus orchestration metadata. Observability is the scoped exception. To answer "why did this fail, and has it failed before?" you need the line, the similar group, the span tree, and the last resolution, together.

Telemetry is not App Data and not source. A secret in message, traceback, or context is stored. See Security & compliance.

Pipeline

producer  --POST /api/logs|metrics|product/ingest[/batch]-->  lumen-api.orkestia.dev
                |  optional: Kafka topic lumen.logs (same JSON as HTTP logs)
                v
         normalize / redact / ingest rules (drop | sample)
                v
         persist  →  201 { id, received_at }  |  200 { dropped: true }  |  429 quota
                v
         SHA-256 fingerprint (sync) + queue
                v
         processor: error groups · alert rules · lumen.alert.fired (ticket)
                v
         GET /api/logs · /error-groups · /traces/{id} · /metrics · Lumen MCP
SignalWriteRead
LogsPOST /api/logs/ingest[/batch]GET /api/logs (default last 24h, limit ≤ 500), /logs/stats, /logs/pulse
Error groupsDerived, not a write APIGET /api/error-groups; PATCH …/resolve, ignore, reopen, assign, severity
Traces / spansUUID trace_id and span_id on logs, or POST /api/traces (org session)GET /api/traces/{id}, /spans/stats, /slow, /errors
MetricsPOST /api/metrics/ingest[/batch]GET /api/metrics, /aggregate, /names
PulsePOST /api/product/ingest[/batch] with lump_GET /api/product/events, /stats

Batch body is { "items": [ … ] } (Pulse also accepts "events"). Max 1000 items; browser Pulse keys cap at 50. Kafka lumen.logs is logs only. channel=ticket alerts publish lumen.alert.fired.

Keys

PrefixScopeHeaderAllowed
lumk_…ingestX-Api-KeyLog and metric writes
lumk_…readX-Api-KeyGET and HEAD only
lump_…productX-Api-Key or Authorization: BearerPulse only

The key binds the org. Do not send organization_uuid with customer keys. Never put a lumk_ ingest key in a browser.

Ingest status codes: 201 stored · 200 dropped · 401 bad or missing key · 403 not provisioned · 422 schema · 429 rate or monthly quota (Retry-After, X-RateLimit-*).

Fingerprint and error groups

A group is the unit of work, not a line. On write:

  1. Ingest rules (drop, sample). Dropped lines never group.
  2. A 64-char lowercase hex SHA-256 is computed synchronously.
  3. A background processor attaches the log to an error group and evaluates alert rules.

Levels that group by default: ERROR, WARNING, CRITICAL. Other levels are stored, not grouped.

Hash order: matching fingerprint rule → sha256(fingerprint_key); else error_type|message_norm or location|message_norm; else weak project|level|message_norm. Dynamic tokens in message are collapsed before hashing.

The line is searchable immediately. The group and any alert appear after the processor runs. Full field list: Send data.

Structural attribution

FieldEffect
workflow_id, workflow_type, workflow_state_id, actor_idFilter a run without reconstructing context
trace_id, span_id, parent_span_idUUID required for a real span tree; anything else becomes external_trace_id
project, environment, releaseScope and compare deploys. The first stored project string creates the project
exception_class, error_code, tracebackStronger grouping
sourceDefault "workflow"; collector sends kubernetes / kubernetes-event; Python SDK sends python-sdk

Rules

Applied along the pipeline: ingest (before persist) → fingerprint (hash override) → alert (after group write).

TypeActionsPermanent?
Ingestdrop, sampleDropped data is gone
Fingerprintregex → sha256(fingerprint_key)Reshapes grouping going forward
Alertnew_group, regressed, thresholdChannels: webhook, slack, email, ticket

JSON shapes: Query API, Rules.

Surfaces

SurfaceURLAuth
Provision and planapp.orkestia.dev/governance/lumenOrg admin
Applumen.orkestia.devSame session; routes map 1:1 to the Query API
APIhttps://lumen-api.orkestia.devlumk_, lump_, or org Bearer. Query windows capped at 90 days
MCPhttps://mcp-lumen.orkestia.dev/mcpBearer forwarded. Not mcp.orkestia.dev
HealthGET /healthnone → {"status":"healthy"}

The Lumen MCP

The Lumen MCP is a second server you connect next to the workflow MCP. Its tools are the same reads and triage mutations the app exposes: list_logs, get_trace, error-group listing and resolve_error_group, metrics queries, and more. A typical triage turn:

assistant transcript
list_error_groups(since="24h", status="open")       → ranked groups
get_trace(trace_id)                                 → span tree for the top one
list_logs(fingerprint=…, limit=50)                  → the raw lines
resolve_error_group(group_id, note="fixed in v1.2") → triage mutation

Inventory and setup: Lumen MCP. Connect it the same way as the workflow server, described in Connect an AI assistant.

Privacy

StoredNot stored
Lines, fingerprints, spans, series, triage notes you sent or grouping derivedSource, App Data rows, long-lived cloud credentials

Scrubbing reduces incidental PII; it does not remove a secret you logged. Pair structured fields with ingest drop rules.

SLA Health (preview)

SLA Health is a separate preview: policies over uptime, latency, and error rate, not a Lumen query. Today it is demo-driven. Production health is Lumen pulse, slow spans, and error-group counts.

Ask your AI assistant

prompts, on the Lumen MCP
List the open error groups from the last 24 hours, rank them by occurrence, and show the trace of the top one.

Find logs for workflow <workflow_id> and tell me the first ERROR line and its fingerprint.

Which projects had a regressed error group since the last release?

Resolve error group <group_id> with the note "fixed by PR #42". Confirm before you do it.

For AI agents

RuleDetail
Two serversWorkflow operations go to mcp.orkestia.dev. Telemetry goes to mcp-lumen.orkestia.dev. Do not look for log tools on the workflow server.
Provisioned or notA 403 LUMEN_NOT_PROVISIONED means the org has not enabled Lumen. Say so; do not retry.
Groups, not linesReason about error groups. Use the fingerprint to pull the lines behind one.
Triage writes are mutationsresolve, ignore, assign, and severity changes should be confirmed with the user.
Correlate with runsworkflow_id on a log line links straight back to get_workflow_history on the workflow MCP.

Where to go next

Enable

Provision, plans, 403 and 429 codes, mint lumk_ and lump_.

Send data

Ingest schema, fingerprint, Kafka, Python SDK.

Query API

Filters, mutations, rule JSON.

Collector

DaemonSet → /api/logs/ingest/batch.

Lumen MCP

Tool inventory on mcp-lumen.orkestia.dev.