Files
register-referentie/docs/architecture/adr-0023-observability-stack.md
T
not 43c6c8e0d2
CI / lint (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m9s
CI / unit (pull_request) Successful in 1m17s
CI / frontend (pull_request) Successful in 2m43s
CI / mutation (pull_request) Successful in 5m56s
CI / verify-stack (pull_request) Failing after 8m6s
docs: ADR-0023 + demo note + backlog sync for the observability backplane (refs #122)
ADR-0023 records the Grafana-native stack (Tempo, no collector, config baked into
images); demo-script and BACKLOG describe S-16a and the S-16 split.

refs #122
2026-07-23 12:48:02 +02:00

3.3 KiB

ADR-0023: Grafana-native observability stack (Tempo + Prometheus + Grafana)

  • Status: Accepted
  • Date: 2026-07-23
  • Deciders: Respellion engineering
  • Slice: S-16a (#122), first of the S-16 (#17) split

Context

The PRD calls for "OpenTelemetry traces, Prometheus metrics; a local Grafana with pre-built dashboards" (§80). S-16 was split (CLAUDE.md §13) into a backplane slice (this one), distributed tracing (#123), and metrics + dashboards (#124). The backplane must stand up first: a local, CI-friendly place for traces and metrics to land, viewable in one UI, reaching green health within the 3-minute compose budget.

Two shape decisions are non-obvious enough to record.

Decision

Run a Grafana-native stack — Grafana Tempo (traces) + Prometheus (metrics) + Grafana (UI) — with the services exporting OTLP straight to Tempo (no collector), and ship the config baked into small built images.

Trace backend: Tempo (not Jaeger)

Tempo keeps everything under one Grafana pane alongside metrics (and later logs), which is exactly the "local Grafana with dashboards" the PRD asks for. Jaeger would add a second UI and a second mental model for no benefit at this scale.

No OTLP collector

Tempo ingests OTLP directly (gRPC 4317 / HTTP 4318) and Prometheus scrapes each service's /metrics, so a collector would be a hop that processes nothing. Skipped. If we later need fan-out, tail sampling, or log processing, a collector is an additive change — the services already speak OTLP.

Config baked into built images, not config volumes

The upstream Common Ground modules (OpenZaak, NRC, Keycloak, Flowable) run as verbatim images and get their config streamed into external named volumes by infra/seed-config.sh, because bind mounts don't reach sibling containers on the CI runner (see docs/runbooks/gitea-actions-gotchas.md). The observability tools are not peer modules we must run verbatim, so we take the simpler path: a three-line Dockerfile per tool that COPYs its config in. This reaches sibling containers everywhere (docker, podman, CI) with no seed step, no CFG_VOLS entry, and no Makefile sprawl.

Verified, not assumed

infra/run-observability-check.sh (the verify-observability step, run early in CI verify-stack) asks Grafana to reach both datasources — Prometheus via its health method, Tempo via the datasource proxy (Tempo's Grafana plugin implements no health method) — so the check proves the datasources are actually wired, not merely that containers started. The containers are not in WAIT_SVCS; the check polls Grafana itself, so no in-image healthcheck tool is required.

Consequences

Positive

  • One UI for traces + metrics + (future) logs. Config is versioned in infra/observability/ and self-contained in the images.
  • Backplane is independent of app instrumentation — #123 and #124 build on it.

Negative / costs

  • Three more images built each CI run (kept small; not on the health-gate list).
  • Storage is ephemeral container fs — a demo backplane, not a retention target. Object storage for Tempo / remote-write for Prometheus is a later concern.

Coupling rules touched (CLAUDE.md §8)

None. The stack is passive infrastructure: services push OTLP and expose /metrics; nothing in the stack calls into a service or a peer module.