CI / lint (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m9s
CI / unit (pull_request) Successful in 1m17s
CI / frontend (pull_request) Successful in 2m43s
CI / mutation (pull_request) Successful in 5m56s
CI / verify-stack (pull_request) Failing after 8m6s
ADR-0023 records the Grafana-native stack (Tempo, no collector, config baked into images); demo-script and BACKLOG describe S-16a and the S-16 split. refs #122
75 lines
3.3 KiB
Markdown
75 lines
3.3 KiB
Markdown
# ADR-0023: Grafana-native observability stack (Tempo + Prometheus + Grafana)
|
|
|
|
- **Status:** Accepted
|
|
- **Date:** 2026-07-23
|
|
- **Deciders:** Respellion engineering
|
|
- **Slice:** S-16a (#122), first of the S-16 (#17) split
|
|
|
|
## Context
|
|
|
|
The PRD calls for "OpenTelemetry traces, Prometheus metrics; a local Grafana with
|
|
pre-built dashboards" (§80). S-16 was split (CLAUDE.md §13) into a backplane slice
|
|
(this one), distributed tracing (#123), and metrics + dashboards (#124). The
|
|
backplane must stand up first: a local, CI-friendly place for traces and metrics to
|
|
land, viewable in one UI, reaching green health within the 3-minute compose budget.
|
|
|
|
Two shape decisions are non-obvious enough to record.
|
|
|
|
## Decision
|
|
|
|
**Run a Grafana-native stack — Grafana Tempo (traces) + Prometheus (metrics) +
|
|
Grafana (UI) — with the services exporting OTLP straight to Tempo (no collector),
|
|
and ship the config baked into small built images.**
|
|
|
|
### Trace backend: Tempo (not Jaeger)
|
|
|
|
Tempo keeps everything under one Grafana pane alongside metrics (and later logs),
|
|
which is exactly the "local Grafana with dashboards" the PRD asks for. Jaeger would
|
|
add a second UI and a second mental model for no benefit at this scale.
|
|
|
|
### No OTLP collector
|
|
|
|
Tempo ingests OTLP directly (gRPC 4317 / HTTP 4318) and Prometheus scrapes each
|
|
service's `/metrics`, so a collector would be a hop that processes nothing. Skipped.
|
|
If we later need fan-out, tail sampling, or log processing, a collector is an
|
|
additive change — the services already speak OTLP.
|
|
|
|
### Config baked into built images, not config volumes
|
|
|
|
The upstream Common Ground modules (OpenZaak, NRC, Keycloak, Flowable) run as
|
|
**verbatim** images and get their config streamed into external named volumes by
|
|
`infra/seed-config.sh`, because bind mounts don't reach sibling containers on the
|
|
CI runner (see `docs/runbooks/gitea-actions-gotchas.md`). The observability tools
|
|
are **not** peer modules we must run verbatim, so we take the simpler path: a
|
|
three-line `Dockerfile` per tool that `COPY`s its config in. This reaches sibling
|
|
containers everywhere (docker, podman, CI) with no seed step, no `CFG_VOLS` entry,
|
|
and no Makefile sprawl.
|
|
|
|
### Verified, not assumed
|
|
|
|
`infra/run-observability-check.sh` (the `verify-observability` step, run early in CI
|
|
`verify-stack`) asks Grafana to reach both datasources — Prometheus via its health
|
|
method, Tempo via the datasource proxy (Tempo's Grafana plugin implements no health
|
|
method) — so the check proves the datasources are actually wired, not merely that
|
|
containers started. The containers are not in `WAIT_SVCS`; the check polls Grafana
|
|
itself, so no in-image healthcheck tool is required.
|
|
|
|
## Consequences
|
|
|
|
**Positive**
|
|
|
|
- One UI for traces + metrics + (future) logs. Config is versioned in
|
|
`infra/observability/` and self-contained in the images.
|
|
- Backplane is independent of app instrumentation — #123 and #124 build on it.
|
|
|
|
**Negative / costs**
|
|
|
|
- Three more images built each CI run (kept small; not on the health-gate list).
|
|
- Storage is ephemeral container fs — a demo backplane, not a retention target.
|
|
Object storage for Tempo / remote-write for Prometheus is a later concern.
|
|
|
|
## Coupling rules touched (CLAUDE.md §8)
|
|
|
|
None. The stack is passive infrastructure: services *push* OTLP and *expose*
|
|
`/metrics`; nothing in the stack calls into a service or a peer module.
|