## What & why
`verify-tracing` flaked on `verify-stack` run 722 — `FAIL — no single trace spanned ['bff', 'projection-api']` — and went green on a plain re-run of the same commit. **The trace chain was not broken; Tempo could not ingest:**
```
removing distributor_pool failing healthcheck addr=127.0.0.1:9095
reason="rpc error: code = DeadlineExceeded"
pusher failed to consume trace data err="context canceled" (x18)
```
The root cause is the *mechanism* of the data loss, not whatever caused the stall. Tempo runs **single-binary**, so the distributor and the ingester are the same process and the distributor's ingester pool holds exactly one, in-process, member. dskit nevertheless health-checks that member over loopback gRPC with a **1 s** deadline (`checkinterval: 15s`, confirmed from the running image's `/status/config`). On the shared runner a transient stall blows the deadline, the only ingester is evicted from the pool, and every subsequent push fails until the next check interval — spans silently dropped.
With one in-process ingester the health check can **never** route around a failure. Its only possible effect is to discard data. So it is off:
```yaml
ingester_client:
pool_config:
healthcheckenabled: false
```
This lands at the point where *both* candidate triggers named in #156 (GC pressure near `mem_limit`, CPU contention from the grown stack) turn into lost spans, so **`mem_limit: 400m` is untouched** — raising it on a memory-tight runner risks reintroducing the `verify-e2e` OOM of #144. It also does not paper over anything the way a longer `TRACING_TIMEOUT` would (#156's own note).
Second change: `infra/tracing-check.py` prints `tempo_distributor_ingester_clients` on its failure path. From the check's side, Tempo-dropped-spans and missing instrumentation look identical — that ambiguity is what cost a container-log dive on run 722. A recurrence now names itself.
Closes #156
## Definition of Done
- [x] Linked Gitea issue (#156).
- [ ] **Failing test committed before the implementation — N/A, and deliberately so.** The trigger is runner load, so no deterministic red exists; the "red" is run 722's observed `verify-tracing` failure plus its Tempo logs. Same precedent as d5e5fa2 (#115, Playwright OOM) and 4aafd32 (#147, uWSGI caps). A test asserting the config says what the config says would add no gate: Tempo hard-fails on an unknown key (verified — `field health_check_enabled not found in type client.PoolConfig`), so a typo or a config rename on a Tempo bump already turns `verify-up` red.
- [x] Conventional Commits referencing the issue (`refs #156`).
- [ ] CI green — the point of the change.
- [x] `docker compose up` health unaffected (Tempo is not in `WAIT_SVCS`; config-only change, same image).
- [x] Docs updated — ADR-0023 Consequences.
- [x] ADR — amended **ADR-0023** rather than adding a new one: this is a consequence of that ADR's single-binary Tempo choice, not a new decision (one decision per ADR, §12).
- [x] Demo note — N/A, not user-visible.
## Notes for reviewers
Verified locally against the built image (the flake itself is not locally reproducible — see the runner-load point above):
1. `docker run --rm register-referentie/tempo:dev -config.file=/etc/tempo.yaml -config.verify=true` → parses.
2. `GET /status/config` on the running container → `healthcheckenabled: false` (was `true`).
3. The new diagnostic reads `tempo_distributor_ingester_clients` off a live Tempo.
Worth knowing: that metric is legitimately `0` on an idle Tempo — the pool is populated lazily on first push. It only prints on the failure path of a check that has already generated traffic, so the reading is meaningful there, but don't read a bare `0` on a quiet stack as an eviction.
Follow-up left undone: if `verify-tracing` still flakes after this, the next suspect is the .NET OTLP exporter timeout (#156's last note), not Tempo's memory cap.Reviewed-on: #157
3.9 KiB
ADR-0023: Grafana-native observability stack (Tempo + Prometheus + Grafana)
- Status: Accepted
- Date: 2026-07-23
- Deciders: Respellion engineering
- Slice: S-16a (#122), first of the S-16 (#17) split
Context
The PRD calls for "OpenTelemetry traces, Prometheus metrics; a local Grafana with pre-built dashboards" (§80). S-16 was split (CLAUDE.md §13) into a backplane slice (this one), distributed tracing (#123), and metrics + dashboards (#124). The backplane must stand up first: a local, CI-friendly place for traces and metrics to land, viewable in one UI, reaching green health within the 3-minute compose budget.
Two shape decisions are non-obvious enough to record.
Decision
Run a Grafana-native stack — Grafana Tempo (traces) + Prometheus (metrics) + Grafana (UI) — with the services exporting OTLP straight to Tempo (no collector), and ship the config baked into small built images.
Trace backend: Tempo (not Jaeger)
Tempo keeps everything under one Grafana pane alongside metrics (and later logs), which is exactly the "local Grafana with dashboards" the PRD asks for. Jaeger would add a second UI and a second mental model for no benefit at this scale.
No OTLP collector
Tempo ingests OTLP directly (gRPC 4317 / HTTP 4318) and Prometheus scrapes each
service's /metrics, so a collector would be a hop that processes nothing. Skipped.
If we later need fan-out, tail sampling, or log processing, a collector is an
additive change — the services already speak OTLP.
Config baked into built images, not config volumes
The upstream Common Ground modules (OpenZaak, NRC, Keycloak, Flowable) run as
verbatim images and get their config streamed into external named volumes by
infra/seed-config.sh, because bind mounts don't reach sibling containers on the
CI runner (see docs/runbooks/gitea-actions-gotchas.md). The observability tools
are not peer modules we must run verbatim, so we take the simpler path: a
three-line Dockerfile per tool that COPYs its config in. This reaches sibling
containers everywhere (docker, podman, CI) with no seed step, no CFG_VOLS entry,
and no Makefile sprawl.
Verified, not assumed
infra/run-observability-check.sh (the verify-observability step, run early in CI
verify-stack) asks Grafana to reach both datasources — Prometheus via its health
method, Tempo via the datasource proxy (Tempo's Grafana plugin implements no health
method) — so the check proves the datasources are actually wired, not merely that
containers started. The containers are not in WAIT_SVCS; the check polls Grafana
itself, so no in-image healthcheck tool is required.
Consequences
Positive
- One UI for traces + metrics + (future) logs. Config is versioned in
infra/observability/and self-contained in the images. - Backplane is independent of app instrumentation — #123 and #124 build on it.
Negative / costs
- Three more images built each CI run (kept small; not on the health-gate list).
- Storage is ephemeral container fs — a demo backplane, not a retention target. Object storage for Tempo / remote-write for Prometheus is a later concern.
- Tempo runs single-binary, so its distributor and ingester are one process and
some of its distributed-mode machinery is not just redundant but harmful. Its
ingester-pool health check is disabled (
ingester_client.pool_config) because with a single in-process ingester the check can never route around a failure — a 1s loopback-gRPC deadline missed under CI load only evicted the one ingester and made Tempo drop spans, which is howverify-tracingflaked (#156). Expect the same shape from other distributed-mode knobs if we tune them; the fix is to switch to real multi-ingester Tempo, not to re-enable them here.
Coupling rules touched (CLAUDE.md §8)
None. The stack is passive infrastructure: services push OTLP and expose
/metrics; nothing in the stack calls into a service or a peer module.