## What & why
`verify-tracing` flaked on `verify-stack` run 722 — `FAIL — no single trace spanned ['bff', 'projection-api']` — and went green on a plain re-run of the same commit. **The trace chain was not broken; Tempo could not ingest:**
```
removing distributor_pool failing healthcheck addr=127.0.0.1:9095
reason="rpc error: code = DeadlineExceeded"
pusher failed to consume trace data err="context canceled" (x18)
```
The root cause is the *mechanism* of the data loss, not whatever caused the stall. Tempo runs **single-binary**, so the distributor and the ingester are the same process and the distributor's ingester pool holds exactly one, in-process, member. dskit nevertheless health-checks that member over loopback gRPC with a **1 s** deadline (`checkinterval: 15s`, confirmed from the running image's `/status/config`). On the shared runner a transient stall blows the deadline, the only ingester is evicted from the pool, and every subsequent push fails until the next check interval — spans silently dropped.
With one in-process ingester the health check can **never** route around a failure. Its only possible effect is to discard data. So it is off:
```yaml
ingester_client:
pool_config:
healthcheckenabled: false
```
This lands at the point where *both* candidate triggers named in #156 (GC pressure near `mem_limit`, CPU contention from the grown stack) turn into lost spans, so **`mem_limit: 400m` is untouched** — raising it on a memory-tight runner risks reintroducing the `verify-e2e` OOM of #144. It also does not paper over anything the way a longer `TRACING_TIMEOUT` would (#156's own note).
Second change: `infra/tracing-check.py` prints `tempo_distributor_ingester_clients` on its failure path. From the check's side, Tempo-dropped-spans and missing instrumentation look identical — that ambiguity is what cost a container-log dive on run 722. A recurrence now names itself.
Closes #156
## Definition of Done
- [x] Linked Gitea issue (#156).
- [ ] **Failing test committed before the implementation — N/A, and deliberately so.** The trigger is runner load, so no deterministic red exists; the "red" is run 722's observed `verify-tracing` failure plus its Tempo logs. Same precedent as d5e5fa2 (#115, Playwright OOM) and 4aafd32 (#147, uWSGI caps). A test asserting the config says what the config says would add no gate: Tempo hard-fails on an unknown key (verified — `field health_check_enabled not found in type client.PoolConfig`), so a typo or a config rename on a Tempo bump already turns `verify-up` red.
- [x] Conventional Commits referencing the issue (`refs #156`).
- [ ] CI green — the point of the change.
- [x] `docker compose up` health unaffected (Tempo is not in `WAIT_SVCS`; config-only change, same image).
- [x] Docs updated — ADR-0023 Consequences.
- [x] ADR — amended **ADR-0023** rather than adding a new one: this is a consequence of that ADR's single-binary Tempo choice, not a new decision (one decision per ADR, §12).
- [x] Demo note — N/A, not user-visible.
## Notes for reviewers
Verified locally against the built image (the flake itself is not locally reproducible — see the runner-load point above):
1. `docker run --rm register-referentie/tempo:dev -config.file=/etc/tempo.yaml -config.verify=true` → parses.
2. `GET /status/config` on the running container → `healthcheckenabled: false` (was `true`).
3. The new diagnostic reads `tempo_distributor_ingester_clients` off a live Tempo.
Worth knowing: that metric is legitimately `0` on an idle Tempo — the pool is populated lazily on first push. It only prints on the failure path of a check that has already generated traffic, so the reading is meaningful there, but don't read a bare `0` on a quiet stack as an eviction.
Follow-up left undone: if `verify-tracing` still flakes after this, the next suspect is the .NET OTLP exporter timeout (#156's last note), not Tempo's memory cap.Reviewed-on: #157
83 lines
3.9 KiB
Markdown
83 lines
3.9 KiB
Markdown
# ADR-0023: Grafana-native observability stack (Tempo + Prometheus + Grafana)
|
|
|
|
- **Status:** Accepted
|
|
- **Date:** 2026-07-23
|
|
- **Deciders:** Respellion engineering
|
|
- **Slice:** S-16a (#122), first of the S-16 (#17) split
|
|
|
|
## Context
|
|
|
|
The PRD calls for "OpenTelemetry traces, Prometheus metrics; a local Grafana with
|
|
pre-built dashboards" (§80). S-16 was split (CLAUDE.md §13) into a backplane slice
|
|
(this one), distributed tracing (#123), and metrics + dashboards (#124). The
|
|
backplane must stand up first: a local, CI-friendly place for traces and metrics to
|
|
land, viewable in one UI, reaching green health within the 3-minute compose budget.
|
|
|
|
Two shape decisions are non-obvious enough to record.
|
|
|
|
## Decision
|
|
|
|
**Run a Grafana-native stack — Grafana Tempo (traces) + Prometheus (metrics) +
|
|
Grafana (UI) — with the services exporting OTLP straight to Tempo (no collector),
|
|
and ship the config baked into small built images.**
|
|
|
|
### Trace backend: Tempo (not Jaeger)
|
|
|
|
Tempo keeps everything under one Grafana pane alongside metrics (and later logs),
|
|
which is exactly the "local Grafana with dashboards" the PRD asks for. Jaeger would
|
|
add a second UI and a second mental model for no benefit at this scale.
|
|
|
|
### No OTLP collector
|
|
|
|
Tempo ingests OTLP directly (gRPC 4317 / HTTP 4318) and Prometheus scrapes each
|
|
service's `/metrics`, so a collector would be a hop that processes nothing. Skipped.
|
|
If we later need fan-out, tail sampling, or log processing, a collector is an
|
|
additive change — the services already speak OTLP.
|
|
|
|
### Config baked into built images, not config volumes
|
|
|
|
The upstream Common Ground modules (OpenZaak, NRC, Keycloak, Flowable) run as
|
|
**verbatim** images and get their config streamed into external named volumes by
|
|
`infra/seed-config.sh`, because bind mounts don't reach sibling containers on the
|
|
CI runner (see `docs/runbooks/gitea-actions-gotchas.md`). The observability tools
|
|
are **not** peer modules we must run verbatim, so we take the simpler path: a
|
|
three-line `Dockerfile` per tool that `COPY`s its config in. This reaches sibling
|
|
containers everywhere (docker, podman, CI) with no seed step, no `CFG_VOLS` entry,
|
|
and no Makefile sprawl.
|
|
|
|
### Verified, not assumed
|
|
|
|
`infra/run-observability-check.sh` (the `verify-observability` step, run early in CI
|
|
`verify-stack`) asks Grafana to reach both datasources — Prometheus via its health
|
|
method, Tempo via the datasource proxy (Tempo's Grafana plugin implements no health
|
|
method) — so the check proves the datasources are actually wired, not merely that
|
|
containers started. The containers are not in `WAIT_SVCS`; the check polls Grafana
|
|
itself, so no in-image healthcheck tool is required.
|
|
|
|
## Consequences
|
|
|
|
**Positive**
|
|
|
|
- One UI for traces + metrics + (future) logs. Config is versioned in
|
|
`infra/observability/` and self-contained in the images.
|
|
- Backplane is independent of app instrumentation — #123 and #124 build on it.
|
|
|
|
**Negative / costs**
|
|
|
|
- Three more images built each CI run (kept small; not on the health-gate list).
|
|
- Storage is ephemeral container fs — a demo backplane, not a retention target.
|
|
Object storage for Tempo / remote-write for Prometheus is a later concern.
|
|
- Tempo runs **single-binary**, so its distributor and ingester are one process and
|
|
some of its distributed-mode machinery is not just redundant but harmful. Its
|
|
ingester-pool health check is disabled (`ingester_client.pool_config`) because with
|
|
a single in-process ingester the check can never route around a failure — a 1s
|
|
loopback-gRPC deadline missed under CI load only evicted the one ingester and made
|
|
Tempo drop spans, which is how `verify-tracing` flaked (#156). Expect the same
|
|
shape from other distributed-mode knobs if we tune them; the fix is to switch to
|
|
real multi-ingester Tempo, not to re-enable them here.
|
|
|
|
## Coupling rules touched (CLAUDE.md §8)
|
|
|
|
None. The stack is passive infrastructure: services *push* OTLP and *expose*
|
|
`/metrics`; nothing in the stack calls into a service or a peer module.
|