feat(infra): observability backplane — Tempo + Prometheus + Grafana (S-16a, closes #122) (#125)
CI / mutation (push) Successful in 6m22s
CI / verify-stack (push) Successful in 11m53s
CI / lint (push) Successful in 1m24s
CI / build (push) Successful in 1m6s
CI / unit (push) Successful in 1m23s
CI / frontend (push) Successful in 2m54s

## What & why

S-16a, the first of the **S-16 split** (#17 closed → #122/#123/#124, §13). Stands up a local, CI-friendly observability backplane so traces (S-16b) and metrics (S-16c) have somewhere to land, viewable in one Grafana.

- **Grafana Tempo** — OTLP trace ingest (gRPC 4317 / HTTP 4318), local storage.
- **Prometheus** — scrapes itself for now; service `/metrics` targets arrive in S-16c.
- **Grafana** — Tempo + Prometheus datasources auto-provisioned with fixed uids (`tempo`, `prometheus`), exposed on :3000.

All three are small **built images** with config baked in (`infra/observability/`), on the existing `cg` network. **No OTLP collector** (Tempo ingests OTLP directly; Prometheus scrapes) and **no config-volume seeding** — the tools aren't verbatim CG peer modules, so a 3-line `COPY` Dockerfile is the simpler path that still reaches sibling containers on the CI runner (**ADR-0023**).

### Verified, not assumed

`make verify-observability` (new CI `verify-stack` step, run early) asks Grafana to reach both datasources — Prometheus via its health method, Tempo via the datasource proxy (Tempo's plugin implements no health method) — so it proves the datasources are wired, not merely that containers booted. Validated locally against the three containers (no external egress): Grafana healthy, both datasources reachable.

Closes #122

## Definition of Done

- [x] Failing test committed first (`verify-observability` fails with no backplane).
- [x] Implementation makes it pass; verified locally.
- [x] Conventional Commits referencing the issue (`refs #122`).
- [ ] CI green — awaiting Gitea Actions (verify-stack now includes the observability step; `docker compose config` validates locally).
- [ ] `docker compose up` reaches green health within 3 min — new containers are lightweight and off the health-gate list.
- [x] Docs — ADR-0023, demo-script, BACKLOG sync.
- [x] ADR added — `docs/architecture/adr-0023-observability-stack.md`.
- [x] Demo note in `docs/demo-script.md`.

## Notes for reviewers

- **No app changes** — this is pure infra; the five services are untouched (instrumentation is #123/#124).
- **Ports:** Grafana 3000 (admin/admin, anonymous viewer on), Prometheus 9090; Tempo internal to `cg`.
- **CI:** the three containers are added to the failure log-dump list; deliberately **not** added to `WAIT_SVCS` (the check polls Grafana itself, so no in-image healthcheck tool is needed). Trades ~3 small image builds per run.
- **Next:** #123 wires OTLP export + `AddAspNetCoreInstrumentation`/`AddHttpClientInstrumentation` into the five hosts so a request becomes one connected trace in Tempo.

Reviewed-on: #125
This commit was merged in pull request #125.
This commit is contained in:
not
2026-07-23 12:26:22 +00:00
parent 4fe9915816
commit 4274fd30d1
13 changed files with 263 additions and 3 deletions
+39
View File
@@ -524,6 +524,45 @@ services:
condition: service_started
networks: [cg]
# ── Observability backplane (S-16a, ADR-0023) ──────────────────────────────
# Grafana-native stack: Tempo ingests OTLP traces (the .NET services export
# straight to it — no collector hop, S-16b), Prometheus scrapes service
# /metrics (S-16c), and Grafana reads both with datasources auto-provisioned.
# Config is baked into small built images (COPY) rather than streamed into
# external config volumes like the upstream CG modules — these aren't verbatim
# peer images, so a built image is the simpler path that still reaches sibling
# containers on the CI runner. Not in WAIT_SVCS: run-observability-check.sh
# polls Grafana itself, so no in-image healthcheck tool is needed.
tempo:
build:
context: ./observability/tempo
image: register-referentie/tempo:dev
command: ["-config.file=/etc/tempo.yaml"]
networks: [cg]
prometheus:
build:
context: ./observability/prometheus
image: register-referentie/prometheus:dev
ports:
- "9090:9090"
networks: [cg]
grafana:
build:
context: ./observability/grafana
image: register-referentie/grafana:dev
environment:
GF_SECURITY_ADMIN_USER: admin
GF_SECURITY_ADMIN_PASSWORD: admin
GF_AUTH_ANONYMOUS_ENABLED: "true"
ports:
- "3000:3000"
depends_on:
- tempo
- prometheus
networks: [cg]
volumes:
oz-db:
nrc-db:
+4
View File
@@ -0,0 +1,4 @@
# Grafana with datasources baked in via provisioning (S-16a, ADR-0023).
# Dashboards (S-16c, #124) are added under provisioning/dashboards later.
FROM grafana/grafana:11.3.0
COPY provisioning/ /etc/grafana/provisioning/
@@ -0,0 +1,17 @@
# Auto-provisioned datasources (S-16a, ADR-0023). Fixed uids so dashboards (S-16c)
# and the verify-observability check can reference them by a stable id.
apiVersion: 1
datasources:
- name: Prometheus
uid: prometheus
type: prometheus
access: proxy
url: http://prometheus:9090
isDefault: true
- name: Tempo
uid: tempo
type: tempo
access: proxy
url: http://tempo:3200
@@ -0,0 +1,2 @@
FROM prom/prometheus:v2.55.1
COPY prometheus.yml /etc/prometheus/prometheus.yml
@@ -0,0 +1,10 @@
# Prometheus scrape config (S-16a, ADR-0023). For the backplane slice it scrapes
# only itself; the .NET services' /metrics scrape targets are added in S-16c
# (#124) when the services expose metrics.
global:
scrape_interval: 15s
scrape_configs:
- job_name: prometheus
static_configs:
- targets: ['localhost:9090']
+4
View File
@@ -0,0 +1,4 @@
# Tempo with our config baked in — so it reaches sibling containers on the CI
# runner without the external-config-volume dance the upstream CG images need.
FROM grafana/tempo:2.6.1
COPY tempo.yaml /etc/tempo.yaml
+27
View File
@@ -0,0 +1,27 @@
# Grafana Tempo — single-binary, all-in-one, local storage (S-16a, ADR-0023).
# Ingests OTLP directly (services export straight to Tempo; no collector hop).
# Storage is ephemeral container fs — this is a local/CI demo backplane, not a
# retention target. ponytail: local backend, swap for object storage if traces
# must outlive the stack.
server:
http_listen_port: 3200
distributor:
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
ingester:
max_block_duration: 5m
storage:
trace:
backend: local
local:
path: /var/tempo/blocks
wal:
path: /var/tempo/wal
+45
View File
@@ -0,0 +1,45 @@
#!/usr/bin/env bash
#
# S-16a (#122): assert the observability backplane is live against an ALREADY-RUNNING
# stack. Runs curl INSIDE the compose network (like the other verify checks) because
# the stack's published ports aren't on the CI runner's localhost — the stack is a set
# of sibling containers on the host daemon. It asks Grafana to reach its provisioned
# datasources — Prometheus via its health method, Tempo via the datasource proxy (Tempo's
# Grafana plugin implements no health method) — so it proves the datasources are wired,
# not merely that the containers started. Polls, so it tolerates a cold Grafana.
#
# Does NOT manage the stack lifecycle (the caller owns bring-up + teardown).
set -euo pipefail
TIMEOUT="${OBS_TIMEOUT:-60}"
AUTH="${GRAFANA_AUTH:-admin:admin}"
gf="$(docker ps -q --filter 'name=[-_]grafana[-_]' | head -1)"
[ -n "$gf" ] || { echo "ERROR: no running grafana container — bring the stack up first" >&2; exit 1; }
net="$(docker inspect -f '{{range $k,$_ := .NetworkSettings.Networks}}{{$k}}{{"\n"}}{{end}}' "$gf" | head -1)"
gf_ip="$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$gf")"
base="http://$gf_ip:3000"
echo ">> grafana=$gf_ip network=$net"
# Run curl inside a throwaway container on the stack network (reaches services by IP).
net_curl() { docker run --rm --network "$net" curlimages/curl:latest "$@"; }
# poll <description> <grep -E pattern> <curl args...>
poll() {
local desc="$1" pat="$2"; shift 2
local deadline=$(( $(date +%s) + TIMEOUT ))
while :; do
if net_curl -fsS "$@" 2>/dev/null | grep -Eq "$pat"; then echo "$desc"; return 0; fi
if [ "$(date +%s)" -ge "$deadline" ]; then echo "$desc ($*)" >&2; return 1; fi
sleep 3
done
}
echo "Checking observability backplane at $base ..."
poll "Grafana is healthy" \
'"database":[[:space:]]*"ok"' "$base/api/health"
poll "Prometheus datasource reachable" \
'"status":[[:space:]]*"OK"' -u "$AUTH" "$base/api/datasources/uid/prometheus/health"
poll "Tempo datasource reachable (via Grafana proxy)" \
'"version"' -u "$AUTH" "$base/api/datasources/proxy/uid/tempo/api/status/buildinfo"
echo "Observability backplane OK."