feat(obs): golden-signal metrics on /metrics + Prometheus scrape + Grafana dashboard (S-16c, refs #124)
CI / lint (pull_request) Successful in 5m7s
CI / unit (pull_request) Successful in 2m2s
CI / frontend (pull_request) Successful in 3m58s
CI / mutation (pull_request) Successful in 6m30s
CI / verify-stack (pull_request) Successful in 9m34s
CI / build (pull_request) Successful in 4m57s
CI / lint (pull_request) Successful in 5m7s
CI / unit (pull_request) Successful in 2m2s
CI / frontend (pull_request) Successful in 3m58s
CI / mutation (pull_request) Successful in 6m30s
CI / verify-stack (pull_request) Successful in 9m34s
CI / build (pull_request) Successful in 4m57s
Wire OTel metrics into the four remaining .NET services (acl, domain, event-subscriber, projection-api) exactly as the BFF: ASP.NET Core + HttpClient instrumentation + the built-in System.Runtime meter, exposed at /metrics via the Prometheus AspNetCore exporter (ADR-0024). Prometheus scrapes one job per service; Grafana ships a pre-built 'Request path — golden signals' dashboard (traffic/errors/latency/saturation). A verify-metrics CI step proves the endpoints are scraped end to end.
This commit is contained in:
@@ -5,6 +5,34 @@ copy-pasteable walkthrough against a local `make up` stack.
|
||||
|
||||
---
|
||||
|
||||
## S-16c — Prometheus metrics + golden-signal Grafana dashboard (#124, ADR-0023)
|
||||
|
||||
**Outcome:** the five .NET services now expose OpenTelemetry metrics in Prometheus format at `/metrics`
|
||||
— ASP.NET Core + `HttpClient` instrumentation plus the built-in `System.Runtime` meter. Prometheus
|
||||
scrapes each service (one job per service), and a **pre-built Grafana dashboard** — *Request path —
|
||||
golden signals* — plots the four golden signals: **traffic** (req/s), **errors** (5xx/s), **latency**
|
||||
(p95 request duration), and **saturation** (CPU cores in use), split by service. It populates under load.
|
||||
|
||||
```bash
|
||||
# 1. Automated (a CI verify-stack step): generate BFF traffic and assert Prometheus scraped the
|
||||
# golden-signal metric from every service.
|
||||
make verify-metrics # → OK — targets up: [...]; request metric scraped from: [...]
|
||||
|
||||
# 2. By hand: drive the stack, generate some load, then open the dashboard.
|
||||
make up
|
||||
for i in $(seq 1 50); do curl -s localhost:8080/openbaar/register >/dev/null; done # BFF → projection-api
|
||||
open http://localhost:3000 # Grafana → Dashboards → "Request path — golden signals"
|
||||
open http://localhost:9090/targets # Prometheus → every service target UP
|
||||
```
|
||||
|
||||
**The path:** each host adds `.WithMetrics(AddAspNetCoreInstrumentation + AddHttpClientInstrumentation +
|
||||
AddMeter("System.Runtime") + AddPrometheusExporter)` and maps `/metrics`; Prometheus scrapes
|
||||
`<service>:8080/metrics` (config in `infra/observability/prometheus/prometheus.yml`); Grafana ships the
|
||||
dashboard via provisioning against the fixed `prometheus` datasource uid. No metrics are pushed over
|
||||
OTLP — Prometheus pulls, so there is no collector hop (ADR-0023).
|
||||
|
||||
---
|
||||
|
||||
## S-16b — distributed traces across the .NET services (#123, ADR-0023)
|
||||
|
||||
**Outcome:** the five .NET services (BFF, Domain, ACL, projection-api, event-subscriber) now emit
|
||||
|
||||
Reference in New Issue
Block a user