Compare commits
5
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
4ac2c3ff6c | ||
|
|
965782dd95 | ||
|
|
61805d5ce7 | ||
|
|
6771fccf47 | ||
|
|
88338396f6 |
@@ -9,6 +9,12 @@ on:
|
|||||||
permissions:
|
permissions:
|
||||||
contents: read
|
contents: read
|
||||||
|
|
||||||
|
# Supersede stale runs: a new push to the same branch/PR cancels the previous run, so the runner's
|
||||||
|
# concurrency slots aren't spent on commits nobody is waiting for (refs #127).
|
||||||
|
concurrency:
|
||||||
|
group: ci-${{ github.workflow }}-${{ github.ref }}
|
||||||
|
cancel-in-progress: true
|
||||||
|
|
||||||
# Self-hosted runner — see docs/runbooks/ci.md for the runner setup.
|
# Self-hosted runner — see docs/runbooks/ci.md for the runner setup.
|
||||||
# `uses:` are absolute, tag-pinned URLs (CLAUDE.md §8.7 / §15).
|
# `uses:` are absolute, tag-pinned URLs (CLAUDE.md §8.7 / §15).
|
||||||
|
|
||||||
@@ -129,12 +135,20 @@ jobs:
|
|||||||
path: services/bff/StrykerOutput/**/reports/mutation-report.html
|
path: services/bff/StrykerOutput/**/reports/mutation-report.html
|
||||||
if-no-files-found: warn
|
if-no-files-found: warn
|
||||||
|
|
||||||
# One stage for every check that needs the live stack. On the single self-hosted
|
# One stage for every check that needs the live stack. Booting OpenZaak once (instead
|
||||||
# runner jobs run sequentially, so booting OpenZaak once (instead of once per job)
|
# of once per job) is the cheapest layout (issue #58). No setup-dotnet: the ACL test runs
|
||||||
# is the cheapest layout (issue #58). No setup-dotnet: the ACL test runs in a built
|
# in a built image and everything reaches services by container IP. Needs Docker + egress
|
||||||
# image and everything reaches services by container IP. Needs Docker + egress
|
|
||||||
# (base images, nuget, selectielijst.openzaak.nl).
|
# (base images, nuget, selectielijst.openzaak.nl).
|
||||||
|
#
|
||||||
|
# `needs: [mutation]` is NOT a data dependency — it serialises the two memory-heavy jobs so
|
||||||
|
# they never co-schedule now the runner has capacity >1. A concurrent Stryker run + full-stack
|
||||||
|
# bring-up + Playwright browser on one host is what OOMs the e2e (commit d5e5fa2, #126). The
|
||||||
|
# light .NET/frontend jobs have no `needs`, so they still parallelise up to runner capacity.
|
||||||
|
# `if: !cancelled()` keeps verify-stack running even when the mutation ratchet fails (so we don't
|
||||||
|
# lose its signal) while still honouring run cancellation from the concurrency group above.
|
||||||
verify-stack:
|
verify-stack:
|
||||||
|
needs: [mutation]
|
||||||
|
if: ${{ !cancelled() }}
|
||||||
runs-on: ubuntu-latest
|
runs-on: ubuntu-latest
|
||||||
steps:
|
steps:
|
||||||
- uses: https://github.com/actions/checkout@v4
|
- uses: https://github.com/actions/checkout@v4
|
||||||
@@ -154,6 +168,10 @@ jobs:
|
|||||||
run: make verify-domain
|
run: make verify-domain
|
||||||
- name: BFF → Keycloak + domain + projection
|
- name: BFF → Keycloak + domain + projection
|
||||||
run: make verify-bff
|
run: make verify-bff
|
||||||
|
- name: Distributed traces reach Tempo (one connected trace across services)
|
||||||
|
run: TRACING_TIMEOUT=120 make verify-tracing
|
||||||
|
- name: Golden-signal metrics scraped by Prometheus (/metrics on every service)
|
||||||
|
run: METRICS_TIMEOUT=120 make verify-metrics
|
||||||
- name: Self-service e2e (Playwright, login → submit → success)
|
- name: Self-service e2e (Playwright, login → submit → success)
|
||||||
run: make verify-e2e
|
run: make verify-e2e
|
||||||
# Log dump must precede teardown (which removes the containers).
|
# Log dump must precede teardown (which removes the containers).
|
||||||
|
|||||||
@@ -57,3 +57,4 @@ vitest.config.*.timestamp*
|
|||||||
tests/e2e/node_modules/
|
tests/e2e/node_modules/
|
||||||
tests/e2e/test-results/
|
tests/e2e/test-results/
|
||||||
tests/e2e/playwright-report/
|
tests/e2e/playwright-report/
|
||||||
|
__pycache__/
|
||||||
|
|||||||
+1
-1
@@ -260,7 +260,7 @@ Split (issue #11 closed) into two independently-demoable slices per §13 — the
|
|||||||
Split into independently deployable sub-slices (CLAUDE.md §13):
|
Split into independently deployable sub-slices (CLAUDE.md §13):
|
||||||
|
|
||||||
- **S-16a** (#122) · Observability backplane — Grafana Tempo + Prometheus + Grafana in compose, datasources auto-provisioned (ADR-0023). No collector; config baked into built images.
|
- **S-16a** (#122) · Observability backplane — Grafana Tempo + Prometheus + Grafana in compose, datasources auto-provisioned (ADR-0023). No collector; config baked into built images.
|
||||||
- **S-16b** (#123) · Distributed traces across the five .NET services (OTLP → Tempo; nginx `traceparent` passthrough). Depends on S-16a.
|
- **S-16b** (#123) · Distributed traces across the five .NET services (OTLP → Tempo; traceparent propagates via the typed HttpClients). Depends on S-16a. ✅
|
||||||
- **S-16c** (#124) · Prometheus metrics + golden-signal Grafana dashboards. Depends on S-16a.
|
- **S-16c** (#124) · Prometheus metrics + golden-signal Grafana dashboards. Depends on S-16a.
|
||||||
|
|
||||||
### S-17 · Quartz.NET scheduler — herregistratie reminder sweep ✅
|
### S-17 · Quartz.NET scheduler — herregistratie reminder sweep ✅
|
||||||
|
|||||||
@@ -43,7 +43,7 @@ export DOCKER_HOST := unix://$(PODMAN_SOCK)
|
|||||||
endif
|
endif
|
||||||
endif
|
endif
|
||||||
|
|
||||||
.PHONY: ci lint build unit mutation frontend integration verify verify-up verify-acl verify-nrc verify-projection verify-bff verify-domain verify-observability verify-notifications smoke up down local verify-local local-down changelog openzaak-up openzaak-smoke openzaak-seed openzaak-down stack-up stack-smoke stack-down keycloak-up keycloak-smoke keycloak-down flowable-up flowable-smoke flowable-down help
|
.PHONY: ci lint build unit mutation frontend integration verify verify-up verify-acl verify-nrc verify-projection verify-bff verify-domain verify-observability verify-tracing verify-metrics verify-notifications smoke up down local verify-local local-down changelog openzaak-up openzaak-smoke openzaak-seed openzaak-down stack-up stack-smoke stack-down keycloak-up keycloak-smoke keycloak-down flowable-up flowable-smoke flowable-down help
|
||||||
|
|
||||||
## ci: run the full pipeline — lint, build, unit, mutation, frontend, verify (mirrors Gitea Actions)
|
## ci: run the full pipeline — lint, build, unit, mutation, frontend, verify (mirrors Gitea Actions)
|
||||||
## `verify` is the live-stack stage (full stack up once → ACL + notification checks).
|
## `verify` is the live-stack stage (full stack up once → ACL + notification checks).
|
||||||
@@ -175,6 +175,16 @@ verify-e2e:
|
|||||||
verify-observability:
|
verify-observability:
|
||||||
bash infra/run-observability-check.sh
|
bash infra/run-observability-check.sh
|
||||||
|
|
||||||
|
## verify-tracing: assert one connected distributed trace spans the .NET services in Tempo
|
||||||
|
## (S-16b), against the already-running stack.
|
||||||
|
verify-tracing:
|
||||||
|
bash infra/run-tracing-check.sh
|
||||||
|
|
||||||
|
## verify-metrics: assert the services expose /metrics and Prometheus scrapes the golden
|
||||||
|
## signals (S-16c), against the already-running stack.
|
||||||
|
verify-metrics:
|
||||||
|
bash infra/run-metrics-check.sh
|
||||||
|
|
||||||
## verify: local mirror of the CI verify-stack job — full stack up once, all checks,
|
## verify: local mirror of the CI verify-stack job — full stack up once, all checks,
|
||||||
## tear down (always). For fast single-concern local iteration use `integration`
|
## tear down (always). For fast single-concern local iteration use `integration`
|
||||||
## (oz-only) or `verify-notifications` (oz+nrc) instead.
|
## (oz-only) or `verify-notifications` (oz+nrc) instead.
|
||||||
|
|||||||
@@ -0,0 +1,53 @@
|
|||||||
|
# ADR-0024: Expose OTel metrics with the (prerelease) Prometheus AspNetCore exporter
|
||||||
|
|
||||||
|
- **Status:** Accepted
|
||||||
|
- **Date:** 2026-07-24
|
||||||
|
- **Deciders:** Respellion engineering
|
||||||
|
- **Slice:** S-16c (#124), last of the S-16 (#17) split
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
ADR-0023 already fixed the shape of metrics collection: **Prometheus scrapes each
|
||||||
|
service's `/metrics`** (pull, no collector). S-16c implements it. That needs a package
|
||||||
|
that turns the OpenTelemetry `MeterProvider` into a Prometheus scrape endpoint inside
|
||||||
|
ASP.NET Core. The canonical one is `OpenTelemetry.Exporter.Prometheus.AspNetCore`
|
||||||
|
(`AddPrometheusExporter()` + `app.MapPrometheusScrapingEndpoint()`).
|
||||||
|
|
||||||
|
The catch: that exporter has **never had a stable release** — the whole OTel .NET
|
||||||
|
Prometheus exporter line is versioned `-beta` (we pin `1.17.0-beta.1`, matched to the
|
||||||
|
`1.17.0` core we already use). Adding it is a new dependency (CLAUDE.md §14), and taking
|
||||||
|
a prerelease package into all five services is the decision worth recording.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
**Add `OpenTelemetry.Exporter.Prometheus.AspNetCore` `1.17.0-beta.1` to the five .NET
|
||||||
|
services and expose `/metrics` with it.**
|
||||||
|
|
||||||
|
- What it gives us: the OTel-native pull endpoint, so the meters we already register for
|
||||||
|
tracing-adjacent instrumentation surface as Prometheus text with zero extra plumbing.
|
||||||
|
- What we'd write to replace it: a hand-rolled `IMetricsListener`/`MeterListener` that
|
||||||
|
formats Prometheus exposition text — real work, and a reimplementation of a widely-used
|
||||||
|
library for no gain.
|
||||||
|
- Risk it adds: a prerelease API that can shift between betas. Contained: it is only
|
||||||
|
wired in `Program.cs` (two calls per service, excluded from mutation), the version is
|
||||||
|
pinned, and `verify-metrics` proves the endpoint + scrape actually work each CI run.
|
||||||
|
|
||||||
|
The alternative — pushing metrics over OTLP to a collector that re-exposes them — was
|
||||||
|
already rejected in ADR-0023 (no collector hop). Not revisited here.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
**Positive**
|
||||||
|
|
||||||
|
- Golden-signal metrics on `/metrics` with the standard OTel names
|
||||||
|
(`http_server_request_duration_seconds`, `dotnet_*`), scraped straight by Prometheus.
|
||||||
|
- No collector, no bespoke exposition code.
|
||||||
|
|
||||||
|
**Negative / costs**
|
||||||
|
|
||||||
|
- A `-beta` package in production services. Mitigated by the pin + the `verify-metrics`
|
||||||
|
CI gate; upgrading tracks the OTel core version bumps.
|
||||||
|
|
||||||
|
## Coupling rules touched (CLAUDE.md §8)
|
||||||
|
|
||||||
|
None. Metrics are passive: Prometheus pulls; no service calls into the stack.
|
||||||
@@ -5,6 +5,61 @@ copy-pasteable walkthrough against a local `make up` stack.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## S-16c — Prometheus metrics + golden-signal Grafana dashboard (#124, ADR-0023)
|
||||||
|
|
||||||
|
**Outcome:** the five .NET services now expose OpenTelemetry metrics in Prometheus format at `/metrics`
|
||||||
|
— ASP.NET Core + `HttpClient` instrumentation plus the built-in `System.Runtime` meter. Prometheus
|
||||||
|
scrapes each service (one job per service), and a **pre-built Grafana dashboard** — *Request path —
|
||||||
|
golden signals* — plots the four golden signals: **traffic** (req/s), **errors** (5xx/s), **latency**
|
||||||
|
(p95 request duration), and **saturation** (CPU cores in use), split by service. It populates under load.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 1. Automated (a CI verify-stack step): generate BFF traffic and assert Prometheus scraped the
|
||||||
|
# golden-signal metric from every service.
|
||||||
|
make verify-metrics # → OK — targets up: [...]; request metric scraped from: [...]
|
||||||
|
|
||||||
|
# 2. By hand: drive the stack, generate some load, then open the dashboard.
|
||||||
|
make up
|
||||||
|
for i in $(seq 1 50); do curl -s localhost:8080/openbaar/register >/dev/null; done # BFF → projection-api
|
||||||
|
open http://localhost:3000 # Grafana → Dashboards → "Request path — golden signals"
|
||||||
|
open http://localhost:9090/targets # Prometheus → every service target UP
|
||||||
|
```
|
||||||
|
|
||||||
|
**The path:** each host adds `.WithMetrics(AddAspNetCoreInstrumentation + AddHttpClientInstrumentation +
|
||||||
|
AddMeter("System.Runtime") + AddPrometheusExporter)` and maps `/metrics`; Prometheus scrapes
|
||||||
|
`<service>:8080/metrics` (config in `infra/observability/prometheus/prometheus.yml`); Grafana ships the
|
||||||
|
dashboard via provisioning against the fixed `prometheus` datasource uid. No metrics are pushed over
|
||||||
|
OTLP — Prometheus pulls, so there is no collector hop (ADR-0023).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## S-16b — distributed traces across the .NET services (#123, ADR-0023)
|
||||||
|
|
||||||
|
**Outcome:** the five .NET services (BFF, Domain, ACL, projection-api, event-subscriber) now emit
|
||||||
|
OpenTelemetry traces — ASP.NET Core + `HttpClient` auto-instrumentation, exported over OTLP to Tempo.
|
||||||
|
Because every cross-service call goes through a typed `HttpClient`, the W3C `traceparent` propagates for
|
||||||
|
free, so a request is **one connected trace** across the services (bff → domain → acl → openzaak;
|
||||||
|
bff → projection-api). `/health` is filtered out. No browser-side instrumentation yet, so the trace
|
||||||
|
begins at the BFF; the async Flowable-poll boundary is a separate trace (ADR-0023).
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 1. Automated (a CI verify-stack step): generate BFF traffic and assert Tempo holds one trace
|
||||||
|
# spanning multiple services.
|
||||||
|
make verify-tracing # → OK — trace <id> spans ['bff', 'projection-api']
|
||||||
|
|
||||||
|
# 2. By hand: drive the stack, then explore traces in Grafana.
|
||||||
|
make up
|
||||||
|
curl -s localhost:8080/openbaar/register >/dev/null # BFF → projection-api
|
||||||
|
open http://localhost:3000 # Grafana → Explore → Tempo → Search → service.name = bff → open a trace
|
||||||
|
```
|
||||||
|
|
||||||
|
**The path:** each host wires `AddOpenTelemetry().WithTracing(AddAspNetCoreInstrumentation +
|
||||||
|
AddHttpClientInstrumentation + AddOtlpExporter)`; `OTEL_SERVICE_NAME` / `OTEL_EXPORTER_OTLP_ENDPOINT`
|
||||||
|
come from compose; spans export to **tempo:4317** and render in Grafana against the provisioned Tempo
|
||||||
|
datasource.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## S-16a — observability backplane: Tempo + Prometheus + Grafana (#122, ADR-0023)
|
## S-16a — observability backplane: Tempo + Prometheus + Grafana (#122, ADR-0023)
|
||||||
|
|
||||||
**Outcome:** the compose stack now includes a Grafana-native observability backplane — **Tempo** (OTLP
|
**Outcome:** the compose stack now includes a Grafana-native observability backplane — **Tempo** (OTLP
|
||||||
|
|||||||
@@ -296,6 +296,10 @@ services:
|
|||||||
dockerfile: Dockerfile
|
dockerfile: Dockerfile
|
||||||
image: register-referentie/acl:dev
|
image: register-referentie/acl:dev
|
||||||
environment:
|
environment:
|
||||||
|
# OpenTelemetry traces → Tempo (S-16b, ADR-0023).
|
||||||
|
OTEL_EXPORTER_OTLP_ENDPOINT: http://tempo:4317
|
||||||
|
OTEL_EXPORTER_OTLP_PROTOCOL: grpc
|
||||||
|
OTEL_SERVICE_NAME: acl
|
||||||
# Overridable so verify-domain can point the ACL at the same OpenZaak host that
|
# Overridable so verify-domain can point the ACL at the same OpenZaak host that
|
||||||
# owns the seeded zaaktype URL (host-consistent zaak creation, ADR-0009).
|
# owns the seeded zaaktype URL (host-consistent zaak creation, ADR-0009).
|
||||||
Acl__OpenZaak__BaseUrl: ${ACL_OPENZAAK_BASEURL:-http://openzaak:8000/}
|
Acl__OpenZaak__BaseUrl: ${ACL_OPENZAAK_BASEURL:-http://openzaak:8000/}
|
||||||
@@ -334,6 +338,10 @@ services:
|
|||||||
dockerfile: Dockerfile
|
dockerfile: Dockerfile
|
||||||
image: register-referentie/domain:dev
|
image: register-referentie/domain:dev
|
||||||
environment:
|
environment:
|
||||||
|
# OpenTelemetry traces → Tempo (S-16b, ADR-0023).
|
||||||
|
OTEL_EXPORTER_OTLP_ENDPOINT: http://tempo:4317
|
||||||
|
OTEL_EXPORTER_OTLP_PROTOCOL: grpc
|
||||||
|
OTEL_SERVICE_NAME: domain
|
||||||
Flowable__BaseUrl: http://flowable-rest:8080/flowable-rest/
|
Flowable__BaseUrl: http://flowable-rest:8080/flowable-rest/
|
||||||
Flowable__Username: rest-admin
|
Flowable__Username: rest-admin
|
||||||
Flowable__Password: test
|
Flowable__Password: test
|
||||||
@@ -360,6 +368,10 @@ services:
|
|||||||
dockerfile: Dockerfile
|
dockerfile: Dockerfile
|
||||||
image: register-referentie/bff:dev
|
image: register-referentie/bff:dev
|
||||||
environment:
|
environment:
|
||||||
|
# OpenTelemetry traces → Tempo (S-16b, ADR-0023).
|
||||||
|
OTEL_EXPORTER_OTLP_ENDPOINT: http://tempo:4317
|
||||||
|
OTEL_EXPORTER_OTLP_PROTOCOL: grpc
|
||||||
|
OTEL_SERVICE_NAME: bff
|
||||||
# The BFF is the portals' only backend; it validates digid tokens and fans out (ADR-0010).
|
# The BFF is the portals' only backend; it validates digid tokens and fans out (ADR-0010).
|
||||||
# Keycloak (start-dev) derives the issuer from the request host, so the BFF authority and the
|
# Keycloak (start-dev) derives the issuer from the request host, so the BFF authority and the
|
||||||
# verify token request both use keycloak:8080 to keep the issuer consistent.
|
# verify token request both use keycloak:8080 to keep the issuer consistent.
|
||||||
@@ -412,6 +424,10 @@ services:
|
|||||||
dockerfile: services/event-subscriber/Dockerfile
|
dockerfile: services/event-subscriber/Dockerfile
|
||||||
image: register-referentie/event-subscriber:dev
|
image: register-referentie/event-subscriber:dev
|
||||||
environment:
|
environment:
|
||||||
|
# OpenTelemetry traces → Tempo (S-16b, ADR-0023).
|
||||||
|
OTEL_EXPORTER_OTLP_ENDPOINT: http://tempo:4317
|
||||||
|
OTEL_EXPORTER_OTLP_PROTOCOL: grpc
|
||||||
|
OTEL_SERVICE_NAME: event-subscriber
|
||||||
ConnectionStrings__Projection: Host=projection-db;Database=projection;Username=projection;Password=projection
|
ConnectionStrings__Projection: Host=projection-db;Database=projection;Username=projection;Password=projection
|
||||||
# The subscriber enriches the projection with each zaak's reference (identificatie) by asking
|
# The subscriber enriches the projection with each zaak's reference (identificatie) by asking
|
||||||
# the ACL — the only code allowed to read ZGW (§8.1, #78).
|
# the ACL — the only code allowed to read ZGW (§8.1, #78).
|
||||||
@@ -441,6 +457,10 @@ services:
|
|||||||
dockerfile: services/projection-api/Dockerfile
|
dockerfile: services/projection-api/Dockerfile
|
||||||
image: register-referentie/projection-api:dev
|
image: register-referentie/projection-api:dev
|
||||||
environment:
|
environment:
|
||||||
|
# OpenTelemetry traces → Tempo (S-16b, ADR-0023).
|
||||||
|
OTEL_EXPORTER_OTLP_ENDPOINT: http://tempo:4317
|
||||||
|
OTEL_EXPORTER_OTLP_PROTOCOL: grpc
|
||||||
|
OTEL_SERVICE_NAME: projection-api
|
||||||
ConnectionStrings__Projection: Host=projection-db;Database=projection;Username=projection;Password=projection
|
ConnectionStrings__Projection: Host=projection-db;Database=projection;Username=projection;Password=projection
|
||||||
ports:
|
ports:
|
||||||
- "8120:8080"
|
- "8120:8080"
|
||||||
@@ -538,12 +558,16 @@ services:
|
|||||||
context: ./observability/tempo
|
context: ./observability/tempo
|
||||||
image: register-referentie/tempo:dev
|
image: register-referentie/tempo:dev
|
||||||
command: ["-config.file=/etc/tempo.yaml"]
|
command: ["-config.file=/etc/tempo.yaml"]
|
||||||
|
# Cap the backplane's footprint so it can't starve the app stack + the Playwright browser on the
|
||||||
|
# memory-tight CI runner (verify-e2e OOM history, commit d5e5fa2). Generous vs idle (~150M).
|
||||||
|
mem_limit: 400m
|
||||||
networks: [cg]
|
networks: [cg]
|
||||||
|
|
||||||
prometheus:
|
prometheus:
|
||||||
build:
|
build:
|
||||||
context: ./observability/prometheus
|
context: ./observability/prometheus
|
||||||
image: register-referentie/prometheus:dev
|
image: register-referentie/prometheus:dev
|
||||||
|
mem_limit: 400m
|
||||||
ports:
|
ports:
|
||||||
- "9090:9090"
|
- "9090:9090"
|
||||||
networks: [cg]
|
networks: [cg]
|
||||||
@@ -552,6 +576,7 @@ services:
|
|||||||
build:
|
build:
|
||||||
context: ./observability/grafana
|
context: ./observability/grafana
|
||||||
image: register-referentie/grafana:dev
|
image: register-referentie/grafana:dev
|
||||||
|
mem_limit: 512m
|
||||||
environment:
|
environment:
|
||||||
GF_SECURITY_ADMIN_USER: admin
|
GF_SECURITY_ADMIN_USER: admin
|
||||||
GF_SECURITY_ADMIN_PASSWORD: admin
|
GF_SECURITY_ADMIN_PASSWORD: admin
|
||||||
|
|||||||
Executable
+75
@@ -0,0 +1,75 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""S-16c (#124): prove the golden-signal metrics pipeline works end to end.
|
||||||
|
|
||||||
|
Generate anonymous BFF traffic (GET /openbaar/register — no auth, no OpenZaak egress),
|
||||||
|
then query Prometheus and assert (1) every .NET service's scrape target is UP, and (2)
|
||||||
|
the http.server.request.duration histogram is actually being scraped — i.e. the services
|
||||||
|
expose /metrics AND Prometheus collects it, which is exactly what the golden-signal
|
||||||
|
dashboard reads.
|
||||||
|
|
||||||
|
Stdlib only (urllib/json) so it runs in a bare python:3-slim container in-network.
|
||||||
|
"""
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
import time
|
||||||
|
import urllib.error
|
||||||
|
import urllib.parse
|
||||||
|
import urllib.request
|
||||||
|
|
||||||
|
BFF = os.environ["BFF"] # http://<bff-ip>:8080
|
||||||
|
PROM = os.environ["PROMETHEUS"] # http://<prometheus-ip>:9090
|
||||||
|
TIMEOUT = int(os.environ.get("METRICS_TIMEOUT", "90"))
|
||||||
|
SERVICES = {"acl", "domain", "bff", "event-subscriber", "projection-api"}
|
||||||
|
|
||||||
|
|
||||||
|
def _get(url):
|
||||||
|
with urllib.request.urlopen(url, timeout=10) as r:
|
||||||
|
return r.read()
|
||||||
|
|
||||||
|
|
||||||
|
def generate_traffic():
|
||||||
|
for _ in range(3):
|
||||||
|
try:
|
||||||
|
_get(f"{BFF}/openbaar/register")
|
||||||
|
except urllib.error.HTTPError:
|
||||||
|
pass # a non-2xx still records an http.server metric
|
||||||
|
|
||||||
|
|
||||||
|
def query(promql):
|
||||||
|
q = urllib.parse.quote(promql)
|
||||||
|
try:
|
||||||
|
data = json.loads(_get(f"{PROM}/api/v1/query?query={q}"))
|
||||||
|
except Exception:
|
||||||
|
return []
|
||||||
|
return data.get("data", {}).get("result", [])
|
||||||
|
|
||||||
|
|
||||||
|
def jobs_up():
|
||||||
|
return {r["metric"].get("job") for r in query("up == 1")}
|
||||||
|
|
||||||
|
|
||||||
|
def jobs_with_request_metric():
|
||||||
|
return {r["metric"].get("job")
|
||||||
|
for r in query("http_server_request_duration_seconds_count")}
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
deadline = time.time() + TIMEOUT
|
||||||
|
while time.time() < deadline:
|
||||||
|
generate_traffic()
|
||||||
|
up = jobs_up()
|
||||||
|
scraped = jobs_with_request_metric()
|
||||||
|
if SERVICES.issubset(up) and SERVICES.issubset(scraped):
|
||||||
|
print(f"OK — targets up: {sorted(up & SERVICES)}; "
|
||||||
|
f"request metric scraped from: {sorted(scraped & SERVICES)}")
|
||||||
|
return 0
|
||||||
|
time.sleep(3)
|
||||||
|
print(f"FAIL — up: {sorted(jobs_up() & SERVICES)}; "
|
||||||
|
f"request metric from: {sorted(jobs_with_request_metric() & SERVICES)}; "
|
||||||
|
f"expected all of {sorted(SERVICES)}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main())
|
||||||
@@ -1,4 +1,4 @@
|
|||||||
# Grafana with datasources baked in via provisioning (S-16a, ADR-0023).
|
# Grafana with datasources + the golden-signals dashboard baked in via provisioning
|
||||||
# Dashboards (S-16c, #124) are added under provisioning/dashboards later.
|
# (S-16a/S-16c, ADR-0023). Everything under provisioning/ is copied in below.
|
||||||
FROM grafana/grafana:11.3.0
|
FROM grafana/grafana:11.3.0
|
||||||
COPY provisioning/ /etc/grafana/provisioning/
|
COPY provisioning/ /etc/grafana/provisioning/
|
||||||
|
|||||||
@@ -0,0 +1,13 @@
|
|||||||
|
# Dashboard provider (S-16c, ADR-0023): Grafana loads every *.json in this folder as a
|
||||||
|
# read-only, code-owned dashboard. The golden-signals board is versioned here, not
|
||||||
|
# clicked together in the UI.
|
||||||
|
apiVersion: 1
|
||||||
|
|
||||||
|
providers:
|
||||||
|
- name: register-referentie
|
||||||
|
type: file
|
||||||
|
disableDeletion: true
|
||||||
|
allowUiUpdates: false
|
||||||
|
options:
|
||||||
|
path: /etc/grafana/provisioning/dashboards
|
||||||
|
foldersFromFilesStructure: false
|
||||||
@@ -0,0 +1,87 @@
|
|||||||
|
{
|
||||||
|
"uid": "golden-signals",
|
||||||
|
"title": "Request path — golden signals",
|
||||||
|
"tags": ["s-16c", "golden-signals"],
|
||||||
|
"timezone": "browser",
|
||||||
|
"schemaVersion": 39,
|
||||||
|
"version": 1,
|
||||||
|
"editable": true,
|
||||||
|
"refresh": "10s",
|
||||||
|
"time": { "from": "now-15m", "to": "now" },
|
||||||
|
"templating": {
|
||||||
|
"list": [
|
||||||
|
{
|
||||||
|
"name": "job",
|
||||||
|
"type": "query",
|
||||||
|
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||||
|
"query": "label_values(http_server_request_duration_seconds_count, job)",
|
||||||
|
"includeAll": true,
|
||||||
|
"multi": true,
|
||||||
|
"current": { "text": "All", "value": "$__all" },
|
||||||
|
"refresh": 2
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"panels": [
|
||||||
|
{
|
||||||
|
"id": 1,
|
||||||
|
"title": "Traffic — requests/sec",
|
||||||
|
"type": "timeseries",
|
||||||
|
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||||
|
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 0 },
|
||||||
|
"fieldConfig": { "defaults": { "unit": "reqps", "custom": { "drawStyle": "line", "fillOpacity": 10 } }, "overrides": [] },
|
||||||
|
"targets": [
|
||||||
|
{
|
||||||
|
"refId": "A",
|
||||||
|
"expr": "sum by (job) (rate(http_server_request_duration_seconds_count{job=~\"$job\"}[$__rate_interval]))",
|
||||||
|
"legendFormat": "{{job}}"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": 2,
|
||||||
|
"title": "Errors — 5xx responses/sec",
|
||||||
|
"type": "timeseries",
|
||||||
|
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||||
|
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 0 },
|
||||||
|
"fieldConfig": { "defaults": { "unit": "reqps", "custom": { "drawStyle": "line", "fillOpacity": 10 }, "color": { "mode": "fixed", "fixedColor": "red" } }, "overrides": [] },
|
||||||
|
"targets": [
|
||||||
|
{
|
||||||
|
"refId": "A",
|
||||||
|
"expr": "sum by (job) (rate(http_server_request_duration_seconds_count{job=~\"$job\",http_response_status_code=~\"5..\"}[$__rate_interval]))",
|
||||||
|
"legendFormat": "{{job}}"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": 3,
|
||||||
|
"title": "Latency — p95 request duration",
|
||||||
|
"type": "timeseries",
|
||||||
|
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||||
|
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 8 },
|
||||||
|
"fieldConfig": { "defaults": { "unit": "s", "custom": { "drawStyle": "line", "fillOpacity": 10 } }, "overrides": [] },
|
||||||
|
"targets": [
|
||||||
|
{
|
||||||
|
"refId": "A",
|
||||||
|
"expr": "histogram_quantile(0.95, sum by (job, le) (rate(http_server_request_duration_seconds_bucket{job=~\"$job\"}[$__rate_interval])))",
|
||||||
|
"legendFormat": "{{job}} p95"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": 4,
|
||||||
|
"title": "Saturation — CPU cores in use",
|
||||||
|
"type": "timeseries",
|
||||||
|
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||||
|
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 8 },
|
||||||
|
"fieldConfig": { "defaults": { "unit": "none", "custom": { "drawStyle": "line", "fillOpacity": 10 } }, "overrides": [] },
|
||||||
|
"targets": [
|
||||||
|
{
|
||||||
|
"refId": "A",
|
||||||
|
"expr": "sum by (job) (rate(dotnet_process_cpu_time_seconds_total{job=~\"$job\"}[$__rate_interval]))",
|
||||||
|
"legendFormat": "{{job}}"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -1,6 +1,7 @@
|
|||||||
# Prometheus scrape config (S-16a, ADR-0023). For the backplane slice it scrapes
|
# Prometheus scrape config (S-16c, ADR-0023). Each .NET service exposes OTel metrics
|
||||||
# only itself; the .NET services' /metrics scrape targets are added in S-16c
|
# at /metrics (Prometheus text format); one scrape job per service, so the service is
|
||||||
# (#124) when the services expose metrics.
|
# identified by the `job` label in the golden-signal dashboard. Targets are reached by
|
||||||
|
# compose service name on the shared `cg` network (internal port 8080).
|
||||||
global:
|
global:
|
||||||
scrape_interval: 15s
|
scrape_interval: 15s
|
||||||
|
|
||||||
@@ -8,3 +9,19 @@ scrape_configs:
|
|||||||
- job_name: prometheus
|
- job_name: prometheus
|
||||||
static_configs:
|
static_configs:
|
||||||
- targets: ['localhost:9090']
|
- targets: ['localhost:9090']
|
||||||
|
|
||||||
|
- job_name: acl
|
||||||
|
static_configs:
|
||||||
|
- targets: ['acl:8080']
|
||||||
|
- job_name: domain
|
||||||
|
static_configs:
|
||||||
|
- targets: ['domain:8080']
|
||||||
|
- job_name: bff
|
||||||
|
static_configs:
|
||||||
|
- targets: ['bff:8080']
|
||||||
|
- job_name: event-subscriber
|
||||||
|
static_configs:
|
||||||
|
- targets: ['event-subscriber:8080']
|
||||||
|
- job_name: projection-api
|
||||||
|
static_configs:
|
||||||
|
- targets: ['projection-api:8080']
|
||||||
|
|||||||
Executable
+28
@@ -0,0 +1,28 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
#
|
||||||
|
# S-16c (#124): assert the golden-signal metrics pipeline works — the .NET services expose
|
||||||
|
# /metrics and Prometheus scrapes them — against an ALREADY-RUNNING full stack. Runs the
|
||||||
|
# driver in a python:3-slim container on the stack network (services reached by container IP;
|
||||||
|
# the runner can't reach published ports — gitea-actions-gotchas.md §5/§6). Does NOT manage
|
||||||
|
# the stack lifecycle.
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||||
|
|
||||||
|
ip() { docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$1"; }
|
||||||
|
|
||||||
|
bff="$(docker ps -q --filter 'name=[-_]bff[-_]' | head -1)"
|
||||||
|
prom="$(docker ps -q --filter 'name=[-_]prometheus[-_]' | head -1)"
|
||||||
|
[ -n "$bff" ] && [ -n "$prom" ] || { echo "ERROR: bff and/or prometheus not running — bring the stack up first" >&2; exit 1; }
|
||||||
|
net="$(docker inspect -f '{{range $k,$_ := .NetworkSettings.Networks}}{{$k}}{{"\n"}}{{end}}' "$bff" | head -1)"
|
||||||
|
bff_ip="$(ip "$bff")"; prom_ip="$(ip "$prom")"
|
||||||
|
echo ">> network=$net bff=$bff_ip prometheus=$prom_ip"
|
||||||
|
|
||||||
|
cid="$(docker create --network "$net" \
|
||||||
|
-e "BFF=http://$bff_ip:8080" -e "PROMETHEUS=http://$prom_ip:9090" \
|
||||||
|
-e "METRICS_TIMEOUT=${METRICS_TIMEOUT:-90}" \
|
||||||
|
python:3-slim python /metrics-check.py)"
|
||||||
|
docker cp "$here/metrics-check.py" "$cid:/metrics-check.py" >/dev/null
|
||||||
|
rc=0; docker start -a "$cid" || rc=$?
|
||||||
|
docker rm -f "$cid" >/dev/null
|
||||||
|
exit $rc
|
||||||
Executable
+27
@@ -0,0 +1,27 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
#
|
||||||
|
# S-16b (#123): assert one connected distributed trace spans the .NET services in Tempo,
|
||||||
|
# against an ALREADY-RUNNING full stack. Runs the driver in a python:3-slim container on the
|
||||||
|
# stack network (services reached by container IP; the runner can't reach published ports —
|
||||||
|
# gitea-actions-gotchas.md §5/§6). Does NOT manage the stack lifecycle.
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||||
|
|
||||||
|
ip() { docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$1"; }
|
||||||
|
|
||||||
|
bff="$(docker ps -q --filter 'name=[-_]bff[-_]' | head -1)"
|
||||||
|
tempo="$(docker ps -q --filter 'name=[-_]tempo[-_]' | head -1)"
|
||||||
|
[ -n "$bff" ] && [ -n "$tempo" ] || { echo "ERROR: bff and/or tempo not running — bring the stack up first" >&2; exit 1; }
|
||||||
|
net="$(docker inspect -f '{{range $k,$_ := .NetworkSettings.Networks}}{{$k}}{{"\n"}}{{end}}' "$bff" | head -1)"
|
||||||
|
bff_ip="$(ip "$bff")"; tempo_ip="$(ip "$tempo")"
|
||||||
|
echo ">> network=$net bff=$bff_ip tempo=$tempo_ip"
|
||||||
|
|
||||||
|
cid="$(docker create --network "$net" \
|
||||||
|
-e "BFF=http://$bff_ip:8080" -e "TEMPO=http://$tempo_ip:3200" \
|
||||||
|
-e "TRACING_TIMEOUT=${TRACING_TIMEOUT:-90}" \
|
||||||
|
python:3-slim python /tracing-check.py)"
|
||||||
|
docker cp "$here/tracing-check.py" "$cid:/tracing-check.py" >/dev/null
|
||||||
|
rc=0; docker start -a "$cid" || rc=$?
|
||||||
|
docker rm -f "$cid" >/dev/null
|
||||||
|
exit $rc
|
||||||
Executable
+81
@@ -0,0 +1,81 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""S-16b (#123): prove distributed tracing works end to end.
|
||||||
|
|
||||||
|
Generate anonymous BFF traffic (GET /openbaar/register, which the BFF serves by
|
||||||
|
calling projection-api — no auth, no OpenZaak egress), then query Tempo and assert
|
||||||
|
that ONE trace contains spans from both `bff` and `projection-api`. That proves the
|
||||||
|
services export OTLP to Tempo AND that the W3C traceparent propagates across the
|
||||||
|
HttpClient hop, stitching the request into a single connected trace.
|
||||||
|
|
||||||
|
Stdlib only (urllib/json) so it runs in a bare python:3-slim container in-network.
|
||||||
|
"""
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
import time
|
||||||
|
import urllib.error
|
||||||
|
import urllib.parse
|
||||||
|
import urllib.request
|
||||||
|
|
||||||
|
BFF = os.environ["BFF"] # http://<bff-ip>:8080
|
||||||
|
TEMPO = os.environ["TEMPO"] # http://<tempo-ip>:3200
|
||||||
|
TIMEOUT = int(os.environ.get("TRACING_TIMEOUT", "90"))
|
||||||
|
WANT = {"bff", "projection-api"} # the two services that must share one trace
|
||||||
|
|
||||||
|
|
||||||
|
def _get(url):
|
||||||
|
with urllib.request.urlopen(url, timeout=10) as r:
|
||||||
|
return r.read()
|
||||||
|
|
||||||
|
|
||||||
|
def generate_traffic():
|
||||||
|
# A non-2xx still produces spans; only total unreachability of the BFF is fatal.
|
||||||
|
for _ in range(3):
|
||||||
|
try:
|
||||||
|
_get(f"{BFF}/openbaar/register")
|
||||||
|
except urllib.error.HTTPError:
|
||||||
|
pass
|
||||||
|
|
||||||
|
|
||||||
|
def search_trace_ids():
|
||||||
|
q = urllib.parse.quote('{ resource.service.name = "bff" }')
|
||||||
|
try:
|
||||||
|
data = json.loads(_get(f"{TEMPO}/api/search?q={q}&limit=50"))
|
||||||
|
except Exception:
|
||||||
|
return []
|
||||||
|
return [t["traceID"] for t in data.get("traces", [])]
|
||||||
|
|
||||||
|
|
||||||
|
def services_in_trace(trace_id):
|
||||||
|
try:
|
||||||
|
data = json.loads(_get(f"{TEMPO}/api/traces/{trace_id}"))
|
||||||
|
except Exception:
|
||||||
|
return set()
|
||||||
|
names = set()
|
||||||
|
for batch in data.get("batches", []):
|
||||||
|
for attr in batch.get("resource", {}).get("attributes", []):
|
||||||
|
if attr.get("key") == "service.name":
|
||||||
|
names.add(attr.get("value", {}).get("stringValue"))
|
||||||
|
return names
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
deadline = time.time() + TIMEOUT
|
||||||
|
generate_traffic()
|
||||||
|
seen = set()
|
||||||
|
while time.time() < deadline:
|
||||||
|
for tid in search_trace_ids():
|
||||||
|
names = services_in_trace(tid)
|
||||||
|
seen |= names
|
||||||
|
if WANT.issubset(names):
|
||||||
|
print(f"OK — trace {tid} spans {sorted(names)}")
|
||||||
|
return 0
|
||||||
|
time.sleep(3)
|
||||||
|
generate_traffic()
|
||||||
|
print(f"FAIL — no single trace spanned {sorted(WANT)}; services seen: {sorted(seen)}",
|
||||||
|
file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main())
|
||||||
@@ -5,6 +5,14 @@
|
|||||||
<ProjectReference Include="..\Acl.Infrastructure\Acl.Infrastructure.csproj" />
|
<ProjectReference Include="..\Acl.Infrastructure\Acl.Infrastructure.csproj" />
|
||||||
</ItemGroup>
|
</ItemGroup>
|
||||||
|
|
||||||
|
<ItemGroup>
|
||||||
|
<PackageReference Include="OpenTelemetry.Exporter.OpenTelemetryProtocol" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Exporter.Prometheus.AspNetCore" Version="1.17.0-beta.1" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Extensions.Hosting" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Instrumentation.AspNetCore" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Instrumentation.Http" Version="1.17.0" />
|
||||||
|
</ItemGroup>
|
||||||
|
|
||||||
<PropertyGroup>
|
<PropertyGroup>
|
||||||
<TargetFramework>net10.0</TargetFramework>
|
<TargetFramework>net10.0</TargetFramework>
|
||||||
<Nullable>enable</Nullable>
|
<Nullable>enable</Nullable>
|
||||||
|
|||||||
@@ -1,8 +1,31 @@
|
|||||||
using Acl.Application;
|
using Acl.Application;
|
||||||
using Acl.Infrastructure;
|
using Acl.Infrastructure;
|
||||||
|
using OpenTelemetry.Metrics;
|
||||||
|
using OpenTelemetry.Resources;
|
||||||
|
using OpenTelemetry.Trace;
|
||||||
|
|
||||||
var builder = WebApplication.CreateBuilder(args);
|
var builder = WebApplication.CreateBuilder(args);
|
||||||
|
|
||||||
|
// OpenTelemetry tracing (S-16b, ADR-0023): auto-instrument incoming ASP.NET Core requests and
|
||||||
|
// outgoing HttpClient calls (the ACL → OpenZaak hop), exported over OTLP to Tempo. Service name +
|
||||||
|
// OTLP endpoint come from OTEL_* env (compose); the exporter no-ops when Tempo is unreachable.
|
||||||
|
builder.Services.AddOpenTelemetry()
|
||||||
|
.ConfigureResource(r => r.AddService(
|
||||||
|
builder.Configuration["OTEL_SERVICE_NAME"] ?? builder.Environment.ApplicationName))
|
||||||
|
.WithTracing(tracing => tracing
|
||||||
|
.AddAspNetCoreInstrumentation(o => o.Filter = ctx => ctx.Request.Path != "/health")
|
||||||
|
.AddHttpClientInstrumentation()
|
||||||
|
.AddOtlpExporter())
|
||||||
|
// OpenTelemetry metrics (S-16c, ADR-0023): golden signals for the request path —
|
||||||
|
// http.server.request.duration (traffic/errors/latency) + http.client.* for downstream hops, plus
|
||||||
|
// the built-in System.Runtime meter for saturation (GC, CPU, thread pool). Prometheus scrapes these
|
||||||
|
// from /metrics (mapped below); metrics aren't pushed over OTLP, so no collector hop (ADR-0023).
|
||||||
|
.WithMetrics(metrics => metrics
|
||||||
|
.AddAspNetCoreInstrumentation()
|
||||||
|
.AddHttpClientInstrumentation()
|
||||||
|
.AddMeter("System.Runtime")
|
||||||
|
.AddPrometheusExporter());
|
||||||
|
|
||||||
builder.Services.AddSingleton<IClock, SystemClock>();
|
builder.Services.AddSingleton<IClock, SystemClock>();
|
||||||
builder.Services.AddSingleton(sp => sp.GetRequiredService<IConfiguration>()
|
builder.Services.AddSingleton(sp => sp.GetRequiredService<IConfiguration>()
|
||||||
.GetSection("Acl:Defaults").Get<AclDefaults>()
|
.GetSection("Acl:Defaults").Get<AclDefaults>()
|
||||||
@@ -19,6 +42,9 @@ var app = builder.Build();
|
|||||||
|
|
||||||
app.MapGet("/health", () => "Healthy");
|
app.MapGet("/health", () => "Healthy");
|
||||||
|
|
||||||
|
// Prometheus scrape endpoint (S-16c): exposes the OTel metrics above in Prometheus text format.
|
||||||
|
app.MapPrometheusScrapingEndpoint();
|
||||||
|
|
||||||
// The ACL's single operation, exposed as a service endpoint.
|
// The ACL's single operation, exposed as a service endpoint.
|
||||||
app.MapPost("/zaken", async (OpenZaakRequest body, AclService acl, CancellationToken ct) =>
|
app.MapPost("/zaken", async (OpenZaakRequest body, AclService acl, CancellationToken ct) =>
|
||||||
{
|
{
|
||||||
|
|||||||
@@ -10,6 +10,11 @@
|
|||||||
<!-- OIDC/JWT validation of Keycloak-issued tokens (ADR-0010) and OpenAPI generation. -->
|
<!-- OIDC/JWT validation of Keycloak-issued tokens (ADR-0010) and OpenAPI generation. -->
|
||||||
<PackageReference Include="Microsoft.AspNetCore.Authentication.JwtBearer" Version="10.0.8" />
|
<PackageReference Include="Microsoft.AspNetCore.Authentication.JwtBearer" Version="10.0.8" />
|
||||||
<PackageReference Include="Microsoft.AspNetCore.OpenApi" Version="10.0.8" />
|
<PackageReference Include="Microsoft.AspNetCore.OpenApi" Version="10.0.8" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Exporter.OpenTelemetryProtocol" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Exporter.Prometheus.AspNetCore" Version="1.17.0-beta.1" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Extensions.Hosting" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Instrumentation.AspNetCore" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Instrumentation.Http" Version="1.17.0" />
|
||||||
</ItemGroup>
|
</ItemGroup>
|
||||||
|
|
||||||
</Project>
|
</Project>
|
||||||
|
|||||||
@@ -3,9 +3,33 @@ using System.Text.Json;
|
|||||||
using System.Text.Json.Serialization;
|
using System.Text.Json.Serialization;
|
||||||
using Bff.Api;
|
using Bff.Api;
|
||||||
using Microsoft.AspNetCore.Authentication.JwtBearer;
|
using Microsoft.AspNetCore.Authentication.JwtBearer;
|
||||||
|
using OpenTelemetry.Metrics;
|
||||||
|
using OpenTelemetry.Resources;
|
||||||
|
using OpenTelemetry.Trace;
|
||||||
|
|
||||||
var builder = WebApplication.CreateBuilder(args);
|
var builder = WebApplication.CreateBuilder(args);
|
||||||
|
|
||||||
|
// OpenTelemetry tracing (S-16b, ADR-0023): auto-instrument incoming ASP.NET Core requests and
|
||||||
|
// outgoing HttpClient calls (BFF → Domain, BFF → projection-api), exported over OTLP to Tempo, so a
|
||||||
|
// portal request is one connected trace across the services. Service name + OTLP endpoint come from
|
||||||
|
// OTEL_* env (compose); the exporter no-ops when Tempo is unreachable. /health is filtered out.
|
||||||
|
builder.Services.AddOpenTelemetry()
|
||||||
|
.ConfigureResource(r => r.AddService(
|
||||||
|
builder.Configuration["OTEL_SERVICE_NAME"] ?? builder.Environment.ApplicationName))
|
||||||
|
.WithTracing(tracing => tracing
|
||||||
|
.AddAspNetCoreInstrumentation(o => o.Filter = ctx => ctx.Request.Path != "/health")
|
||||||
|
.AddHttpClientInstrumentation()
|
||||||
|
.AddOtlpExporter())
|
||||||
|
// OpenTelemetry metrics (S-16c, ADR-0023): the golden signals for the request path —
|
||||||
|
// http.server.request.duration (traffic/errors/latency) + http.client.* for the downstream hops,
|
||||||
|
// plus the built-in System.Runtime meter for saturation (GC, CPU, thread pool). Prometheus scrapes
|
||||||
|
// these from /metrics (mapped below); no OTLP push for metrics, so no collector hop (ADR-0023).
|
||||||
|
.WithMetrics(metrics => metrics
|
||||||
|
.AddAspNetCoreInstrumentation()
|
||||||
|
.AddHttpClientInstrumentation()
|
||||||
|
.AddMeter("System.Runtime")
|
||||||
|
.AddPrometheusExporter());
|
||||||
|
|
||||||
var keycloakAuthority = builder.Configuration["Keycloak:Authority"]
|
var keycloakAuthority = builder.Configuration["Keycloak:Authority"]
|
||||||
?? throw new InvalidOperationException("Missing configuration 'Keycloak:Authority'");
|
?? throw new InvalidOperationException("Missing configuration 'Keycloak:Authority'");
|
||||||
// Behandelaars authenticate against a *different* Keycloak realm (medewerker) than citizens (digid),
|
// Behandelaars authenticate against a *different* Keycloak realm (medewerker) than citizens (digid),
|
||||||
@@ -68,6 +92,9 @@ app.UseAuthentication();
|
|||||||
app.UseAuthorization();
|
app.UseAuthorization();
|
||||||
|
|
||||||
app.MapHealthChecks("/health");
|
app.MapHealthChecks("/health");
|
||||||
|
|
||||||
|
// Prometheus scrape endpoint (S-16c): exposes the OTel metrics above in Prometheus text format.
|
||||||
|
app.MapPrometheusScrapingEndpoint();
|
||||||
app.MapOpenApi();
|
app.MapOpenApi();
|
||||||
|
|
||||||
// Self-service submit: requires a valid digid token; the bsn comes from the token, not the body,
|
// Self-service submit: requires a valid digid token; the bsn comes from the token, not the body,
|
||||||
|
|||||||
@@ -0,0 +1,28 @@
|
|||||||
|
using System.Net;
|
||||||
|
using Microsoft.AspNetCore.Mvc.Testing;
|
||||||
|
|
||||||
|
namespace Bff.Tests;
|
||||||
|
|
||||||
|
/// <summary>
|
||||||
|
/// S-16c (#124): the service exposes OTel HTTP-server metrics in Prometheus text format at /metrics,
|
||||||
|
/// so Prometheus can scrape the golden signals (traffic, errors, latency) for the request path.
|
||||||
|
/// </summary>
|
||||||
|
public class MetricsEndpointTests(WebApplicationFactory<Program> factory)
|
||||||
|
: IClassFixture<WebApplicationFactory<Program>>
|
||||||
|
{
|
||||||
|
[Fact]
|
||||||
|
public async Task Metrics_endpoint_exposes_http_server_request_duration_after_traffic()
|
||||||
|
{
|
||||||
|
var client = factory.CreateClient();
|
||||||
|
|
||||||
|
// One request produces an http.server.request.duration measurement...
|
||||||
|
await client.GetAsync("/health");
|
||||||
|
|
||||||
|
// ...which the /metrics scrape endpoint then exposes in Prometheus text format.
|
||||||
|
var response = await client.GetAsync("/metrics");
|
||||||
|
|
||||||
|
Assert.Equal(HttpStatusCode.OK, response.StatusCode);
|
||||||
|
var body = await response.Content.ReadAsStringAsync();
|
||||||
|
Assert.Contains("http_server_request_duration", body);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -6,6 +6,11 @@
|
|||||||
</ItemGroup>
|
</ItemGroup>
|
||||||
|
|
||||||
<ItemGroup>
|
<ItemGroup>
|
||||||
|
<PackageReference Include="OpenTelemetry.Exporter.OpenTelemetryProtocol" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Exporter.Prometheus.AspNetCore" Version="1.17.0-beta.1" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Extensions.Hosting" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Instrumentation.AspNetCore" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Instrumentation.Http" Version="1.17.0" />
|
||||||
<PackageReference Include="Quartz.Extensions.Hosting" Version="3.18.2" />
|
<PackageReference Include="Quartz.Extensions.Hosting" Version="3.18.2" />
|
||||||
</ItemGroup>
|
</ItemGroup>
|
||||||
|
|
||||||
|
|||||||
@@ -1,10 +1,35 @@
|
|||||||
using Big.Application;
|
using Big.Application;
|
||||||
using Big.Domain;
|
using Big.Domain;
|
||||||
using Big.Infrastructure;
|
using Big.Infrastructure;
|
||||||
|
using OpenTelemetry.Metrics;
|
||||||
|
using OpenTelemetry.Resources;
|
||||||
|
using OpenTelemetry.Trace;
|
||||||
using Quartz;
|
using Quartz;
|
||||||
|
|
||||||
var builder = WebApplication.CreateBuilder(args);
|
var builder = WebApplication.CreateBuilder(args);
|
||||||
|
|
||||||
|
// OpenTelemetry tracing (S-16b, ADR-0023): auto-instrument incoming ASP.NET Core requests and
|
||||||
|
// outgoing HttpClient calls, exported over OTLP to Tempo, so a request is one connected trace across
|
||||||
|
// the services. Service name + OTLP endpoint come from OTEL_* env (compose); the exporter no-ops
|
||||||
|
// harmlessly when Tempo is unreachable (e.g. a service run standalone). /health is filtered out so
|
||||||
|
// liveness polls don't flood the traces.
|
||||||
|
builder.Services.AddOpenTelemetry()
|
||||||
|
.ConfigureResource(r => r.AddService(
|
||||||
|
builder.Configuration["OTEL_SERVICE_NAME"] ?? builder.Environment.ApplicationName))
|
||||||
|
.WithTracing(tracing => tracing
|
||||||
|
.AddAspNetCoreInstrumentation(o => o.Filter = ctx => ctx.Request.Path != "/health")
|
||||||
|
.AddHttpClientInstrumentation()
|
||||||
|
.AddOtlpExporter())
|
||||||
|
// OpenTelemetry metrics (S-16c, ADR-0023): golden signals for the request path —
|
||||||
|
// http.server.request.duration (traffic/errors/latency) + http.client.* for downstream hops, plus
|
||||||
|
// the built-in System.Runtime meter for saturation (GC, CPU, thread pool). Prometheus scrapes these
|
||||||
|
// from /metrics (mapped below); metrics aren't pushed over OTLP, so no collector hop (ADR-0023).
|
||||||
|
.WithMetrics(metrics => metrics
|
||||||
|
.AddAspNetCoreInstrumentation()
|
||||||
|
.AddHttpClientInstrumentation()
|
||||||
|
.AddMeter("System.Runtime")
|
||||||
|
.AddPrometheusExporter());
|
||||||
|
|
||||||
// Options bound from configuration (compose sets Flowable__* and Acl__* env vars).
|
// Options bound from configuration (compose sets Flowable__* and Acl__* env vars).
|
||||||
builder.Services.AddSingleton(sp => sp.GetRequiredService<IConfiguration>()
|
builder.Services.AddSingleton(sp => sp.GetRequiredService<IConfiguration>()
|
||||||
.GetSection("Flowable").Get<FlowableOptions>()
|
.GetSection("Flowable").Get<FlowableOptions>()
|
||||||
@@ -69,6 +94,9 @@ var app = builder.Build();
|
|||||||
|
|
||||||
app.MapGet("/health", () => "Healthy");
|
app.MapGet("/health", () => "Healthy");
|
||||||
|
|
||||||
|
// Prometheus scrape endpoint (S-16c): exposes the OTel metrics above in Prometheus text format.
|
||||||
|
app.MapPrometheusScrapingEndpoint();
|
||||||
|
|
||||||
// Submit a registration. The aggregate is created (INGEDIEND) and the registratie process started;
|
// Submit a registration. The aggregate is created (INGEDIEND) and the registratie process started;
|
||||||
// the zaak is opened later, off the request path, by the worker — so this returns 202 Accepted with
|
// the zaak is opened later, off the request path, by the worker — so this returns 202 Accepted with
|
||||||
// a location to read the registration's progress (ADR-0009, eventual consistency).
|
// a location to read the registration's progress (ADR-0009, eventual consistency).
|
||||||
|
|||||||
@@ -5,6 +5,14 @@
|
|||||||
<ProjectReference Include="..\..\projection-api\Projection.ReadModel\Projection.ReadModel.csproj" />
|
<ProjectReference Include="..\..\projection-api\Projection.ReadModel\Projection.ReadModel.csproj" />
|
||||||
</ItemGroup>
|
</ItemGroup>
|
||||||
|
|
||||||
|
<ItemGroup>
|
||||||
|
<PackageReference Include="OpenTelemetry.Exporter.OpenTelemetryProtocol" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Exporter.Prometheus.AspNetCore" Version="1.17.0-beta.1" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Extensions.Hosting" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Instrumentation.AspNetCore" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Instrumentation.Http" Version="1.17.0" />
|
||||||
|
</ItemGroup>
|
||||||
|
|
||||||
<PropertyGroup>
|
<PropertyGroup>
|
||||||
<TargetFramework>net10.0</TargetFramework>
|
<TargetFramework>net10.0</TargetFramework>
|
||||||
<Nullable>enable</Nullable>
|
<Nullable>enable</Nullable>
|
||||||
|
|||||||
@@ -1,9 +1,32 @@
|
|||||||
using System.Text.Json;
|
using System.Text.Json;
|
||||||
using EventSubscriber.Application;
|
using EventSubscriber.Application;
|
||||||
|
using OpenTelemetry.Metrics;
|
||||||
|
using OpenTelemetry.Resources;
|
||||||
|
using OpenTelemetry.Trace;
|
||||||
using Projection.ReadModel;
|
using Projection.ReadModel;
|
||||||
|
|
||||||
var builder = WebApplication.CreateBuilder(args);
|
var builder = WebApplication.CreateBuilder(args);
|
||||||
|
|
||||||
|
// OpenTelemetry tracing (S-16b, ADR-0023): auto-instrument the incoming NRC notification callback and
|
||||||
|
// the outgoing ACL enrichment call, exported over OTLP to Tempo. Service name + OTLP endpoint come
|
||||||
|
// from OTEL_* env (compose); the exporter no-ops when Tempo is unreachable.
|
||||||
|
builder.Services.AddOpenTelemetry()
|
||||||
|
.ConfigureResource(r => r.AddService(
|
||||||
|
builder.Configuration["OTEL_SERVICE_NAME"] ?? builder.Environment.ApplicationName))
|
||||||
|
.WithTracing(tracing => tracing
|
||||||
|
.AddAspNetCoreInstrumentation(o => o.Filter = ctx => ctx.Request.Path != "/health")
|
||||||
|
.AddHttpClientInstrumentation()
|
||||||
|
.AddOtlpExporter())
|
||||||
|
// OpenTelemetry metrics (S-16c, ADR-0023): golden signals for the request path —
|
||||||
|
// http.server.request.duration (traffic/errors/latency) + http.client.* for downstream hops, plus
|
||||||
|
// the built-in System.Runtime meter for saturation (GC, CPU, thread pool). Prometheus scrapes these
|
||||||
|
// from /metrics (mapped below); metrics aren't pushed over OTLP, so no collector hop (ADR-0023).
|
||||||
|
.WithMetrics(metrics => metrics
|
||||||
|
.AddAspNetCoreInstrumentation()
|
||||||
|
.AddHttpClientInstrumentation()
|
||||||
|
.AddMeter("System.Runtime")
|
||||||
|
.AddPrometheusExporter());
|
||||||
|
|
||||||
var connectionString = builder.Configuration.GetConnectionString("Projection")
|
var connectionString = builder.Configuration.GetConnectionString("Projection")
|
||||||
?? throw new InvalidOperationException("Missing connection string 'ConnectionStrings:Projection'");
|
?? throw new InvalidOperationException("Missing connection string 'ConnectionStrings:Projection'");
|
||||||
// The exact Authorization header value Open Notificaties sends on each abonnement callback.
|
// The exact Authorization header value Open Notificaties sends on each abonnement callback.
|
||||||
@@ -28,6 +51,9 @@ await app.Services.MigrateProjectionAsync();
|
|||||||
|
|
||||||
app.MapGet("/health", () => "Healthy");
|
app.MapGet("/health", () => "Healthy");
|
||||||
|
|
||||||
|
// Prometheus scrape endpoint (S-16c): exposes the OTel metrics above in Prometheus text format.
|
||||||
|
app.MapPrometheusScrapingEndpoint();
|
||||||
|
|
||||||
// The NRC abonnement callback. Open Notificaties POSTs a notification here; we project it.
|
// The NRC abonnement callback. Open Notificaties POSTs a notification here; we project it.
|
||||||
// Auth-on-callback is mandatory: the auth check runs *before* the body is read, so NRC's
|
// Auth-on-callback is mandatory: the auth check runs *before* the body is read, so NRC's
|
||||||
// registration probe (a POST without the configured Authorization, and without a valid
|
// registration probe (a POST without the configured Authorization, and without a valid
|
||||||
|
|||||||
@@ -1,8 +1,31 @@
|
|||||||
using Microsoft.EntityFrameworkCore;
|
using Microsoft.EntityFrameworkCore;
|
||||||
|
using OpenTelemetry.Metrics;
|
||||||
|
using OpenTelemetry.Resources;
|
||||||
|
using OpenTelemetry.Trace;
|
||||||
using Projection.ReadModel;
|
using Projection.ReadModel;
|
||||||
|
|
||||||
var builder = WebApplication.CreateBuilder(args);
|
var builder = WebApplication.CreateBuilder(args);
|
||||||
|
|
||||||
|
// OpenTelemetry tracing (S-16b, ADR-0023): auto-instrument incoming ASP.NET Core requests, exported
|
||||||
|
// over OTLP to Tempo, so a BFF → projection-api read is one connected trace. Service name + OTLP
|
||||||
|
// endpoint come from OTEL_* env (compose); the exporter no-ops when Tempo is unreachable.
|
||||||
|
builder.Services.AddOpenTelemetry()
|
||||||
|
.ConfigureResource(r => r.AddService(
|
||||||
|
builder.Configuration["OTEL_SERVICE_NAME"] ?? builder.Environment.ApplicationName))
|
||||||
|
.WithTracing(tracing => tracing
|
||||||
|
.AddAspNetCoreInstrumentation(o => o.Filter = ctx => ctx.Request.Path != "/health")
|
||||||
|
.AddHttpClientInstrumentation()
|
||||||
|
.AddOtlpExporter())
|
||||||
|
// OpenTelemetry metrics (S-16c, ADR-0023): golden signals for the request path —
|
||||||
|
// http.server.request.duration (traffic/errors/latency) + http.client.* for downstream hops, plus
|
||||||
|
// the built-in System.Runtime meter for saturation (GC, CPU, thread pool). Prometheus scrapes these
|
||||||
|
// from /metrics (mapped below); metrics aren't pushed over OTLP, so no collector hop (ADR-0023).
|
||||||
|
.WithMetrics(metrics => metrics
|
||||||
|
.AddAspNetCoreInstrumentation()
|
||||||
|
.AddHttpClientInstrumentation()
|
||||||
|
.AddMeter("System.Runtime")
|
||||||
|
.AddPrometheusExporter());
|
||||||
|
|
||||||
var connectionString = builder.Configuration.GetConnectionString("Projection")
|
var connectionString = builder.Configuration.GetConnectionString("Projection")
|
||||||
?? throw new InvalidOperationException("Missing connection string 'ConnectionStrings:Projection'");
|
?? throw new InvalidOperationException("Missing connection string 'ConnectionStrings:Projection'");
|
||||||
|
|
||||||
@@ -17,6 +40,9 @@ await app.Services.MigrateProjectionAsync();
|
|||||||
|
|
||||||
app.MapGet("/health", () => "Healthy");
|
app.MapGet("/health", () => "Healthy");
|
||||||
|
|
||||||
|
// Prometheus scrape endpoint (S-16c): exposes the OTel metrics above in Prometheus text format.
|
||||||
|
app.MapPrometheusScrapingEndpoint();
|
||||||
|
|
||||||
// The read side of the projection. Public-safe field filtering is tightened in S-09; for now
|
// The read side of the projection. Public-safe field filtering is tightened in S-09; for now
|
||||||
// the minimal projection only carries id + status (bsn/naam deferred — ADR-0008).
|
// the minimal projection only carries id + status (bsn/naam deferred — ADR-0008).
|
||||||
app.MapGet("/register", async (ProjectionDbContext db, CancellationToken ct) =>
|
app.MapGet("/register", async (ProjectionDbContext db, CancellationToken ct) =>
|
||||||
|
|||||||
@@ -4,6 +4,14 @@
|
|||||||
<ProjectReference Include="..\Projection.ReadModel\Projection.ReadModel.csproj" />
|
<ProjectReference Include="..\Projection.ReadModel\Projection.ReadModel.csproj" />
|
||||||
</ItemGroup>
|
</ItemGroup>
|
||||||
|
|
||||||
|
<ItemGroup>
|
||||||
|
<PackageReference Include="OpenTelemetry.Exporter.OpenTelemetryProtocol" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Exporter.Prometheus.AspNetCore" Version="1.17.0-beta.1" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Extensions.Hosting" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Instrumentation.AspNetCore" Version="1.17.0" />
|
||||||
|
<PackageReference Include="OpenTelemetry.Instrumentation.Http" Version="1.17.0" />
|
||||||
|
</ItemGroup>
|
||||||
|
|
||||||
<PropertyGroup>
|
<PropertyGroup>
|
||||||
<TargetFramework>net10.0</TargetFramework>
|
<TargetFramework>net10.0</TargetFramework>
|
||||||
<Nullable>enable</Nullable>
|
<Nullable>enable</Nullable>
|
||||||
|
|||||||
Reference in New Issue
Block a user