Compare commits

..
Author SHA1 Message Date
not 681654233e ci: parallelise jobs at runner capacity >1, keep the two heavy jobs apart (closes #127)
CI / lint (pull_request) Successful in 2m5s
CI / build (pull_request) Successful in 1m33s
CI / unit (pull_request) Successful in 1m48s
CI / frontend (pull_request) Successful in 3m57s
CI / mutation (pull_request) Successful in 7m14s
CI / verify-stack (pull_request) Successful in 9m1s
The six jobs have no data dependencies, so they now schedule concurrently on the
capacity-2 runner. Guard against the two memory-heavy jobs (mutation + verify-stack)
co-scheduling and re-triggering the e2e OOM (#126): order verify-stack after mutation
via needs, with if: !cancelled() so it still runs when the ratchet fails. Add a
concurrency group so a new push supersedes the previous run instead of wasting a slot.

closes #127
2026-07-23 16:41:57 +02:00
not 88338396f6 feat(obs): distributed traces across the .NET services (S-16b, closes #123) (#126)
CI / verify-stack (push) Successful in 12m13s
CI / build (push) Successful in 1m50s
CI / lint (push) Successful in 1m58s
CI / unit (push) Successful in 2m8s
CI / frontend (push) Successful in 4m29s
CI / mutation (push) Successful in 11m51s
## What & why

S-16b, second of the S-16 split, on top of the #125 backplane. The five .NET services now emit OpenTelemetry traces so a request is **one connected trace** across them.

- Each host wires `AddOpenTelemetry().WithTracing(...)` with `AddAspNetCoreInstrumentation` (incoming) + `AddHttpClientInstrumentation` (outgoing) + `AddOtlpExporter` to **Tempo**.
- Because every cross-service call already goes through a typed `HttpClient` (§8 boundaries), the W3C `traceparent` propagates with no manual code — bff → domain → acl → openzaak and bff → projection-api stitch into a single trace.
- Service name + OTLP endpoint come from `OTEL_*` env set per app service in compose. `/health` is filtered out so liveness polls don't flood the traces.

No new ADR — ADR-0023 already records the stack + the two documented gaps (browser-side tracing is out of scope, so the trace begins at the BFF; the async Flowable-poll boundary is a separate trace).

Closes #123

## Definition of Done

- [x] Failing test committed first (`verify-tracing` fails with no instrumentation).
- [x] Implementation makes it pass — **validated locally end to end**: a real connected trace spanning `bff` + `projection-api` was found in Tempo (BFF→projection→db + Tempo subset, no OpenZaak/egress).
- [x] Conventional Commits referencing the issue (`refs #123`).
- [ ] CI green — awaiting Gitea Actions (verify-tracing added to verify-stack after verify-bff).
- [x] `docker compose up` health unaffected — services boot healthy even when Tempo is unreachable (exporter no-ops; verified).
- [x] Docs — demo-script + BACKLOG.
- [x] ADR — none needed (covered by ADR-0023).

## Notes for reviewers

- **Per-service wiring, no shared lib:** the block is duplicated across the five hosts by design — services don't share code across boundaries here (§8), same as the duplicated typed clients.
- **Packages:** OpenTelemetry.Extensions.Hosting / Instrumentation.AspNetCore / Instrumentation.Http / Exporter.OpenTelemetryProtocol, all 1.17.0, pinned per-csproj (no central props file).
- **The check** generates anonymous BFF→projection traffic (no auth, no OpenZaak), then queries Tempo (TraceQL search → fetch trace → assert both service.names present) from a python:3-slim container in-network — same idiom as run-projection-check.sh.
- **Next:** #124 (S-16c) adds `/metrics` + Prometheus scrape targets + golden-signal Grafana dashboards.

Reviewed-on: #126
2026-07-23 14:38:26 +00:00
+18 -4
View File
@@ -9,6 +9,12 @@ on:
permissions: permissions:
contents: read contents: read
# Supersede stale runs: a new push to the same branch/PR cancels the previous run, so the runner's
# concurrency slots aren't spent on commits nobody is waiting for (refs #127).
concurrency:
group: ci-${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
# Self-hosted runner — see docs/runbooks/ci.md for the runner setup. # Self-hosted runner — see docs/runbooks/ci.md for the runner setup.
# `uses:` are absolute, tag-pinned URLs (CLAUDE.md §8.7 / §15). # `uses:` are absolute, tag-pinned URLs (CLAUDE.md §8.7 / §15).
@@ -129,12 +135,20 @@ jobs:
path: services/bff/StrykerOutput/**/reports/mutation-report.html path: services/bff/StrykerOutput/**/reports/mutation-report.html
if-no-files-found: warn if-no-files-found: warn
# One stage for every check that needs the live stack. On the single self-hosted # One stage for every check that needs the live stack. Booting OpenZaak once (instead
# runner jobs run sequentially, so booting OpenZaak once (instead of once per job) # of once per job) is the cheapest layout (issue #58). No setup-dotnet: the ACL test runs
# is the cheapest layout (issue #58). No setup-dotnet: the ACL test runs in a built # in a built image and everything reaches services by container IP. Needs Docker + egress
# image and everything reaches services by container IP. Needs Docker + egress
# (base images, nuget, selectielijst.openzaak.nl). # (base images, nuget, selectielijst.openzaak.nl).
#
# `needs: [mutation]` is NOT a data dependency — it serialises the two memory-heavy jobs so
# they never co-schedule now the runner has capacity >1. A concurrent Stryker run + full-stack
# bring-up + Playwright browser on one host is what OOMs the e2e (commit d5e5fa2, #126). The
# light .NET/frontend jobs have no `needs`, so they still parallelise up to runner capacity.
# `if: !cancelled()` keeps verify-stack running even when the mutation ratchet fails (so we don't
# lose its signal) while still honouring run cancellation from the concurrency group above.
verify-stack: verify-stack:
needs: [mutation]
if: ${{ !cancelled() }}
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- uses: https://github.com/actions/checkout@v4 - uses: https://github.com/actions/checkout@v4