2 Commits
Author SHA1 Message Date
not 1dd8bd4e1b S-24/#25 · Helm chart + Kubernetes deployment, and Caddy for the portals (#166) (#167)
CI / lint (push) Successful in 1m17s
CI / build (push) Successful in 1m12s
CI / unit (push) Successful in 1m26s
CI / frontend (push) Successful in 2m58s
CI / mutation (push) Successful in 9m1s
CI / verify-stack (push) Successful in 8m53s
## What & why

Two changes, made and verified together on a real cluster.

**S-24 / #25 — a Helm chart for the platform.** One chart, `infra/helm/big-reference`,
whose `values.yaml` is a near-literal transcription of `infra/docker-compose.yml`, rendered
by three generic templates (Deployment, Job, Service) over a `workloads` map. Adding a
service is a values edit. `make k8s-lint` renders and schema-checks the whole stack without
a cluster. The issue asked for a *sketch*; this is deployed and verified end to end (see
below), which is more than it asked for — the part it asked for that is **not** here is the
production-posture write-up (HA, secrets, backup), see Known gaps.

**#166 — Caddy replaces nginx in the portals.** nginx resolves a variable `proxy_pass`
upstream itself, using only the `resolver` directive and never `/etc/resolv.conf`'s search
domains. That had cost two workarounds in one script: rewriting the resolver address for
rootless podman, and injecting a full FQDN so the bare `bff` name could resolve on
Kubernetes. Caddy dials per request through the system resolver, so `reverse_proxy
bff:8080` works on every engine unchanged; `apps/portal-nginx-resolver.sh` and the chart's
`BFF_HOST` env are deleted.

Closes #25
Closes #166

## Definition of Done

- [x] Linked Gitea issue (above).
- [x] Failing test committed before the implementation — twice: the Caddyfile contract test
      before the Caddyfiles, `make k8s-lint` before the chart.
- [x] Implementation makes the test pass.
- [x] Conventional Commits referencing the issues (`refs #25` / `refs #166`).
- [ ] CI green — awaiting the run on this PR (`make k8s-lint`, `dotnet format` and the new
      unit self-check pass locally; the compose e2e and mutation lanes are CI's).
- [ ] `docker compose up` from a fresh clone reaches green health checks within 3 minutes —
      the portal images were rebuilt and verified standalone, but a full `make up` run has
      not been done on this branch. Please confirm in review or let CI's smoke test speak.
- [x] Docs updated — `docs/runbooks/kubernetes-talos.md` (new), `frontend-decisions.md`,
      `demo-script.md`, and the docs that named nginx.
- [x] ADR added — ADR-0033 (chart) and ADR-0034 (Caddy).
- [ ] Demo note in `docs/demo-script.md` — not added: the deployment target is not a
      user-visible slice, and the Caddy swap is invisible to the demo script beyond the
      wording fix included here.

## How it was verified

Brought up from scratch on a single-node Talos v1.14.0 VM (6 vCPU / 10 GB, virtio disk)
under virt-manager: **29 pods ready and four bootstrap Jobs complete in under three
minutes, zero restarts**, using ~4.4 GB of the VM's 10 GB.

- Full Common Ground path: portal Caddy → BFF → domain → Flowable → ACL → OpenZaak +
  Objecten → NRC → event-subscriber → projection → public register (`INGEDIEND`, reference
  matching the submitted registration).
- Werkbak read with an MFA'd medewerker token → 200.
- The browser flow driven with Playwright against `http://localhost:30140`: secure context,
  `crypto.subtle` present, Keycloak form reached, login completed, **no console errors**.
- Routing checked against a stub BFF: SPA fallback serves deep links, each portal proxies
  its own groups, and a portal does *not* proxy a neighbour's group.

## Notes for reviewers

Three bugs this shook out, each fixed at the cause rather than the symptom:

1. **`command` vs `args`.** Compose's `command:` replaces the image CMD; Kubernetes'
   replaces the ENTRYPOINT. Transcribing one to the other broke every upstream image that
   relies on its entrypoint — postgres refused to run as root, Keycloak tried to exec
   `start-dev`. The chart now `fail`s at render time on `command`.
2. **Concurrent migrations.** Both `/setup_configuration.sh` and `/start.sh` run
   `manage.py migrate`; compose serialises them with `depends_on`, Kubernetes has no such
   edge, so the init Job and its web pod raced (`relation "zgw_consumers_service" already
   exists`). The four Django services now do both steps in order in the web pod — which
   also deletes four workloads.
3. **`emptyDir` databases are wiped by any pod-template change.** `make k8s-reseed` now
   also restarts `event-subscriber` and `projection-api`, which create the projection
   schema on start and otherwise keep writing to a schema-less database.

Known gaps / follow-ups:

- **Secrets.** `values.yaml` carries the dev credentials in plain text (`admin/admin`, the
  ZGW client secret, the two Objecten tokens) and the chart has no `Secret` objects. Fine
  for a laptop demo, and exactly what #25's "production posture" ADR should address — I
  suggest a follow-up issue rather than stretching this PR.
- **No CI gate for the chart yet.** `make k8s-lint` exists but is not wired into
  `.gitea/workflows/ci.yaml`, and nothing enforces that the chart and the compose file stay
  in step. Worth a small follow-up.
- **This is two slices in one PR.** They were built and verified together and the diff is
  entangled (the chart was written against Caddy from the start), so splitting now would
  mean re-creating an nginx-shaped chart to throw away. Happy to split if you'd rather.
- **Rebased onto #161** (merged as #165) rather than merged, to keep the history linear.
  One conflict, in the `unit:` target where both branches add a self-check line — resolved
  by keeping both. #161's `infra/host-browser.yml` arrived with
  `/usr/share/nginx/html/config.json` and is fixed to `/usr/share/caddy/` inside the
  `feat(portals)` commit, so no commit on this branch leaves that overlay pointing at a
  path the images no longer have.Reviewed-on: #167
2026-09-10 08:53:58 +00:00
not 8b206a005f S-26/#162 · Werkbak refreshes itself when a registration is ready for beoordeling (#164)
CI / build (push) Successful in 1m8s
CI / lint (push) Successful in 1m23s
CI / unit (push) Successful in 1m27s
CI / frontend (push) Successful in 3m8s
CI / mutation (push) Successful in 6m13s
CI / verify-stack (push) Successful in 10m12s
## What & why

The behandel werkbak now **refreshes itself** while it is open, so a registration that reaches
beoordeling after the behandelaar opened the page shows up on its own — no reload.

`interval(WERKBAK_REFRESH_MS)` (5 s) re-reads the existing BFF endpoint, scoped to the page with
`takeUntilDestroyed()`. A *background* read leaves the rows and states on screen alone until it has
an answer, so a tick never flashes the loading state over rows being read and one failed poll never
swaps the list for the error alert; a read that comes back also clears an earlier failure, so the
view recovers on its own rather than needing the very reload this slice removes.

No new endpoint, dependency or server-side state, and no service boundary moves — rxjs and
`GET /behandel/werkbak` are both already here. **ADR-0032** records why polling rather than a pushed
stream: nothing notifies the BFF either, so SSE/WebSockets would poll the domain *inside* the BFF for
the same freshness, plus connection lifecycle, nginx buffering and a stateful BFF. Proposal: #163.

Closes #162

## Definition of Done

- [x] Linked Gitea issue (above).
- [x] Failing test committed before the implementation.
- [x] Implementation makes the test pass; refactor commit if structure improved.
- [x] Conventional Commits referencing the issue (`refs #162`).
- [ ] CI green — all Gitea Actions jobs.
- [x] `docker compose up` from a fresh clone reaches green health checks within 3 minutes (unchanged; only the behandel bundle differs).
- [x] Docs updated if behaviour, contracts, or operations changed.
- [x] ADR added in `docs/architecture/` (ADR-0032).
- [x] Demo note in `docs/demo-script.md` (user-visible).

## Notes for reviewers

**The e2e is the real acceptance test, and it took two goes to make it one.** Simply dropping the
`staff.reload()` from the happy path proved nothing: the werkbak was visited *after* the documents
were supplied, so the row was already there at page load. The spec now logs the behandelaar in
**first**, asserts the row is not there yet, and only then has the citizen supply the documents that
route it to Beoordelen — so the row can only reach that already-open, never-reloaded page via the
refresh. Verified both ways against a live stack: with the interval stubbed out it fails at
`Goedkeuren <ref> … element(s) not found` after 30 s; with it, the behandel nginx logs the poll that
delivers the row. The page is foregrounded before the assertion because Chromium throttles timers in
a hidden tab.

**Ceiling (named in the ADR):** a fixed 5 s interval, per open page, that keeps polling in a
background tab; each tick costs one Flowable task query plus a store read per open task. Upgrade
path: publish task events from the domain, then swap the `interval` for a stream — the endpoint
contract and the rendering stay put. Gate on `document.visibilityState` first if request volume is
the concern.

**Two housekeeping notes, neither blocking:**
- #162 is on **no milestone** (DoD item 1). It is portal UX, so it fits neither *Data Governance*
  nor *Production Posture* cleanly — your call where it lands.
- The issue titles itself **S-26**, which already belongs to the self-service resume slice (#111,
  `BACKLOG.md`). Everything here references **#162**; worth renumbering the title if the S-ids are
  meant to stay unique. `BACKLOG.md` is untouched for the same reason (it mirrors the active
  milestone, and this slice is on none).Reviewed-on: #164
2026-09-04 09:34:14 +00:00