## What & why Two changes, made and verified together on a real cluster. **S-24 / #25 — a Helm chart for the platform.** One chart, `infra/helm/big-reference`, whose `values.yaml` is a near-literal transcription of `infra/docker-compose.yml`, rendered by three generic templates (Deployment, Job, Service) over a `workloads` map. Adding a service is a values edit. `make k8s-lint` renders and schema-checks the whole stack without a cluster. The issue asked for a *sketch*; this is deployed and verified end to end (see below), which is more than it asked for — the part it asked for that is **not** here is the production-posture write-up (HA, secrets, backup), see Known gaps. **#166 — Caddy replaces nginx in the portals.** nginx resolves a variable `proxy_pass` upstream itself, using only the `resolver` directive and never `/etc/resolv.conf`'s search domains. That had cost two workarounds in one script: rewriting the resolver address for rootless podman, and injecting a full FQDN so the bare `bff` name could resolve on Kubernetes. Caddy dials per request through the system resolver, so `reverse_proxy bff:8080` works on every engine unchanged; `apps/portal-nginx-resolver.sh` and the chart's `BFF_HOST` env are deleted. Closes #25 Closes #166 ## Definition of Done - [x] Linked Gitea issue (above). - [x] Failing test committed before the implementation — twice: the Caddyfile contract test before the Caddyfiles, `make k8s-lint` before the chart. - [x] Implementation makes the test pass. - [x] Conventional Commits referencing the issues (`refs #25` / `refs #166`). - [ ] CI green — awaiting the run on this PR (`make k8s-lint`, `dotnet format` and the new unit self-check pass locally; the compose e2e and mutation lanes are CI's). - [ ] `docker compose up` from a fresh clone reaches green health checks within 3 minutes — the portal images were rebuilt and verified standalone, but a full `make up` run has not been done on this branch. Please confirm in review or let CI's smoke test speak. - [x] Docs updated — `docs/runbooks/kubernetes-talos.md` (new), `frontend-decisions.md`, `demo-script.md`, and the docs that named nginx. - [x] ADR added — ADR-0033 (chart) and ADR-0034 (Caddy). - [ ] Demo note in `docs/demo-script.md` — not added: the deployment target is not a user-visible slice, and the Caddy swap is invisible to the demo script beyond the wording fix included here. ## How it was verified Brought up from scratch on a single-node Talos v1.14.0 VM (6 vCPU / 10 GB, virtio disk) under virt-manager: **29 pods ready and four bootstrap Jobs complete in under three minutes, zero restarts**, using ~4.4 GB of the VM's 10 GB. - Full Common Ground path: portal Caddy → BFF → domain → Flowable → ACL → OpenZaak + Objecten → NRC → event-subscriber → projection → public register (`INGEDIEND`, reference matching the submitted registration). - Werkbak read with an MFA'd medewerker token → 200. - The browser flow driven with Playwright against `http://localhost:30140`: secure context, `crypto.subtle` present, Keycloak form reached, login completed, **no console errors**. - Routing checked against a stub BFF: SPA fallback serves deep links, each portal proxies its own groups, and a portal does *not* proxy a neighbour's group. ## Notes for reviewers Three bugs this shook out, each fixed at the cause rather than the symptom: 1. **`command` vs `args`.** Compose's `command:` replaces the image CMD; Kubernetes' replaces the ENTRYPOINT. Transcribing one to the other broke every upstream image that relies on its entrypoint — postgres refused to run as root, Keycloak tried to exec `start-dev`. The chart now `fail`s at render time on `command`. 2. **Concurrent migrations.** Both `/setup_configuration.sh` and `/start.sh` run `manage.py migrate`; compose serialises them with `depends_on`, Kubernetes has no such edge, so the init Job and its web pod raced (`relation "zgw_consumers_service" already exists`). The four Django services now do both steps in order in the web pod — which also deletes four workloads. 3. **`emptyDir` databases are wiped by any pod-template change.** `make k8s-reseed` now also restarts `event-subscriber` and `projection-api`, which create the projection schema on start and otherwise keep writing to a schema-less database. Known gaps / follow-ups: - **Secrets.** `values.yaml` carries the dev credentials in plain text (`admin/admin`, the ZGW client secret, the two Objecten tokens) and the chart has no `Secret` objects. Fine for a laptop demo, and exactly what #25's "production posture" ADR should address — I suggest a follow-up issue rather than stretching this PR. - **No CI gate for the chart yet.** `make k8s-lint` exists but is not wired into `.gitea/workflows/ci.yaml`, and nothing enforces that the chart and the compose file stay in step. Worth a small follow-up. - **This is two slices in one PR.** They were built and verified together and the diff is entangled (the chart was written against Caddy from the start), so splitting now would mean re-creating an nginx-shaped chart to throw away. Happy to split if you'd rather. - **Rebased onto #161** (merged as #165) rather than merged, to keep the history linear. One conflict, in the `unit:` target where both branches add a self-check line — resolved by keeping both. #161's `infra/host-browser.yml` arrived with `/usr/share/nginx/html/config.json` and is fixed to `/usr/share/caddy/` inside the `feat(portals)` commit, so no commit on this branch leaves that overlay pointing at a path the images no longer have.Reviewed-on: #167
80 lines
3.9 KiB
Markdown
80 lines
3.9 KiB
Markdown
# ADR-0032: The werkbak refreshes itself by polling, not by a pushed stream
|
|
|
|
- **Status:** Accepted
|
|
- **Date:** 2026-09-04
|
|
- **Deciders:** Respellion engineering
|
|
- **Slice:** #162 (proposal #163). The issue titles it S-26; that id already belongs to
|
|
the self-service resume slice (#111), so #162 is the identifier that counts.
|
|
|
|
## Context
|
|
|
|
The werkbak (S-12) is a read of the open Flowable `Beoordelen` tasks: portal → BFF
|
|
`GET /behandel/werkbak` → domain `Werkbak` query → workflow engine, each task enriched
|
|
from its aggregate. A registration reaches `Beoordelen` **asynchronously**, only once the
|
|
citizen supplies its documents and the DMN routes it (S-10a) — so it appears in a werkbak
|
|
that is already open, and until now a behandelaar had to reload the page to see it.
|
|
|
|
Three forces shape the mechanism:
|
|
|
|
- **Nothing notifies anyone.** The trigger lives in Flowable. The domain does not publish
|
|
task events, and there is no bus between the domain and the BFF.
|
|
- **The BFF is stateless** and sits behind each portal's reverse proxy.
|
|
- **This is the repo's first live-updating view**, so the choice sets a precedent.
|
|
|
|
## Decision
|
|
|
|
**The werkbak page re-reads the existing BFF endpoint on a fixed interval
|
|
(`WERKBAK_REFRESH_MS`, 5 s) while it is open. No new endpoint, dependency or server-side
|
|
state.**
|
|
|
|
The refresh is a *background* read: it leaves the rows and the loading/failure states
|
|
untouched until it has an answer, so a tick never flashes a spinner over rows a
|
|
behandelaar is reading and a single failed poll never swaps the list for the error alert.
|
|
A read that comes back also clears an earlier failure, so the view recovers on its own —
|
|
the same reload this slice set out to remove would otherwise be needed to escape a
|
|
transient error. Only a foreground read (on open, after a decision) speaks for whether the
|
|
werkbak is readable at all.
|
|
|
|
### Why not SSE or WebSockets
|
|
|
|
Neither buys freshness here, because **nothing notifies the BFF either**:
|
|
|
|
- **SSE** (`text/event-stream`) would mean a new streaming endpoint whose handler polls the
|
|
domain and forwards diffs — the same latency, plus connection lifecycle, proxy
|
|
buffering, and auth on a long-lived connection.
|
|
- **WebSocket/SignalR** adds a dependency (CLAUDE.md §13) and makes the BFF stateful and
|
|
sticky-session-bound. A genuine push path would *also* need the domain to publish task
|
|
events. Warranted by high-frequency, bidirectional or fan-out-heavy traffic; the werkbak
|
|
is none of those.
|
|
|
|
Polling meets the acceptance ("a registration can be seen in the werkbak once it is ready
|
|
for review") in a handful of lines inside one component.
|
|
|
|
- ponytail ceiling: a fixed 5 s interval, per open page, that keeps polling in a
|
|
background tab. Each tick costs one Flowable task query plus a store read per open task.
|
|
- Upgrade path: publish task events from the domain, then swap the component's `interval`
|
|
for a stream. The endpoint contract and the component's rendering stay as they are;
|
|
gate on `document.visibilityState` first if request volume is the concern.
|
|
|
|
## Consequences
|
|
|
|
**Positive**
|
|
|
|
- The outcome is delivered with no new endpoint, dependency, or server-side state, and no
|
|
service boundary moves.
|
|
- Self-healing: a transient read failure no longer strands the view until a manual reload.
|
|
- The e2e got *simpler* — the happy path waits for the werkbak row without reloading the
|
|
page, which is itself the live-refresh assertion.
|
|
|
|
**Negative / costs**
|
|
|
|
- Staleness is bounded by one interval (≤5 s) rather than instant.
|
|
- One `GET /behandel/werkbak` per open werkbak per interval, including in hidden tabs.
|
|
- The precedent is polling; a future view with genuinely high-frequency updates will have
|
|
to revisit this (see the upgrade path above).
|
|
|
|
## Coupling rules touched (CLAUDE.md §8)
|
|
|
|
None. The poll reuses the existing portal → BFF → domain read path: §8.3 (portals talk
|
|
only to the BFF) and §8.2 (only the Workflow Client talks to Flowable) are unchanged.
|