ADR-0011 through ADR-0034, six of the seven runbooks (OpenZaak, Open Notificaties, Keycloak, Flowable, Kubernetes-on-Talos, Gitea Actions gotchas) and synthetic-data.md are now in mkdocs.yml's nav, which makes `make lint` green again. Also fills in ADR-0033's `Slice:` header, which said "none yet" — #25 closed it.
10 KiB
ADR-0033: Kubernetes deployment is one values-driven Helm chart, not a chart per service
- Status: Accepted
- Date: 2026-09-04
- Deciders: Respellion engineering
- Slice: #25 (S-24) — raised directly as a deployment-target request and matched to that issue afterwards; see the "Process note" at the end
Context
The stack is defined once, in infra/docker-compose.yml: 30-odd containers made of six
upstream Common Ground modules (OpenZaak, Open Notificaties, Objecten, Objecttypen,
Keycloak, Flowable), their databases and workers, five .NET services, four portals, six
one-shot bootstrap containers, and an observability backplane (off by default here). Compose is the
CI-canonical stack: make verify and every verify-* script drive it.
We now also want the stack on Kubernetes — first target a single-node Talos VM on a laptop. Four properties of this particular stack shape the answer:
- The upstream images are used verbatim and read their configuration from a mounted
directory (
setup_configuration/data.yaml, Keycloak realm exports, BPMN/DMN). Compose streams those files into external volumes (infra/seed-config.sh) because bind mounts don't reach sibling containers on the CI runner. Kubernetes needs the same files as ConfigMaps — from somewhere. - Django's
URLValidatorrejects single-label hosts. Compose works around it by handing the ACL and the seeds a container IP (ADR-0009, ADR-0020, ADR-0029, and theobjecten.localnetwork alias). In Kubernetes a Service FQDN is already multi-label, so the workaround has a natural replacement — but the hosts have to line up exactly, since Objecten reflects the request Host into the URLs it publishes to NRC. - The OIDC issuer must be one string for both the browser and the BFF (ADR-0010).
infra/host-browser.ymlalready solved this for a host browser: pinKC_HOSTNAME, keep backchannel discovery in-cluster, and mount aconfig.jsonper portal. - Nothing here is highly available. One replica of everything, on one node.
Decision
One chart — infra/helm/big-reference — whose values.yaml is a near-literal
transcription of the compose file, rendered by three generic templates (Deployment, Job,
Service) over a workloads map. Adding a service is a values edit.
Consequences of that shape, each chosen deliberately:
- Config files are not copied into the chart.
infra/helm/seed-configmaps.shcreates the ConfigMaps from the files that already live in the repo — the Kubernetes sibling ofinfra/seed-config.sh. The chart therefore needsmake k8s-seedbeforehelm install, which is the same two-step dance compose already has. - Bootstrap one-shots become Jobs, with no ordering mechanism. Every one is idempotent
(ADR-0020); each waits for the TCP ports it needs via a busybox init container and
Kubernetes retries the rest.
make k8s-reseedre-runs them. - The four Django services apply their own
setup_configuration—args: [sh, -c, "/setup_configuration.sh && exec /start.sh"]— instead of getting a separate*-initJob like compose. Both of those image scripts runmanage.py migrate, and compose serialises them withdepends_on: service_completed_successfully; Kubernetes has no such edge, so a Job and its web pod migrate the same database concurrently and Django dies with "relation zgw_consumers_service already exists". Running the two steps in order inside the one container leaves exactly one migrator per database, and deletes four workloads. args, nevercommand. Compose'scommand:replaces the image's CMD; Kubernetes'command:replaces its ENTRYPOINT. Transcribing one to the other silently broke every upstream image that relies on its entrypoint — postgres ran as root and refused to start, Keycloak tried to execstart-devas a binary. The chart nowfails at render time if a workload setscommand, because the symptom (a crashloop three layers down) is nothing like the cause.- Published ports are NodePorts. No ingress controller, no LoadBalancer, no TLS. The
four portals are the exception in use, not in wiring: PKCE needs
crypto.subtle, which browsers expose only in a secure context, so a portal has to be reached overlocalhost(make k8s-portalsforwards them) or eventually over HTTPS..Values.hostis therefore "the address the browser uses", not "the node's address" — it pins Keycloak's issuer and each portal'sconfig.json, and both must agree with the URL bar (ADR-0010). - Databases are
emptyDirby default, so the stack comes up on a cluster with no CSI driver; settingpersistence.storageClassswitches every database to a PVC. - Only two hosts become FQDNs — OpenZaak (for the ACL and the zaaktype seed) and
Objecten (for the ACL's register writes), the two that Django validates as URLs.
Everything else keeps the short compose service name, because the upstream
setup_configurationfiles name those and Objecten matches an objecttype URL against the one it was configured with. The portals used to be a third case — nginx'sresolvernever appends search domains, so the barebffupstream could not resolve on Kubernetes — which ADR-0034 removed by serving them with Caddy, whose resolver honours/etc/resolv.conf. - Compose stays CI-canonical. The chart is a second deployment target, not a replacement; the acceptance, verify and e2e lanes are unchanged.
Alternatives considered
-
A chart per service, or an umbrella of 30 subcharts. The conventional layout, and roughly 1,500 lines of near-identical YAML for a stack where 28 of 30 workloads are "one pod, one image, some env". It buys independent versioning we don't want (the stack is demoed as a whole) and costs the eye-diffability against the compose file that keeps the two stacks honest.
-
kompose convert. One-shot generation, no ongoing artefact to maintain — but it drops exactly the parts that carry the design (init ordering, the config volumes, the issuer pinning) and produces output nobody owns. -
Bitnami PostgreSQL/Redis subcharts. Six more dependencies (CLAUDE.md §13) and a second way of expressing the same three-line database.
-
ingress-nginx with hostname routing. Needs a controller,
/etc/hostsentries and a matching issuer host; NodePorts need none of it and reuse the mechanisminfra/host-browser.ymlalready proves. -
A registry on the laptop (the obvious home for images built there). Talos cannot side-load an image, so a registry is required either way — but reaching one on the host means opening an inbound port on firewalld's
libvirtzone, which needs root, and pushing to it over plain HTTP means aninsecure-registriesentry in the Docker daemon, which needs root again.infra/helm/registry.yamlruns the registry in the cluster on a NodePort instead: pushing laptop → node is outbound and unfiltered, the node pulls from its own NodePort, anddocker save | crane push --insecureneeds no daemon configuration. Cost: one more (throwaway,emptyDir) workload, and a re-push if its pod is replaced. -
Helm hooks (
pre-install/post-install) for bootstrap ordering. Hooks run after--wait, which would deadlock: OpenZaak's readiness needs the migrations that the hook is supposed to run. Idempotent Jobs plus retries need no such sequencing. -
ponytail ceiling: single-node assumptions are baked in — one replica per workload,
Recreaterollouts, ReadWriteOnce volumes, no PodDisruptionBudgets, no resource requests or limits (a laptop VM schedules everything or nothing), plain HTTP. Upgrade path for a real cluster: add requests/limits per workload (the field is already passed through), swap NodePorts for an Ingress with TLS, and give the databases a real StorageClass — none of which changes the workload graph.
Consequences
Positive
- One file to read to see what the cluster runs, and it lines up with the compose file line for line.
- The compose IP workarounds disappear: cluster DNS supplies multi-label hosts.
make k8s-lintrenders and schema-checks the whole stack without a cluster.- The config inputs have exactly one home (the repo) for both stacks — no fork to drift.
Negative / costs
- A second deployment description to keep in step with compose. Nothing enforces that today; a drift check belongs in CI (follow-up).
helm installalone is not enough — the ConfigMaps must be seeded first, and a missing one surfaces asContainerCreating, not as a clear error.- Generic templates mean a values typo can render valid-but-wrong YAML;
k8s-lintcatches schema errors, not intent. - The verify/e2e lanes do not run against the chart, so the Kubernetes path is verified by hand (docs/runbooks/kubernetes-talos.md §5) rather than by CI.
- The chart deviates from compose in four places now (args, self-configuring Django pods, FQDN hosts, NodePorts). Each is forced by the platform and commented where it appears, but it is four more things that can drift.
Coupling rules touched (CLAUDE.md §8)
None. The chart deploys the same graph: portals reach only the BFF (§8.3), only the ACL holds ZGW credentials (§8.1), only the Workflow Client talks to Flowable (§8.2), each service keeps its own database (§8.5). No workload gained a peer it didn't have in compose.
Verified
Brought up from scratch on a single-node Talos v1.14.0 VM (6 vCPU / 10 GB, virtio disk)
under virt-manager: 29 pods ready and four bootstrap Jobs complete in under three minutes,
with zero restarts, using ~4.4 GB of the VM's 10 GB. The smoke test in the runbook's §5
walks the whole path — portal proxy → BFF → domain → Flowable → ACL → OpenZaak + Objecten →
NRC → event-subscriber → projection → public register — plus a werkbak read with an
MFA'd medewerker token. The browser flow itself was driven with Playwright against
http://localhost:30140: secure context, PKCE, Keycloak form, login, no console errors.
Process note
CLAUDE.md §14 wants the ADR proposal issue opened before the code, and §7 wants a slice issue behind the work. This landed the other way round — chart first, on request. The issue and the CI drift check are the outstanding follow-ups.