Commit Graph
6 Commits
Author SHA1 Message Date
notandClaude Opus 5 e8cb1ec7e9 docs(k8s): how to reach the deployed portals from a laptop (refs #175)
The five forwards are not optional and not independent: the portals' OIDC
authority is pinned to localhost:30180, so forwarding the portal without
Keycloak gets ERR_CONNECTION_REFUSED on the discovery document and an opaque
"[object Object]" in the console. One ssh replaces `make k8s-portals` for the
lab server, and needs no kubeconfig.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 16:19:02 +02:00
notandClaude Opus 5 de6db7d35b docs(k8s): split the taint fix from the registry patch (refs #175)
Talos 1.14 rejects a `patch mc` that sets `cluster.allowSchedulingOnControlPlanes`
— the field left the v1alpha1 schema, the way `machine.install` did — and the
rejection discards the rest of the patch with it, so the registry mirror never
lands and the failure moves from Pending pods to ImagePullBackOff without ever
saying so. Document the two as separate steps, and correct the claim that the
taint has to be patched away: `kubectl taint` is what §1 prescribes and it holds
until the node re-registers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 16:06:24 +02:00
notandClaude Opus 5 d17e79959b ci(deploy): say why a deploy stalled, and document the cluster prerequisite (refs #175)
The first run against the lab server's VM timed out on `kubectl rollout status`
for the in-cluster registry with nothing but "timed out waiting for the
condition". The cause was three commands up the runbook: that VM was installed
from a stock Talos config, so the only node still carries the control-plane
taint and no pod can schedule — and the missing registry mirror would have
failed the image pulls right after.

`rollout status` can only ever report the symptom, so dump the whole cluster's
pods and the recent events on failure instead of `big`'s pods alone; the
scheduler's "untolerated taint" message is the answer and it lives in the
events. Runbook §9 now opens with the machine-config patch the deploy assumes,
as one applied-live patch rather than a `kubectl taint` that the controller
undoes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 16:01:47 +02:00
notandClaude Opus 5 fc036d53d5 ci(deploy): deploy the stack to Talos on merge to main (refs #175)
CI / k8s (pull_request) Successful in 5s
CI / lint (pull_request) Successful in 1m25s
CI / build (pull_request) Successful in 1m22s
CI / unit (pull_request) Successful in 1m24s
CI / frontend (pull_request) Successful in 1m45s
CI / mutation (pull_request) Successful in 3m1s
CI / verify-stack (pull_request) Successful in 6m22s
The chart has been deployable by hand since #25 and linted in CI since #168;
this makes a merged PR actually ship it to the lab server's Talos VM.

Neither the Kubernetes API nor the in-cluster registry is publicly reachable,
so the job forwards 6443, 30500 and 30141 over the same SSH hop into the Fedora
host that the Gitea-runner pipeline uses. That splits the registry into two
names for one store: images are pushed through the tunnel to localhost:30500,
and the node pulls them from its own NodePort — the address its registry-mirror
patch trusts over plain HTTP.

It deploys with `make k8s-reseed` rather than `make k8s-up`: the bootstrap Jobs
are idempotent, and deleting them first is what stops a changed Job template
from wedging `helm upgrade`. The nine deployments are then rolled explicitly,
because `dev` is a mutable tag and helm sees an unchanged pod template.

PR CI is the merge gate, so this workflow does not re-run the checks. Deploys
queue instead of cancelling: a `helm upgrade` killed half-way leaves the release
in `pending-upgrade` and needs unwedging by hand.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 15:40:42 +02:00
not 9d7e8e5b65 ci(k8s): gate the Helm chart in CI + a compose↔chart drift check (closes #168) (#171)
CI / k8s (push) Successful in 5s
CI / lint (push) Successful in 1m27s
CI / build (push) Successful in 1m22s
CI / unit (push) Successful in 1m12s
CI / frontend (push) Successful in 2m7s
CI / mutation (push) Successful in 3m9s
CI / verify-stack (push) Successful in 6m20s
## What & why

The Helm chart landed in #167 with two gaps written into ADR-0033: `make k8s-lint` existed
but no CI job ran it, and *"a second deployment description to keep in step with compose —
nothing enforces that today; a drift check belongs in CI (follow-up)"*. Both are closed here.

**`make k8s-drift`** (`infra/helm/check-drift.py`, stdlib only) compares what each stack
actually deploys rather than diffing two files that differ by design: workload names and
resolved container images, taken from `docker compose config --format json` and a rendered
chart. The six differences that exist today are declared in `DEVIATIONS` with the reason
each was forced — the four `*-init` Django services folded into their web pods, and the two
bootstrap Jobs compose runs from the host — so only a *new* difference fails.

**A `k8s` CI job** runs `k8s-lint` then `k8s-drift` on every push and PR. No cluster, no
marketplace action: helm is fetched as the pinned static binary the Talos runbook already
gives developers.

Closes #168

## Definition of Done

- [x] Linked Gitea issue (above).
- [x] Failing test committed before the implementation — the red commit reports all six
      real differences; the green commit declares them.
- [x] Implementation makes the test pass.
- [x] Conventional Commits referencing the issue (`refs #168`).
- [x] CI green — awaiting the run on this PR (`make k8s-lint` and `make k8s-drift` pass locally).
- [x] `docker compose up` unaffected — no service, image or compose file is touched.
- [x] Docs updated — `docs/runbooks/ci.md` (job table + the one place local and CI now
      differ), `docs/runbooks/kubernetes-talos.md` §7/§"not ported", and ADR-0033's cost note.
- [x] No ADR needed: no new dependency (python stdlib, and helm/docker were already
      prerequisites of the `k8s-*` targets), no boundary moved, no §8 rule bent.
- [x] Not user-visible, so no demo note.

## Notes for reviewers

Verified by hand that both drift classes fail the check, not just that it passes today:

- bumping `OPENZAAK_TAG` in compose alone → reports `openzaak` and `oz-celery` with both
  image strings;
- adding a workload to `values.yaml` alone → reports it by name.

Deliberate limits (there is a `ponytail:` note in the script):

- **Names and images only**, as sets — no per-workload env, ports or volumes. Those differ
  by design in four documented places, so comparing them would mean re-encoding every
  deviation field by field for very little more signal.
- **The three observability workloads are rendered with `enabled=true`** by the check, even
  though both stacks default them off, so their images can't drift unwatched.
- **`k8s-lint`/`k8s-drift` are not in `make ci`**, to avoid making `helm` a hard
  prerequisite for everyone. That is now the only local/CI difference; it's called out in
  `docs/runbooks/ci.md`.

Follow-ups filed while reviewing the chart, not addressed here: #169 (the published docs
omit every ADR after 0010 and all runbooks but `ci.md`) and #170 (the production-posture
ADR #25 asked for — secrets are still plain text in `values.yaml`).Reviewed-on: #171
2026-09-18 13:25:24 +00:00
not 1dd8bd4e1b S-24/#25 · Helm chart + Kubernetes deployment, and Caddy for the portals (#166) (#167)
CI / lint (push) Successful in 1m17s
CI / build (push) Successful in 1m12s
CI / unit (push) Successful in 1m26s
CI / frontend (push) Successful in 2m58s
CI / mutation (push) Successful in 9m1s
CI / verify-stack (push) Successful in 8m53s
## What & why

Two changes, made and verified together on a real cluster.

**S-24 / #25 — a Helm chart for the platform.** One chart, `infra/helm/big-reference`,
whose `values.yaml` is a near-literal transcription of `infra/docker-compose.yml`, rendered
by three generic templates (Deployment, Job, Service) over a `workloads` map. Adding a
service is a values edit. `make k8s-lint` renders and schema-checks the whole stack without
a cluster. The issue asked for a *sketch*; this is deployed and verified end to end (see
below), which is more than it asked for — the part it asked for that is **not** here is the
production-posture write-up (HA, secrets, backup), see Known gaps.

**#166 — Caddy replaces nginx in the portals.** nginx resolves a variable `proxy_pass`
upstream itself, using only the `resolver` directive and never `/etc/resolv.conf`'s search
domains. That had cost two workarounds in one script: rewriting the resolver address for
rootless podman, and injecting a full FQDN so the bare `bff` name could resolve on
Kubernetes. Caddy dials per request through the system resolver, so `reverse_proxy
bff:8080` works on every engine unchanged; `apps/portal-nginx-resolver.sh` and the chart's
`BFF_HOST` env are deleted.

Closes #25
Closes #166

## Definition of Done

- [x] Linked Gitea issue (above).
- [x] Failing test committed before the implementation — twice: the Caddyfile contract test
      before the Caddyfiles, `make k8s-lint` before the chart.
- [x] Implementation makes the test pass.
- [x] Conventional Commits referencing the issues (`refs #25` / `refs #166`).
- [ ] CI green — awaiting the run on this PR (`make k8s-lint`, `dotnet format` and the new
      unit self-check pass locally; the compose e2e and mutation lanes are CI's).
- [ ] `docker compose up` from a fresh clone reaches green health checks within 3 minutes —
      the portal images were rebuilt and verified standalone, but a full `make up` run has
      not been done on this branch. Please confirm in review or let CI's smoke test speak.
- [x] Docs updated — `docs/runbooks/kubernetes-talos.md` (new), `frontend-decisions.md`,
      `demo-script.md`, and the docs that named nginx.
- [x] ADR added — ADR-0033 (chart) and ADR-0034 (Caddy).
- [ ] Demo note in `docs/demo-script.md` — not added: the deployment target is not a
      user-visible slice, and the Caddy swap is invisible to the demo script beyond the
      wording fix included here.

## How it was verified

Brought up from scratch on a single-node Talos v1.14.0 VM (6 vCPU / 10 GB, virtio disk)
under virt-manager: **29 pods ready and four bootstrap Jobs complete in under three
minutes, zero restarts**, using ~4.4 GB of the VM's 10 GB.

- Full Common Ground path: portal Caddy → BFF → domain → Flowable → ACL → OpenZaak +
  Objecten → NRC → event-subscriber → projection → public register (`INGEDIEND`, reference
  matching the submitted registration).
- Werkbak read with an MFA'd medewerker token → 200.
- The browser flow driven with Playwright against `http://localhost:30140`: secure context,
  `crypto.subtle` present, Keycloak form reached, login completed, **no console errors**.
- Routing checked against a stub BFF: SPA fallback serves deep links, each portal proxies
  its own groups, and a portal does *not* proxy a neighbour's group.

## Notes for reviewers

Three bugs this shook out, each fixed at the cause rather than the symptom:

1. **`command` vs `args`.** Compose's `command:` replaces the image CMD; Kubernetes'
   replaces the ENTRYPOINT. Transcribing one to the other broke every upstream image that
   relies on its entrypoint — postgres refused to run as root, Keycloak tried to exec
   `start-dev`. The chart now `fail`s at render time on `command`.
2. **Concurrent migrations.** Both `/setup_configuration.sh` and `/start.sh` run
   `manage.py migrate`; compose serialises them with `depends_on`, Kubernetes has no such
   edge, so the init Job and its web pod raced (`relation "zgw_consumers_service" already
   exists`). The four Django services now do both steps in order in the web pod — which
   also deletes four workloads.
3. **`emptyDir` databases are wiped by any pod-template change.** `make k8s-reseed` now
   also restarts `event-subscriber` and `projection-api`, which create the projection
   schema on start and otherwise keep writing to a schema-less database.

Known gaps / follow-ups:

- **Secrets.** `values.yaml` carries the dev credentials in plain text (`admin/admin`, the
  ZGW client secret, the two Objecten tokens) and the chart has no `Secret` objects. Fine
  for a laptop demo, and exactly what #25's "production posture" ADR should address — I
  suggest a follow-up issue rather than stretching this PR.
- **No CI gate for the chart yet.** `make k8s-lint` exists but is not wired into
  `.gitea/workflows/ci.yaml`, and nothing enforces that the chart and the compose file stay
  in step. Worth a small follow-up.
- **This is two slices in one PR.** They were built and verified together and the diff is
  entangled (the chart was written against Caddy from the start), so splitting now would
  mean re-creating an nginx-shaped chart to throw away. Happy to split if you'd rather.
- **Rebased onto #161** (merged as #165) rather than merged, to keep the history linear.
  One conflict, in the `unit:` target where both branches add a self-check line — resolved
  by keeping both. #161's `infra/host-browser.yml` arrived with
  `/usr/share/nginx/html/config.json` and is fixed to `/usr/share/caddy/` inside the
  `feat(portals)` commit, so no commit on this branch leaves that overlay pointing at a
  path the images no longer have.Reviewed-on: #167
2026-09-10 08:53:58 +00:00