Compare commits

..
6 Commits
Author SHA1 Message Date
not e6aaed7c8c docs(k8s): ADR-0033 + the Talos deployment runbook (refs #25)
CI / build (pull_request) Successful in 1m9s
CI / lint (pull_request) Successful in 1m27s
CI / unit (pull_request) Successful in 1m37s
CI / frontend (pull_request) Successful in 3m16s
CI / mutation (pull_request) Successful in 6m20s
CI / verify-stack (pull_request) Successful in 9m58s
ADR-0033 records why one values-driven chart rather than 30 subcharts, the four
platform-forced deviations from compose, and the alternatives (kompose, bitnami
subcharts, ingress-nginx, Helm hooks for ordering, a laptop-side registry).

The runbook is the walkthrough as actually performed on a single-node Talos v1.14
VM under virt-manager, including the parts that bite: virt-manager ejecting the
install ISO on first shutdown, Talos 1.14 moving the install disk into its own
config document, the control-plane taint, and why the portals must be reached
over localhost (crypto.subtle needs a secure context for PKCE).
2026-09-04 17:51:21 +02:00
not 7a5840149c feat(k8s): Helm chart for the whole stack on a single-node cluster (refs #25)
One chart whose values.yaml is a near-literal transcription of
infra/docker-compose.yml, rendered by three generic templates (Deployment, Job,
Service) over a `workloads` map — so the two stacks can be diffed by eye instead
of by archaeology, and adding a service is a values edit.

Platform-forced deviations, each commented where it appears:
- `args`, never `command`: compose replaces the image CMD, Kubernetes replaces the
  ENTRYPOINT. The chart fails to render on `command`, because the symptom (postgres
  refusing to run as root, Keycloak exec-ing `start-dev`) is nothing like the cause.
- The four Django services apply their own setup_configuration in the web pod
  rather than in a separate init Job: both scripts migrate, and without compose's
  depends_on they race the same database.
- OpenZaak and Objecten are addressed by service FQDN, because Django rejects a
  single-label host in a URL — the reason compose passes container IPs around.
- NodePorts, no ingress; databases are emptyDir until persistence.storageClass is
  set, so the stack comes up on a cluster with no CSI driver.

The upstream config inputs stay in the repo and become ConfigMaps via
infra/helm/seed-configmaps.sh — the Kubernetes sibling of infra/seed-config.sh —
so the compose stack and the chart cannot fork. infra/helm/registry.yaml runs an
in-cluster registry because Talos cannot side-load an image and a laptop-side one
needs a root-level firewall change.
2026-09-04 17:51:21 +02:00
not 916d671d49 test(k8s): gate the Helm chart with a render + schema check (refs #25)
`make k8s-lint` runs `helm lint` plus a full `helm template`, so a values typo or a
malformed resource is caught without a cluster — the only automated check the chart
can have while CI has no Kubernetes to deploy into.

Red: there is no chart to lint yet.
2026-09-04 17:51:21 +02:00
not 2d40c84e2c docs(portals): ADR-0034 — Caddy serves the portals (refs #166)
Records the decision, the directive-order footgun that shapes the Caddyfiles, and
the measured cost (the images grew 75.7 MB → 90.6 MB). Also updates the three
frontend-decisions entries and the two other docs that named nginx.
2026-09-04 17:51:21 +02:00
not 4edcf00267 feat(portals): serve each portal with Caddy instead of nginx (refs #166)
nginx resolves a variable `proxy_pass` upstream itself, using only the `resolver`
directive and never the search domains in /etc/resolv.conf. That cost two
workarounds in one script: rewriting the resolver address for rootless podman
(Docker's 127.0.0.11 is wrong there), and injecting a full FQDN so the bare `bff`
name could resolve on Kubernetes at all.

Caddy dials its upstream per request through the system resolver, which reads
nameserver *and* search domains, so `reverse_proxy bff:8080` resolves on every
engine with no per-engine configuration — and it still starts before the BFF
exists and picks up its restarts. Both workarounds are deleted with the script.

Routing uses mutually-exclusive `handle` blocks, not a bare `try_files`: Caddy
sorts rewrites *before* reverse_proxy, so a top-level SPA fallback would rewrite
every API path to /index.html before the proxy saw it.
2026-09-04 17:51:21 +02:00
not 51d99855d1 test(portals): assert each portal proxies only its own endpoint group (refs #166)
The four portal proxy configs are near-identical, so a copy-paste slip is cheap to
introduce and expensive to find: proxying another portal's endpoint group hands a
browser an endpoint its token is not for, and the failure surfaces as a 401 three
services away. Asserts each portal proxies exactly its own groups to the BFF and
keeps the SPA fallback for Angular's client-side routes.

Red: the Caddyfiles it reads do not exist yet.
2026-09-04 17:50:50 +02:00
10 changed files with 8 additions and 381 deletions
-21
View File
@@ -41,27 +41,6 @@ jobs:
nuget-${{ runner.os }}-
- run: make lint
# The Helm chart's only automated gate: it renders and schema-checks the whole
# stack, and checks it still describes the same stack as the compose file
# (ADR-0033). No cluster involved — see docs/runbooks/kubernetes-talos.md.
k8s:
runs-on: ubuntu-latest
steps:
- uses: https://github.com/actions/checkout@v4
# helm as its pinned static binary rather than a marketplace action: one URL,
# the same one the Talos runbook §0 gives a developer, and no third-party
# action to vet (CLAUDE.md §13). The drift check also needs `docker compose`,
# which the runner already has (see docs/runbooks/ci.md).
- name: Install helm
run: |
mkdir -p "$HOME/.local/bin"
curl -sSL https://get.helm.sh/helm-v3.16.4-linux-amd64.tar.gz \
| tar xz -O linux-amd64/helm > "$HOME/.local/bin/helm"
chmod +x "$HOME/.local/bin/helm"
echo "$HOME/.local/bin" >> "$GITHUB_PATH"
- run: make k8s-lint
- run: make k8s-drift
build:
runs-on: ubuntu-latest
steps:
-110
View File
@@ -1,110 +0,0 @@
name: Deploy to Talos
# A merge to main ships the stack to the Talos cluster on the lab server
# (docs/runbooks/kubernetes-talos.md §9). PR CI is the merge gate, so main is
# green by construction — this workflow only deploys.
on:
push:
branches: [main]
workflow_dispatch:
permissions:
contents: read
# Queue deploys, never cancel one: a helm upgrade killed half-way leaves the
# release in `pending-upgrade` and the next run has to be unwedged by hand.
concurrency:
group: deploy-talos
cancel-in-progress: false
jobs:
deploy:
runs-on: ubuntu-latest
env:
# The Talos VM as seen from the Fedora host (libvirt guest IP), and the
# address a browser uses to reach the cluster. `localhost` is deliberate:
# the portals' PKCE needs a secure context, so they are reached over
# `kubectl port-forward` — runbook §5. Override with repo variables.
TALOS_VM_IP: ${{ vars.TALOS_VM_IP }}
TALOS_HOST: ${{ vars.TALOS_HOST }}
steps:
- uses: https://github.com/actions/checkout@v4
# Pinned static binaries, the same URLs the Talos runbook §0 gives a
# developer and the same helm the `k8s` CI job uses — no action to vet.
- name: Install kubectl, helm and crane
run: |
set -euo pipefail
bin="$HOME/.local/bin"; mkdir -p "$bin"
curl -sSLo "$bin/kubectl" https://dl.k8s.io/release/v1.37.0/bin/linux/amd64/kubectl
curl -sSL https://get.helm.sh/helm-v3.16.4-linux-amd64.tar.gz | tar xz -O linux-amd64/helm > "$bin/helm"
curl -sSL https://github.com/google/go-containerregistry/releases/download/v0.20.2/go-containerregistry_Linux_x86_64.tar.gz | tar xz -O crane > "$bin/crane"
chmod +x "$bin"/{kubectl,helm,crane}
echo "$bin" >> "$GITHUB_PATH"
# The cluster's API and its registry are only reachable through the Fedora
# host, so forward both to the runner. 30141 is the openbaar portal, for
# the smoke at the end.
- name: Tunnel the Talos API + registry through the Fedora host
env:
SSH_KEY: ${{ secrets.TALOS_SSH_KEY }}
run: |
set -euo pipefail
: "${TALOS_VM_IP:=192.168.122.173}"
umask 077
printf '%s\n' "$SSH_KEY" > ~/.ssh_talos
ssh -i ~/.ssh_talos -o StrictHostKeyChecking=no -o IdentitiesOnly=yes \
-o ExitOnForwardFailure=yes -p 6667 -f -N \
-L 6443:$TALOS_VM_IP:6443 \
-L 30500:$TALOS_VM_IP:30500 \
-L 30141:$TALOS_VM_IP:30141 \
user@labs.respellion.tech
# The kubeconfig's server must be https://127.0.0.1:6443 — Talos puts
# 127.0.0.1 in the apiserver cert SANs, so TLS verification still holds
# through the tunnel.
- name: Write the kubeconfig
env:
KUBECONFIG_B64: ${{ secrets.TALOS_KUBECONFIG }}
run: |
set -euo pipefail
umask 077
base64 -d <<< "$KUBECONFIG_B64" > "$RUNNER_TEMP/kubeconfig"
echo "KUBECONFIG=$RUNNER_TEMP/kubeconfig" >> "$GITHUB_ENV"
kubectl --kubeconfig "$RUNNER_TEMP/kubeconfig" get nodes
# Idempotent; also makes a first deploy onto a bare cluster work. The
# registry's storage is an emptyDir, so a replaced pod loses the images —
# which the push in the next step puts back anyway.
- name: Ensure the in-cluster registry
run: make k8s-registry
# Push through the tunnel (localhost), pull from the node's own NodePort
# (the address in the Talos registry-mirror patch) — same registry, two
# names, so the two `make` calls get different K8S_REGISTRY values.
- name: Build and push the images
run: make k8s-images K8S_REGISTRY=localhost:30500
# k8s-reseed = seed configmaps + helm upgrade + re-run the bootstrap jobs.
# The jobs are idempotent, and deleting them first is what keeps a changed
# Job template from wedging the upgrade (`cannot patch … with kind Job`).
- name: Deploy the chart
run: make k8s-reseed TALOS_HOST=${TALOS_HOST:-localhost} K8S_REGISTRY=${TALOS_VM_IP:-192.168.122.173}:30500
# `dev` is a mutable tag and helm sees an unchanged pod template, so the
# new images only land on a restart (pullPolicy is already Always).
- name: Roll the services onto the new images
run: |
set -euo pipefail
svcs="acl domain bff event-subscriber projection-api self-service openbaar behandel beheer"
kubectl -n big rollout restart deploy $svcs
kubectl -n big rollout status --timeout=300s deploy $svcs
# Proves portal → Caddy → BFF → projection end to end. An empty register is
# a pass; a 502 or a timeout is not.
- name: Smoke the public register
run: curl -fsS --retry 10 --retry-delay 6 --retry-all-errors http://localhost:30141/openbaar/register
- name: Pods on failure
if: failure()
run: kubectl -n big get pods,jobs || true
+1 -11
View File
@@ -43,7 +43,7 @@ export DOCKER_HOST := unix://$(PODMAN_SOCK)
endif
endif
.PHONY: ci lint build unit mutation frontend integration verify verify-up verify-acl verify-nrc verify-projection verify-bff verify-domain verify-observability verify-tracing verify-metrics verify-objecttypen verify-objecten verify-registerrecord verify-objecten-notifications verify-notifications smoke up down local verify-local local-down changelog openzaak-up openzaak-smoke openzaak-seed openzaak-down stack-up stack-smoke stack-down keycloak-up keycloak-smoke keycloak-down flowable-up flowable-smoke flowable-down k8s-lint k8s-drift k8s-registry k8s-images k8s-seed k8s-up k8s-reseed k8s-portals k8s-down k8s-purge help
.PHONY: ci lint build unit mutation frontend integration verify verify-up verify-acl verify-nrc verify-projection verify-bff verify-domain verify-observability verify-tracing verify-metrics verify-objecttypen verify-objecten verify-registerrecord verify-objecten-notifications verify-notifications smoke up down local verify-local local-down changelog openzaak-up openzaak-smoke openzaak-seed openzaak-down stack-up stack-smoke stack-down keycloak-up keycloak-smoke keycloak-down flowable-up flowable-smoke flowable-down k8s-lint k8s-registry k8s-images k8s-seed k8s-up k8s-reseed k8s-portals k8s-down k8s-purge help
## ci: run the full pipeline — lint, build, unit, mutation, frontend, verify (mirrors Gitea Actions)
## `verify` is the live-stack stage (full stack up once → ACL + notification checks).
@@ -64,9 +64,6 @@ frontend:
## lint: verify formatting (no changes)
lint:
dotnet format $(SLN) --verify-no-changes
# Only pages in mkdocs.yml's nav are published, and mkdocs keeps a build green
# when one is missing — so the nav is checked here rather than not at all.
python3 infra/check-docs-nav.py
## build: release build
build:
@@ -353,13 +350,6 @@ k8s-lint:
helm lint $(K8S_CHART)
helm template big $(K8S_CHART) -n $(K8S_NS) --set images.registry=registry.invalid:5000 >/dev/null
## k8s-drift: fail if compose and the Helm chart describe different stacks
# Compose is CI-canonical (ADR-0033) and the chart is a transcription of it; this
# compares what each one deploys — workload names and resolved images. Needs
# `docker compose` and `helm`, no cluster.
k8s-drift:
python3 infra/helm/check-drift.py
## k8s-registry: deploy the in-cluster image registry (NodePort 30500)
k8s-registry:
kubectl apply -f infra/helm/registry.yaml
@@ -3,8 +3,8 @@
- **Status:** Accepted
- **Date:** 2026-09-04
- **Deciders:** Respellion engineering
- **Slice:** #25 (S-24) — raised directly as a deployment-target request and matched to
that issue afterwards; see the "Process note" at the end
- **Slice:** _(none yet — raised directly as a deployment-target request; see
"Process note" at the end)_
## Context
@@ -126,9 +126,8 @@ Consequences of that shape, each chosen deliberately:
**Negative / costs**
- A second deployment description to keep in step with compose. `make k8s-drift` (#168)
now enforces the part that bites — the workload set and the resolved images, with the
four deviations below declared — but not per-workload env, ports or volumes.
- A second deployment description to keep in step with compose. Nothing enforces that
today; a drift check belongs in CI (follow-up).
- `helm install` alone is not enough — the ConfigMaps must be seeded first, and a missing
one surfaces as `ContainerCreating`, not as a clear error.
- Generic templates mean a values typo can render valid-but-wrong YAML; `k8s-lint` catches
-2
View File
@@ -14,8 +14,6 @@ should teach.
In Dutch; the strategic framing lives in `Respellion/innovation-lab`.
- **[Working in Gitea](gitea-workflow.md)** — issues, milestones, branches, PRs.
- **[CI runbook](runbooks/ci.md)** — the pipeline and the `make ci` local gate.
- **[Kubernetes on Talos](runbooks/kubernetes-talos.md)** — the second deployment target:
one Helm chart, a single-node cluster, and the parts that bite (ADR-0033).
## Quickstart
+2 -10
View File
@@ -2,10 +2,8 @@
> **Status: active.** The workflow `.gitea/workflows/ci.yaml` runs on Gitea's
> hosted `ubuntu-latest` runner — no self-hosted runner required.
> **`make ci` is still the local gate** — it runs the same checks via the same
> `make` targets, with one exception: the `k8s` job's targets are not in `make ci`,
> because `helm` is optional for everyone not deploying to Kubernetes. Run
> `make k8s-lint k8s-drift` by hand after touching the chart or the compose file.
> **`make ci` is still the local gate** — it runs the exact same checks
> (the workflow calls the same `make` targets).
## The pipeline
@@ -18,8 +16,6 @@ and CI cannot drift:
| `lint` | `make lint``dotnet format … --verify-no-changes` | .NET 10 SDK |
| `build` | `make build``dotnet build … -c Release` | .NET 10 SDK |
| `unit` | `make unit``dotnet test … -c Release --filter "Category!=Integration"` | .NET 10 SDK |
| `frontend` | `make frontend` → Nx lint/test/build for the four portals | pnpm + Node |
| `k8s` | `make k8s-lint` (render + schema-check the Helm chart) → `make k8s-drift` (chart still describes the same stack as `infra/docker-compose.yml`) | pinned `helm` binary + `docker compose` |
| `mutation` | `make mutation``dotnet tool restore``dotnet stryker` (ACL); uploads the HTML report as an artifact | .NET 10 SDK |
| `verify-stack` | the single live-stack stage — steps: `make verify-up` (full stack up + health, the DoD smoke) → `make verify-acl` (ACL ↔ OpenZaak) → `make verify-nrc` (OpenZaak → NRC delivery) → `make down` | container engine + egress (base images, nuget, `selectielijst.openzaak.nl`) |
@@ -31,10 +27,6 @@ and CI cannot drift:
> services by **container IP** (the runner can't reach published ports — see
> [gitea-actions-gotchas.md §5/§6](gitea-actions-gotchas.md)).
A second workflow, `.gitea/workflows/deploy.yaml`, deploys the stack to the Talos
cluster on the lab server when a PR is merged to `main` — see
[kubernetes-talos.md §9](kubernetes-talos.md) for its secrets and the SSH tunnel it needs.
All `uses:` references are absolute, tag-pinned URLs (`https://github.com/actions/checkout@v4`,
`https://github.com/actions/setup-dotnet@v4`) per CLAUDE.md §8.7 and §15 — Gitea
Actions resolves them from GitHub.
+1 -45
View File
@@ -322,7 +322,6 @@ The PVCs carry `helm.sh/resource-policy: keep`, so `make k8s-down` leaves the da
```bash
make k8s-lint # render + schema-check the chart, no cluster needed
make k8s-drift # fail if compose and the chart describe different stacks
make k8s-portals # forward the portals + Keycloak to localhost (browser access)
make k8s-images K8S_REGISTRY=... # after changing a service or a portal
make k8s-up TALOS_HOST=... K8S_REGISTRY=...
@@ -360,47 +359,6 @@ immutable, so `helm upgrade` is rejected with `cannot patch "…" with kind Job`
| Pods `Evicted` / `OOMKilled` | the VM is too small (§0) |
| A Job shows `BackoffLimitExceeded` | read it: `kubectl -n big logs job/<name>` |
## 9. Deploying on merge to main
`.gitea/workflows/deploy.yaml` runs the §3–§4 steps against the **lab server's** Talos VM
every time a PR is squash-merged to `main` (and on demand via *Run workflow*). PR CI is the
merge gate, so the workflow deploys without re-running the checks.
The cluster's API and registry are not exposed publicly, so the job forwards them over the
same SSH hop the Gitea-runner pipeline uses:
```
ssh -p 6667 user@labs.respellion.tech -L 6443 -L 30500 -L 30141 → <TALOS_VM_IP>
```
Consequences worth knowing:
- Images are **pushed** to `localhost:30500` (the tunnel) and **pulled** by the node from
`<TALOS_VM_IP>:30500` (its own NodePort, the address in the Talos registry-mirror patch).
Same registry, two names — hence the two `K8S_REGISTRY` values in the workflow.
- It calls `make k8s-reseed`, not `make k8s-up`: the bootstrap Jobs are idempotent, and
deleting them first is what stops a changed Job template from wedging `helm upgrade` (§7).
- `dev` is a mutable tag, so a `rollout restart` of the nine repo deployments is what
actually puts the new images in the pods.
- Deploys **queue** (`cancel-in-progress: false`): a helm upgrade killed half-way leaves the
release in `pending-upgrade`, which has to be unwedged by hand.
Settings, all on the repository in Gitea:
| Kind | Name | What |
|---|---|---|
| Secret | `TALOS_SSH_KEY` | private key for `user@labs.respellion.tech` (the Fedora host) |
| Secret | `TALOS_KUBECONFIG` | base64 of the kubeconfig, **`server: https://127.0.0.1:6443`** — Talos puts `127.0.0.1` in the apiserver cert SANs, so TLS still verifies through the tunnel |
| Variable | `TALOS_VM_IP` | the VM's libvirt address (default `192.168.122.173`) |
| Variable | `TALOS_HOST` | the browser-facing host baked into Keycloak's issuer (default `localhost`, see §5) |
The last step smokes `GET /openbaar/register` through the openbaar portal, which exercises
portal → Caddy → BFF → projection. An empty register passes; a 502 does not.
Not covered: the portals still need `make k8s-portals` (or an SSH forward) to be usable in a
browser, because PKCE needs a secure context (§5). Giving the server a hostname + TLS is the
upgrade path.
## What is not ported
- **Observability** (Tempo, Prometheus, Grafana) is defined but disabled — those are built
@@ -408,7 +366,5 @@ upgrade path.
`K8S_SET='--set workloads.tempo.enabled=true --set workloads.prometheus.enabled=true --set workloads.grafana.enabled=true'`.
The .NET services still export OTLP; the exporter fails harmlessly when Tempo is absent.
- **The verify/e2e lanes.** `make verify*` and the Playwright e2e drive compose, not the
chart. The Kubernetes path is verified with §5's smoke test. CI's `k8s` job runs the two
clusterless checks (`k8s-lint`, `k8s-drift`) on every PR — a values typo or a compose
image bump that skipped the chart fails there, but nothing deploys the chart in CI.
chart. The Kubernetes path is verified with §5's smoke test.
- **Ingress, TLS, and resource requests.** See the ponytail ceiling in ADR-0033.
-30
View File
@@ -1,30 +0,0 @@
#!/usr/bin/env python3
"""Fail when a page under docs/ is missing from mkdocs.yml's nav.
docs/ is the source of truth (CLAUDE.md §12), but only the pages listed in the nav
are published — and mkdocs' own `omitted_files: warn` keeps a build green while
silently dropping them, which is how every ADR after 0010 and every runbook but
ci.md fell off the site.
ponytail: a substring test, not a YAML parse — a page's path either appears in
mkdocs.yml or it doesn't, and that needs no dependency.
"""
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
nav = (ROOT / "mkdocs.yml").read_text()
missing = sorted(
str(page.relative_to(ROOT / "docs"))
for page in (ROOT / "docs").rglob("*.md")
if str(page.relative_to(ROOT / "docs")) not in nav
)
if missing:
print(f"{len(missing)} page(s) under docs/ are not in mkdocs.yml's nav:")
print("\n".join(f" {m}" for m in missing))
sys.exit(1)
print("docs nav complete: every page under docs/ is published")
-116
View File
@@ -1,116 +0,0 @@
#!/usr/bin/env python3
"""Fail when the compose stack and the Helm chart stop describing the same stack.
`infra/docker-compose.yml` is CI-canonical; `infra/helm/big-reference` is a
transcription of it (ADR-0033), and until now nothing kept the two in step — an
upstream image bump or a new service applied to only one of them landed
unnoticed. This compares what each side actually *deploys*, not the two files:
the rendered chart against `docker compose config`. Both tools are already
prerequisites of the `k8s-*` make targets.
Run it with `make k8s-drift`. No cluster needed.
ponytail: names and images only, as sets — no per-workload env/ports/volumes.
Those differ by design in four documented places (ADR-0033), so comparing them
would mean re-encoding every deviation field by field; a tag bump and a missing
service are the drift that actually bites.
"""
import json
import re
import subprocess
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
COMPOSE = ROOT / "infra/docker-compose.yml"
CHART = ROOT / "infra/helm/big-reference"
# The busybox init container that every `waitFor` workload gets exists only in
# the chart (compose has `depends_on`). Rendering it under a sentinel makes it
# filterable without teaching the check what busybox is.
BUSYBOX = "drift-check-ignored-init-image"
# Differences that Kubernetes forces, not drift (ADR-0033). A name listed here is
# expected to be on exactly one side; anything else fails.
DEVIATIONS = {
# The four Django services apply their own setup_configuration in the web pod
# (`args: [sh, -c, "/setup_configuration.sh && exec /start.sh"]`) rather than in a
# separate init Job. Both that script and /start.sh run `manage.py migrate`, and
# Kubernetes has no `depends_on: service_completed_successfully` to serialise them,
# so the Job and its web pod migrated the same database concurrently.
"oz-init": "folded into the openzaak pod",
"nrc-init": "folded into the nrc-web pod",
"objecttypen-init": "folded into the objecttypen pod",
"objecten-init": "folded into the objecten pod",
# Compose seeds these from the host — the verify scripts `docker cp` the two
# scripts into a running container, and docker-compose.local.yml carries
# `local-seed` + `nrc-subscribe` for `make local`. A cluster has no host to seed
# from, so both became Jobs in the chart.
"seed-zaaktype": "compose seeds the catalogus from the host (infra/openzaak/seed_catalogus.py)",
"nrc-subscribe": "compose registers the abonnement from the host (infra/local/register-abonnement.py)",
}
# Workloads the observability backplane adds. Off by default in both stacks'
# defaults, so they are rendered on purpose here — otherwise their images drift
# unwatched.
OBSERVABILITY = ["tempo", "prometheus", "grafana"]
def compose_services() -> dict[str, str]:
"""Service name -> image, with ${TAG:-default} interpolation already applied."""
out = run(["docker", "compose", "-f", str(COMPOSE), "config", "--format", "json"])
return {name: svc.get("image", "") for name, svc in json.loads(out)["services"].items()}
def chart_workloads() -> dict[str, str]:
"""Workload name -> image, read back out of the rendered manifests."""
out = run(
["helm", "template", "big", str(CHART), "-n", "big", "--set", f"images.busybox={BUSYBOX}"]
+ [f"--set=workloads.{w}.enabled=true" for w in OBSERVABILITY]
)
workloads = {}
for doc in out.split("\n---"):
if not re.search(r"^kind: (Deployment|Job)$", doc, re.M):
continue
name = re.search(r"^ name: (\S+)$", doc, re.M)[1]
images = [i for i in re.findall(r"^\s+image: (\S+)$", doc, re.M) if i != BUSYBOX]
workloads[name] = images[0]
return workloads
def run(argv: list[str]) -> str:
proc = subprocess.run(argv, capture_output=True, text=True)
if proc.returncode != 0:
sys.exit(f"{argv[0]} failed:\n{proc.stderr}")
return proc.stdout
def main() -> int:
compose, chart = compose_services(), chart_workloads()
problems = []
for name in sorted(set(compose) - set(chart) - set(DEVIATIONS)):
problems.append(f" {name}: in docker-compose.yml, not in the chart")
for name in sorted(set(chart) - set(compose) - set(DEVIATIONS)):
problems.append(f" {name}: in the chart, not in docker-compose.yml")
for name in sorted(set(compose) & set(chart)):
if compose[name] != chart[name]:
problems.append(f" {name}: compose runs {compose[name]}, the chart runs {chart[name]}")
if problems:
print("compose and the Helm chart describe different stacks:\n" + "\n".join(problems))
print(
"\nPort the change to the other stack, or — if the difference is forced by\n"
"Kubernetes — declare it in DEVIATIONS in this file, with the reason."
)
return 1
print(f"no drift: {len(chart)} workloads, images identical on both stacks")
for name, why in sorted(DEVIATIONS.items()):
print(f" deviation (declared): {name}{why}")
return 0
if __name__ == "__main__":
sys.exit(main())
-31
View File
@@ -32,30 +32,6 @@ nav:
- "ADR-0008: Read projection store": architecture/adr-0008-read-projection-store.md
- "ADR-0009: External-task job worker": architecture/adr-0009-external-task-job-worker.md
- "ADR-0010: BFF OIDC validation": architecture/adr-0010-bff-oidc.md
- "ADR-0011: Approval status flow": architecture/adr-0011-approval-status-flow.md
- "ADR-0012: Citizen reference correlation": architecture/adr-0012-citizen-reference-correlation.md
- "ADR-0013: Behandel-portal wiring": architecture/adr-0013-behandel-portal-wiring.md
- "ADR-0014: Withdrawal cancels the process": architecture/adr-0014-withdrawal-cancels-the-process.md
- "ADR-0015: Beoordeling escalation": architecture/adr-0015-beoordeling-escalation.md
- "ADR-0016: Diploma eligibility DMN": architecture/adr-0016-diploma-eligibility-dmn.md
- "ADR-0017: Document-wait timeout": architecture/adr-0017-document-wait-timeout-cancellation.md
- "ADR-0018: Diploma upload via the ACL": architecture/adr-0018-diploma-upload-via-acl-documenten.md
- "ADR-0019: Zaak cancellation on timeout": architecture/adr-0019-zaak-cancellation-on-timeout.md
- "ADR-0020: Local stack self-seeds": architecture/adr-0020-local-stack-self-seeds.md
- "ADR-0021: Zaaktype by identificatie": architecture/adr-0021-acl-resolves-zaaktype-by-identificatie.md
- "ADR-0022: Quartz scheduler": architecture/adr-0022-quartz-scheduler.md
- "ADR-0023: Observability stack": architecture/adr-0023-observability-stack.md
- "ADR-0024: Prometheus AspNetCore exporter": architecture/adr-0024-prometheus-aspnetcore-exporter.md
- "ADR-0025: BFF reads catalogus via the ACL": architecture/adr-0025-bff-reads-catalogus-via-acl.md
- "ADR-0026: Mutable default-fill store": architecture/adr-0026-mutable-default-fill-store.md
- "ADR-0027: RegisterRecord objecttype": architecture/adr-0027-registerrecord-objecttype-schema.md
- "ADR-0028: Objecten holds the register": architecture/adr-0028-objecten-holds-the-register.md
- "ADR-0029: Objecten publishes to NRC": architecture/adr-0029-objecten-publishes-to-nrc.md
- "ADR-0030: Projection sourced from the register": architecture/adr-0030-projection-sourced-from-the-register.md
- "ADR-0031: MFA on the medewerker realm": architecture/adr-0031-mfa-on-the-medewerker-realm.md
- "ADR-0032: Werkbak live refresh": architecture/adr-0032-werkbak-live-refresh.md
- "ADR-0033: Kubernetes via one Helm chart": architecture/adr-0033-kubernetes-via-one-helm-chart.md
- "ADR-0034: Caddy serves the portals": architecture/adr-0034-caddy-serves-the-portals.md
- FDS-architectuur:
- Overzicht: architecture/fds/README.md
- Componentview (L3): architecture/fds/c4-component-view.md
@@ -70,15 +46,8 @@ nav:
- Working in Gitea: gitea-workflow.md
- Frontend decisions: frontend-decisions.md
- Demo script: demo-script.md
- Synthetic data: synthetic-data.md
- Runbooks:
- CI: runbooks/ci.md
- OpenZaak: runbooks/openzaak.md
- Open Notificaties (NRC): runbooks/opennotificaties.md
- Keycloak: runbooks/keycloak.md
- Flowable: runbooks/flowable.md
- Kubernetes on Talos: runbooks/kubernetes-talos.md
- Gitea Actions gotchas: runbooks/gitea-actions-gotchas.md
markdown_extensions:
- admonition