Files
register-referentie/docs/architecture/adr-0035-public-tls-edge-in-cluster.md
T
notandClaude Opus 5 399d110663
CI / k8s (pull_request) Successful in 6s
CI / build (pull_request) Successful in 1m28s
CI / lint (pull_request) Successful in 1m51s
CI / unit (pull_request) Successful in 1m12s
CI / frontend (pull_request) Successful in 2m11s
CI / mutation (pull_request) Successful in 3m45s
CI / verify-stack (pull_request) Successful in 8m15s
docs(arch): ADR-0035 and the runbook section for publishing the stack (refs #177)
The ADR records why the edge is in the cluster rather than on the Fedora host —
routing and certificates should be state a `helm upgrade` can see — and the
three costs that buys: the host forward nobody in the cluster can repair, the
Let's Encrypt rate limit that makes `persistence.storageClass` non-optional, and
publishing behandel and beheer to the internet behind synthetic accounts.

Runbook §10 is the operational half: the five DNS records, the two firewalld
rules (including the masquerade that makes the return path work), and the
symptoms each missing piece produces.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 16:30:38 +02:00

5.8 KiB

ADR-0035: The public TLS edge is a Caddy deployment in the cluster

  • Status: Accepted
  • Date: 2026-09-18
  • Deciders: Respellion engineering
  • Slice: #177

Context

The stack deploys to a Talos VM on the lab server (ADR-0033, issue #175). Until now it was only usable through five SSH port-forwards: the portals' OIDC flow uses PKCE, PKCE needs crypto.subtle, and browsers expose that only in a secure context — HTTPS or an origin on localhost. A NodePort on the VM's address is neither, so the deployment was pinned to host: localhost and every viewer had to forward all five browser-facing ports (a portal without Keycloak on the same localhost:30180 fails on the discovery document).

That is not a demo anyone can be sent a link to. We want public hostnames with real certificates — and we want the routing and the certificates to be cluster state, not host-side configuration that no helm upgrade can see.

The public IP is on the Fedora host (46.224.220.37); the cluster is a libvirt guest behind it.

Decision

Terminate TLS in the cluster, with a Caddy deployment rendered by the chart (templates/edge.yaml), and give the Fedora host nothing but a layer-4 forward.

  • public.domain is the single switch. Empty — the default, and what compose and CI use — renders nothing: the stack is reached on its NodePorts and host pins the OIDC origin exactly as before. Set it, and the edge appears.
  • public.routes maps a subdomain to an in-cluster service:port. Caddy proxies to the ClusterIP services, so a public deployment does not use the browser-facing NodePorts at all.
  • Caddy obtains and renews certificates itself (ACME HTTP-01). There is no cert-manager.
  • The host forwards :80/:443 to two NodePorts with two firewall-cmd --add-forward-port rules. No TLS, no routing, no per-service knowledge there — adding a portal is a chart change, not a host change.
  • KC_HOSTNAME and the portals' config.json stop being host + NodePort. Both now come from one helper, big.keycloakUrl, so the issuer Keycloak pins and the authority the portals are configured with cannot drift apart (ADR-0010).

Alternatives considered

  • Caddy on the Fedora host. Fewest moving parts — but the routing table and the certificates would live outside the cluster, in a file no deployment touches, and adding a portal would mean editing a host we deploy to over SSH. Rejected on exactly the ground this ADR exists to record.
  • Traefik or ingress-nginx, plus cert-manager. The conventional answer, and the right one for a cluster with many teams and changing hostnames. Here it buys a controller, a set of CRDs and Ingress objects to describe five hostnames that never change — and cert-manager to do what Caddy already does unprompted.
  • A LoadBalancer service (MetalLB). Solves address allocation, which is not the problem; the node has exactly one address and it still is not the public one.
  • Keep the SSH forwards. Free, and genuinely fine for one developer. It is not a demo you can send to someone.

Consequences

Positive

  • No new dependency: the four portals already run caddy:2-alpine (ADR-0034), whose ceiling note called this out — "a real hostname makes TLS a one-line Caddyfile change". This is that change.
  • Routing is cluster state: kubectl -n big get cm caddy-edge-config -o yaml is the whole truth about what is published, and helm upgrade is how it changes.
  • The secure context is real, so TALOS_HOST=localhost and the five forwards disappear — and with them the class of failure where a mismatched issuer logs the user out silently.
  • Nothing changes for compose, CI or a laptop cluster: with public.domain empty the rendered manifests are byte-identical to before.

Negative / costs

  • The host forward is irreducible. Two firewalld rules, applied by hand once, with sudo on a machine our pipeline reaches only over SSH. If someone rebuilds that host, the stack is unreachable until they are re-applied, and nothing in the cluster can tell them so.

  • Certificates need a volume. On the default emptyDir every pod restart asks Let's Encrypt again, and its duplicate-certificate limit is five per week — a handful of restarts and the edge serves an untrusted certificate for a week. persistence.storageClass stops being optional for anything public (runbook §6).

  • All five hostnames are published, including behandel and beheer, which approve registrations and administer the register. They are protected by synthetic accounts with well-known passwords, and by MFA on the medewerker realm (ADR-0031). That is a deliberate choice for a demonstration environment holding synthetic data only, and it is the reason this bullet is in the ADR rather than in a comment: if this stack ever holds anything real, this decision is the first one to revisit.

  • One more workload in the chart with no counterpart in compose — compose has no edge because it has no hostname. The drift check (make k8s-drift) renders the defaults, so it does not see it.

  • auth is load-bearing: big.keycloakUrl builds the issuer from that subdomain, so renaming the key in public.routes without the helper breaks every login. Both carry a comment saying so.

  • ponytail ceiling: one replica, no HSTS, no security headers beyond Caddy's defaults, no rate limiting, and HTTP-01 rather than DNS-01 (so a wildcard certificate is not available). Upgrade path in that order; DNS-01 first if the subdomain list ever grows.

Coupling rules touched (CLAUDE.md §8)

None. §8.3 holds — the browser reaches a portal, the portal reverse-proxies its own BFF group, and the edge is in front of all of it. The edge terminates TLS and routes by hostname; it does not know what any service does.