fix(infra): anchor wait-healthy's container lookup on the compose replica suffix (refs #153)
CI / build (pull_request) Successful in 1m6s
CI / lint (pull_request) Successful in 1m22s
CI / unit (pull_request) Successful in 1m24s
CI / frontend (pull_request) Successful in 3m12s
CI / mutation (pull_request) Successful in 6m20s
CI / verify-stack (pull_request) Failing after 5m43s

Bring-up timed out with

  TIMEOUT: 'objecten' not healthy (status=none)

while the very `docker ps` it dumps showed infra-objecten-1 "Up 9 minutes (healthy)".

`--filter name=` is a substring match, so `objecten` also matches objecten-db,
objecten-redis and (since #152) objecten-celery. `head -1` took whichever docker listed
first; the celery worker declares no healthcheck, so it inspected as status=none and the
wait sat there until the deadline.

Not objecten-specific — `objecttypen` matches objecttypen-db the same way. The bug has been
latent since those services landed and was decided by listing order, which is why it only
surfaced now. Anchored on the replica suffix, matching both docker compose and
podman-compose naming — the same anchoring the verify check scripts already use.
This commit is contained in:
not
2026-08-28 13:01:10 +02:00
parent b496ac9477
commit 744f91a2b2
+7 -3
View File
@@ -15,9 +15,13 @@ set -euo pipefail
timeout="${WAIT_TIMEOUT:-420}" timeout="${WAIT_TIMEOUT:-420}"
deadline=$(( $(date +%s) + timeout )) deadline=$(( $(date +%s) + timeout ))
# compose service name -> container id. The name filter matches both docker # compose service name -> container id. `--filter name=` is a substring match, so it is anchored on
# compose ("infra-openzaak-1") and podman-compose ("infra_openzaak_1") naming. # the compose replica suffix — otherwise 'objecten' also matches objecten-db / objecten-redis /
cid_for() { docker ps -aq --filter "name=$1" | head -1; } # objecten-celery, and 'objecttypen' matches objecttypen-db. Whichever docker listed first won, so a
# service with a sibling that has no healthcheck timed out with status=none while it was in fact
# healthy. The pattern matches both docker compose ("infra-objecten-1") and podman-compose
# ("infra_objecten_1") naming; the same anchoring the verify check scripts use.
cid_for() { docker ps -aq --filter "name=$1[-_][0-9]+\$" | head -1; }
for svc in "$@"; do for svc in "$@"; do
echo "waiting for '$svc' to be healthy (timeout ${timeout}s)..." echo "waiting for '$svc' to be healthy (timeout ${timeout}s)..."