## What & why #161 is really two defects, and the second one is why the first was undiagnosable. **A wedged suite consumed the job, and took the post-mortem with it.** Nothing bounded the Playwright run, so CI stopped the job mid-suite — and `if: always()` does not survive that. Run 739's job metadata shows every step after the e2e as a **0-second failure** stamped at the kill: ``` 14 failure 09:48:17 -> 10:14:54 Self-service e2e (Playwright …) 15 failure 10:14:54 -> 10:14:54 verify-stack check summary ← if: always() 16 failure 10:14:54 -> 10:14:54 e2e spec summary ← if: always() 17 failure 10:14:54 -> 10:14:54 Dump container logs on failure ← if: failure() 18 failure 10:14:54 -> 10:14:54 Tear down ← if: always() ``` So the per-spec summary, the container-log dump and the teardown never ran, and the log lost whatever the killed process had buffered — leaving the single `✘` line the issue was filed from. `globalTimeout` now makes Playwright stop and *report*: the JSON report is written and those steps still get their turn. (A `timeout-minutes` on the job would have reproduced the same failure, so there isn't one.) The "~24-minute gap" is that kill, not necessarily a hang — note run 739 shows `run_attempt: 2`, and `concurrency.cancel-in-progress` kills an in-flight run on any re-run or push. **A login that never got its form ate the 90-second test timeout.** Playwright actions auto-wait until the *test* timeout, not `expect.timeout` — so a portal that serves its page but never bootstraps (its `config.json` fetch or the OIDC discovery behind `authorize()` failed; `main.ts` only `console.error`s) spent 90s to report `locator.fill: Test timeout of 90000ms exceeded`: the symptom, not the cause. That is catalogus.spec's 1.8 minutes. Both Keycloak forms are now asserted visible first, with a 20s budget and a message naming the step that never happened. Verified against a real blank-bootstrap portal — the beheer image served with a `config.json` that is not JSON — which fails in **20.2s** with *"the Keycloak login form never appeared — the portal did not reach Keycloak (check its config.json fetch and the OIDC discovery …)"*. **And the summary now says why.** The per-spec table (#136) rendered a verdict icon and nothing else, so even a surviving summary cost a log dive. Failing specs now carry their first error, flattened for a table cell (ANSI stripped, newlines collapsed, `|` escaped, clipped) — shape verified against a real @playwright/test 1.61 failing report, with a stdlib assert self-check on `make unit`. Closes #161 ## Definition of Done - [x] Linked Gitea issue (above). - [x] Failing test committed before the implementation. - [x] Implementation makes the test pass; refactor commit follows (login helper dedup). - [x] Conventional Commits referencing the issue (`refs #161`). - [ ] CI green — all Gitea Actions jobs. - [x] `docker compose up` from a fresh clone reaches green health checks within 3 minutes (untouched). - [x] Docs updated — `docs/runbooks/gitea-actions-gotchas.md` §9. - [x] ADR — not needed: no boundary, dependency or coupling rule touched (test/CI infra only). - [x] Demo note — not applicable: nothing user-visible. ## Notes for reviewers **What this does not do: identify why the beheerder login failed that once.** The evidence to do that was destroyed by defect 2, which is what this PR fixes. The suite ran green here five times today (catalogus.spec 1.1–5.3s each) — but a local box is not the loaded CI runner, so that is weak evidence and I am not claiming the flake is gone. What changes is that the next occurrence is bounded and self-describing: it fails in 20s naming the failing step, the JSON report survives, and the summary prints the error. Please keep #161 in mind rather than treating this as proof. **Two follow-ups I did not pull into this PR:** - *All four portals show a permanently blank page if their startup fetch fails* — `main.ts` does `fetch('config.json').then(bootstrap).catch(console.error)`, one shot, no UI and no recovery. That is a real product gap (the deliberately-broken portal above is exactly what a user would see) and wants its own slice, not a test-infra PR. - `retries: 1` is untouched. CLAUDE.md §15 says flaky tests are fixed rather than retried, but removing retries while a real flake is unexplained would trade a rare red for a frequent one. Worth revisiting once #161 recurs (or doesn't) with the new diagnostics. The login-helper rename (`medewerker-login.ts` → `keycloak-login.ts`, citizen logins routed through `loginBurger`) is its own no-behaviour-change commit: the three citizen specs each duplicated the same three-line login, so guarding the login path once meant routing them through it first.Reviewed-on: #165
57 lines
3.6 KiB
TypeScript
57 lines
3.6 KiB
TypeScript
import { defineConfig, devices } from '@playwright/test';
|
||
|
||
// The e2e runs inside the compose network (infra/run-e2e-check.sh); baseURL defaults to the
|
||
// self-service service. Keep timeouts generous — the first navigation triggers the DigiD flow.
|
||
const baseURL = process.env.SELF_SERVICE_URL ?? 'http://self-service';
|
||
// The behandel portal is a second origin the happy path visits (staff approve from the werkbak);
|
||
// it needs the same insecure-origin-as-secure treatment as self-service for the PKCE login (below).
|
||
const behandelURL = process.env.BEHANDEL_URL ?? 'http://behandel';
|
||
// The beheer portal is a third medewerker-realm origin (the read-only catalogus viewer, S-15a); it
|
||
// needs the same insecure-origin-as-secure treatment as the others for the PKCE login (below).
|
||
const beheerURL = process.env.BEHEER_URL ?? 'http://beheer';
|
||
|
||
export default defineConfig({
|
||
testDir: '.',
|
||
timeout: 90_000,
|
||
expect: { timeout: 15_000 },
|
||
retries: 1,
|
||
// Bound the whole run, not just each test (#161). A wedged suite used to run until CI killed the
|
||
// job — which also killed the `if: always()` steps that would have said why: the per-spec summary
|
||
// and the container-log dump never ran, leaving a 36-minute job whose entire surviving output was
|
||
// one ✘ line. On `globalTimeout` Playwright stops and *reports*, so the JSON report is written and
|
||
// those steps still run. Generous over the ~1-minute suite: this is a backstop, not a budget.
|
||
globalTimeout: 12 * 60_000,
|
||
// Run the specs serially. Each spec drives a full `channel: 'chromium'` browser, and the e2e
|
||
// shares an 8 GB runner with the entire compose stack (OpenZaak, NRC, Keycloak, Flowable, 4×
|
||
// Postgres, every service + 3 portals). Two parallel browsers exhaust memory and the renderer is
|
||
// OOM-killed mid-action ("Page crashed") — fixing the flakiness at its source rather than leaning
|
||
// on `retries` (CLAUDE.md §15). Only two long-running happy-path specs, so serial costs little.
|
||
workers: 1,
|
||
// `list` for the live log; `json` (→ /e2e/playwright-report.json in the container) is copied out
|
||
// by run-e2e-check.sh and rendered as a per-spec table in the CI job summary (#136).
|
||
reporter: [['list'], ['json', { outputFile: 'playwright-report.json' }]],
|
||
use: {
|
||
baseURL,
|
||
trace: 'on-first-retry',
|
||
// The portal is served over plain HTTP on a non-localhost origin (http://self-service) inside the
|
||
// compose network, so it is NOT a secure context — and Web Crypto (`crypto.subtle`) is undefined
|
||
// there. angular-auth-oidc-client needs SubtleCrypto to build the PKCE code challenge, so
|
||
// `authorize()` throws and the login redirect never fires (the login form never appears). In
|
||
// production the portal runs behind HTTPS, where this works. Rather than terminate TLS in the
|
||
// throwaway e2e stack, tell Chromium to treat this origin as secure — which faithfully emulates
|
||
// the production HTTPS context. This flag is only honoured by the full Chromium build (new
|
||
// headless), not Playwright's default headless-shell, so pin `channel: 'chromium'`.
|
||
channel: 'chromium',
|
||
launchOptions: {
|
||
args: [
|
||
`--unsafely-treat-insecure-origin-as-secure=${baseURL},${behandelURL},${beheerURL}`,
|
||
// Write Chromium's shared memory to /tmp instead of the container's small /dev/shm, so a
|
||
// large DOM/heap can't crash the renderer on the memory-constrained runner (belt-and-braces
|
||
// alongside the single worker above).
|
||
'--disable-dev-shm-usage',
|
||
],
|
||
},
|
||
},
|
||
projects: [{ name: 'chromium', use: { ...devices['Desktop Chrome'] } }],
|
||
});
|