Files
atomic-design-poc/docs/project/backlog/WP-30-ci-perf-followups.md
T
ehoandClaude Sonnet 5 c2e06cc8d7
CI / changes (push) Successful in 30s
CI / lint (push) Successful in 4m0s
CI / frontend (push) Successful in 4m42s
CI / backend (push) Successful in 2m27s
CI / e2e (push) Successful in 3m36s
CI / semgrep (push) Failing after 1m11s
CI / storybook-a11y (push) Successful in 8m38s
CI / api-client-drift (push) Successful in 1m48s
docs(backlog): WP-30 status update — 5 of 6 items landed
Records what's implemented (items 1/3/4/5/6), what's deliberately skipped
this round (item 2, blocked on act_runner access), and that the WP can't be
marked fully done until a real Gitea run confirms the CI-timing/path-filter
behavior this environment can't observe. npm run ci confirmed green locally.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-30 09:36:36 +02:00

5.8 KiB

WP-30 — CI performance follow-ups

Status: in-progress (5 of 6 items implemented + committed; pending a real Gitea run to confirm — see "Status update" below) Phase: follow-on · CI/infra

Why

Tier-1 CI speedups shipped in 708d4c2 (CodeQL off the PR path, Playwright/NuGet caches, npm ci flags) and the demo web image shrank to node:24-slim. These are the remaining options that were deliberately deferred — bigger changes, policy calls, or things that need Gitea runner-admin access. Revisit once there's an actual CI-timing breakdown to prioritise by, or when someone confirms act_runner access.

Constraint carried over: CI runs are not observable from the agent's environment — validate any workflow edit by watching a real Gitea run; ship one change at a time so a red run is easy to bisect and revert. The docker compose images are NOT used by CI (CI = Gitea ubuntu-latest runner image, set on the act_runner host).

Read first

  • .github/workflows/ci.yml (current 6 jobs + the Tier-1 caches already in place).
  • The ci-and-local-gate note (agent memory) — CI traps + what's already done.
  • docker-compose.yml (demo images; node:24-slim done, dotnet SDK still full).

Candidate items (pick per impact once measured)

  1. Skip npm ci install via a node_modules cache. actions/cache on node_modules keyed by package-lock.json hash; on a hit, npm ci is near-instant across the 4 npm jobs. Bigger win than the existing npm-download cache, but a ~777 MB cache with a small staleness risk — best if the runner's cache storage is local/fast. Medium effort, low-medium risk.
  2. Smaller CI runner image.
    • Real fix (needs runner admin): point act_runner's ubuntu-latest (or a new label) at a smaller image with node + dotnet preinstalled. Biggest startup win. Blocked on confirming act_runner access.
    • Repo-only partial: container: node:24-slim on the node-only jobs (frontend, storybook-a11y), dropping setup-node. Doesn't help the node+dotnet jobs (e2e, api-client-drift, backend) — a combined image would need building/pushing (new infra). Risky on act_runner, unverifiable locally → stage alone, last.
  3. Path-filtered jobs. Skip backend on FE-only changes and vice-versa (workflow paths: or dorny/paths-filter). Cuts compute on narrow PRs; watch required-check rules that expect every job to report a status.
  4. Split a fast lint job (lint + format:check + check:tokens) for ~1 min fail-fast feedback — only worth it once item 1 (node_modules cache) lands, else it duplicates npm ci.
  5. Lean deployable backend image (optional, not for the dev demo): multi-stage prod build on mcr.microsoft.com/dotnet/aspnet:10.0 (~220 MB) in a separate docker-compose.prod.yml. The dev docker-compose.yml keeps the SDK image because dotnet run hot-reload needs it.
  6. Semgrep: triage findings + make it blocking. Semgrep replaced CodeQL (GitHub-only, couldn't run on Gitea) and currently runs report-only — a local dry-run found 27 findings, mostly CI/config policy (unpinned GitHub Actions in ci.yml, .npmrc min-release-age) rather than app-code vulns. Triage them (fix or # nosemgrep/.semgrepignore the noise; consider a tighter ruleset than p/default if the GitHub-Actions-policy rules aren't wanted), then add --error to semgrep scan so it's a real gate.

Status update (2026-07-30)

Items 1, 3, 4, 5, 6 implemented, each as its own commit (item 6 526da76, item 1 e46b87b, item 4 e02e8ce, item 3 e7db69d, item 5 see git log -- backend/Dockerfile): triaged real local semgrep findings (25, not the 27 this file remembered — dependabot cooldown, npm min-release-age, every GH Action pinned to SHA, 2 nosemgrep'd ReDoS false positives) and flipped the gate to --error; node_modules cache (skips npm ci entirely on a hit) across all 4 npm jobs; a new fast-fail lint job split out of frontend; a changes job (dorny/paths-filter) gating every downstream job's real steps (not the whole job — the safer "skip steps" variant, so a required-status-check never waits on a job that never ran) on which side changed; an optional backend/Dockerfile + docker-compose.prod.yml (additive, unused by CI or the dev demo).

Item 2 (smaller runner image) deliberately skipped this round — the real fix needs act_runner admin access (unconfirmed), and the repo-only partial (node:24-slim on frontend/storybook-a11y) conflicts with storybook-a11y's deliberately-chosen node:24-bookworm + memory-cap container (verified against a real OOM risk). Revisit once act_runner access is confirmed.

Cannot self-certify GREEN: per this WP's own constraint, CI timing/behavior isn't observable from the agent's environment. Everything above was checked as far as locally possible (YAML parse, actionlint 0 issues, npm run ci, a real docker build) but the actual speedup and the path-filter's interaction with any required-status-check config need a watched Gitea run before this WP can be marked fully done.

Acceptance criteria

  • Each chosen item verified GREEN on a real Gitea run (watched, since it's not observable from the agent env), landed as its own revertable commit.
  • npm run ci still passes locally after any workflow/script change (confirmed 2026-07-30, full run including backend dotnet test/dotnet format and both drift checks).

Out of scope

  • nx/turbo/remote build caching (overkill for this repo size).
  • Rewriting the dev compose into prod images (would lose bind-mount hot reload).

Risks

  • Unverifiable-from-agent workflow edits can only be confirmed on the runner — one change per commit, revert on red.
  • act_runner-level changes depend on infra access this repo doesn't control.