Files
atomic-design-poc/docs/project/backlog/WP-30-ci-perf-followups.md
T
ehoandClaude Opus 4.8 c404995980
CI / frontend (push) Failing after 31s
CI / storybook-a11y (push) Failing after 30s
CI / backend (push) Failing after 30s
CI / e2e (push) Failing after 30s
CI / semgrep (push) Failing after 30s
CI / api-client-drift (push) Failing after 30s
ci: replace CodeQL with Semgrep (Gitea-compatible SAST)
CodeQL is GitHub-only — its analyze step uploads SARIF to GitHub's code-scanning
API and assumes a GitHub Security tab; this CI runs on Gitea only, so the job could
never go green (it had been red since it was added). Replace it with Semgrep OSS, a
plain CLI SAST with no account/platform API, which runs fine on Gitea.

- Remove the codeql job (+ its security-events permission) and the schedule trigger
  (it existed only for codeql; semgrep runs on push + PR).
- Add a semgrep job: setup-python + `pip install semgrep` +
  `semgrep scan --config p/default --config p/csharp --metrics=off`. pip-on-runner
  (not container:) mirrors the other jobs' model; anonymous registry, telemetry off.
- Report-only for now (no --error → job stays green): a local dry-run found 27
  findings, mostly CI/config policy (unpinned actions, .npmrc), not app-code vulns.
  WP-30 tracks triaging them + flipping to --error (a blocking gate).

Verified locally: `semgrep scan` runs clean (exit 0 without --error, 306 rules /
450 files). CI behaviour confirmable only on the Gitea runner — watch the run.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-22 09:34:13 +02:00

70 lines
3.9 KiB
Markdown

# WP-30 — CI performance follow-ups
Status: todo
Phase: follow-on · CI/infra
## Why
Tier-1 CI speedups shipped in `708d4c2` (CodeQL off the PR path, Playwright/NuGet caches,
`npm ci` flags) and the demo web image shrank to `node:24-slim`. These are the remaining
options that were deliberately deferred — bigger changes, policy calls, or things that need
Gitea runner-admin access. Revisit once there's an actual CI-timing breakdown to prioritise by,
or when someone confirms act_runner access.
Constraint carried over: **CI runs are not observable from the agent's environment** — validate
any workflow edit by watching a real Gitea run; ship one change at a time so a red run is easy to
bisect and revert. **The `docker compose` images are NOT used by CI** (CI = Gitea `ubuntu-latest`
runner image, set on the act_runner host).
## Read first
- `.github/workflows/ci.yml` (current 6 jobs + the Tier-1 caches already in place).
- The `ci-and-local-gate` note (agent memory) — CI traps + what's already done.
- `docker-compose.yml` (demo images; `node:24-slim` done, dotnet SDK still full).
## Candidate items (pick per impact once measured)
1. **Skip `npm ci` install via a `node_modules` cache.** `actions/cache` on `node_modules`
keyed by `package-lock.json` hash; on a hit, `npm ci` is near-instant across the 4 npm jobs.
Bigger win than the existing npm-download cache, but a ~777 MB cache with a small staleness
risk — best if the runner's cache storage is local/fast. Medium effort, low-medium risk.
2. **Smaller CI runner image.**
- _Real fix (needs runner admin):_ point act_runner's `ubuntu-latest` (or a new label) at a
smaller image with node + dotnet preinstalled. Biggest startup win. **Blocked on confirming
act_runner access.**
- _Repo-only partial:_ `container: node:24-slim` on the node-only jobs (`frontend`,
`storybook-a11y`), dropping `setup-node`. Doesn't help the node+dotnet jobs (`e2e`,
`api-client-drift`, `backend`) — a combined image would need building/pushing (new infra).
Risky on act_runner, unverifiable locally → stage alone, last.
3. **Path-filtered jobs.** Skip `backend` on FE-only changes and vice-versa (workflow `paths:`
or `dorny/paths-filter`). Cuts compute on narrow PRs; watch required-check rules that expect
every job to report a status.
4. **Split a fast `lint` job** (lint + format:check + check:tokens) for ~1 min fail-fast
feedback — only worth it once item 1 (node_modules cache) lands, else it duplicates `npm ci`.
5. **Lean deployable backend image** (optional, not for the dev demo): multi-stage prod build on
`mcr.microsoft.com/dotnet/aspnet:10.0` (~220 MB) in a separate `docker-compose.prod.yml`. The
dev `docker-compose.yml` keeps the SDK image because `dotnet run` hot-reload needs it.
6. **Semgrep: triage findings + make it blocking.** Semgrep replaced CodeQL (GitHub-only, couldn't
run on Gitea) and currently runs **report-only** — a local dry-run found 27 findings, mostly
CI/config policy (unpinned GitHub Actions in `ci.yml`, `.npmrc` min-release-age) rather than
app-code vulns. Triage them (fix or `# nosemgrep`/`.semgrepignore` the noise; consider a
tighter ruleset than `p/default` if the GitHub-Actions-policy rules aren't wanted), then add
`--error` to `semgrep scan` so it's a real gate.
## Acceptance criteria
- [ ] Each chosen item verified GREEN on a real Gitea run (watched, since it's not observable
from the agent env), landed as its own revertable commit.
- [ ] `npm run ci` still passes locally after any workflow/script change.
## Out of scope
- nx/turbo/remote build caching (overkill for this repo size).
- Rewriting the dev compose into prod images (would lose bind-mount hot reload).
## Risks
- Unverifiable-from-agent workflow edits can only be confirmed on the runner — one change per
commit, revert on red.
- act_runner-level changes depend on infra access this repo doesn't control.