Every pull request to safety.viz waits on one check, “Build, format, and test”. On 9 October it took a median of 22.6 minutes. I timed five other ways of running the same tests on the same runners. Nothing is merged yet.
dev beside the release 1.11 work and ships in that release.dev. It keeps every test, every screenshot comparison and the evidence guard, and it keeps the check’s name, so the ruleset on dev does not change.One recent run on dev, 22 minutes 57 seconds, drawn to scale. The hatched part is work the same run had already done.
npm run evidence:check always runs both suites itself, and the workflow calls it after it has already run them as its own steps.Source: CI run on dev, 9 October 04:04 UTC; ci.yml; scripts/evidence.mjs.
Median length of a successful check, by day, over the last 100 successful runs.
Test time by spec file, from the same run. Purple is the eight real-R tests; blue is the other 398.
A temporary workflow on a scratch branch ran each layout on the same kind of runner the check uses, on the same commit of dev. Times are from the first job starting to the last job finishing.
| Layout | Wall clock | Result | What it showed |
|---|---|---|---|
| Today: one job, suites run twice | 22.6 min | passes | Median of the 9 successful checks on 9 October. |
| One job, suites run once, guard reads the results | 12.1 min | passes | Three runs: 12.0 and 12.1 minutes, and 22.0 when the browser install stalled for 9.5 minutes (see the note further down). The guard itself took under a second. |
| The same, with four browser workers | 13.1 min | passes | No gain. The browser step took 10.8 minutes, held by the one long file. |
| The same, four workers, tests inside a file side by side | 10.6 min | 1 test fails | The local demo spec starts a server; two workers each start one on the same port and one test cannot connect. Would need test changes. |
| Four shards, tests inside a file side by side, then a gate | 8.4 min | passes | Shards are cut by test count, and the first one gets the whole demo app spec: 7.3 minutes against 0.5 for the second. |
| Two browser jobs split by real R, a static job, then a gate | 6.3 min | passes | Three runs: 5.3, 6.3 and 6.5 minutes. The spread is the real-R job, whose start-up time varies from run to run. |
Sources: round 1, round 2, the temporary workflow, everything the scratch branch changes.
npm run evidence:check behaves as it does today.dev is unchanged and no ruleset edit is needed.Added later on 9 October. A second agent, given the scratch branch and the plan but not this page’s conclusions, looked for ways the new check could say green when it should not. It ran the evidence guard against hand-edited results from a real run of the 2,392 unit tests and the real list of 406 browser tests. It pushed nothing, so what it says about GitHub’s own behaviour is read from documentation, not run.
| Finding | Severity | From this change? | What closes it |
|---|---|---|---|
| The gate can report success while a job before it failed | High | Yes | The gate’s first step fails unless each of the three jobs ended in success. Nothing in the workflow may continue on error. |
| GitHub counts a skipped required job as passed | High | Yes, as a hazard; the spike had this part right | The gate always runs. A unit test reads the workflow file and fails if the gate’s shape changes. |
| The guard trusts the results files more than exit codes did | Medium | Made worse | Handed files, the guard also fails on a run marked unsuccessful, on errors, on unexpected or flaky results, and on a test name that appears twice. One test for each. |
| The ruleset accepts a check of this name from any source | Medium | No, it was already so | A ruleset edit that ties the check to GitHub Actions, which is yours alone. It also means changing the program’s ruleset standard, so it is raised with the sign-off and left out of the build. |
| Picking tests by tag: a pattern that matches nothing, or a typo | Low | Yes | One variable feeds both browser jobs. The guard already names any test that ran in neither. |
| Installing the browser without its system packages could shift fonts inside the 2% screenshot limit | Low | Only if adopted | Not adopted here. Both workflows keep the packages and name the runner image. |
| Artifacts downloaded into the checkout could pick up a committed file | Low | Yes | Download to the runner’s temporary directory, fixed names, and an error when an upload finds no file. |
| Pull requests from forks | For information | No new exposure | The check stays on the read-only pull request trigger and never gains write permission. |
dev. After the split their job is the longest of the three by about a minute, so leaving it out would save about a minute and take the comparison with desktop R out of the check.Decided on 9 October; see Decisions at the top. As first put:
gh issue comment 412 -R jwildfire/obot.roadmap --body "Tree signed off: the two requirements (#410 and #411) and their tasks, as filed on 2026-10-09."
Build the faster required check for safety.viz, in release 1.11. The objective is "A fast check" (jwildfire/obot.roadmap#412). Its tree is filed, and my sign-off comment is on the objective; if you cannot find one from my account, stop and ask. Run /requirement-session 410. When it closes with its proof, tell me, and ask before starting requirement 411: it is first on the cut line, and whether it runs depends on where release 1.11 stands that day. Order inside requirement 410: 1. safety.viz#291 first, alone, and merged before anything else starts: the evidence guard reads the results it is handed, and the check runs each suite once. It takes the check from about 22 minutes to about 12 for every other session, so it is worth having on dev today. 2. safety.viz#292: the four jobs behind a gate named "Build, format, and test". 3. obot.roadmap#413 last: the developer guidelines, written from the workflow as merged. Show me the diff. Read before building: - The objective's Boundaries, which hold the constraints, the cut line and what is left out. - The spike and its security assessment: https://jwildfire.github.io/obot.roadmap/reports/sv-ci-speed-spike-2026-10-09/ . Its source is in the obot.roadmap clone at reports/sv-ci-speed-spike-2026-10-09/. - Requirement 410's Design section. It is the specification, and the part headed "What the security assessment of this layout adds" is not optional. - The prototype: the branch spike/ci-speed in safety.viz, checked out at safety.viz/.claude/worktrees/spike-ci-speed. Read .github/workflows/spike-ci-speed.yml and its change to scripts/evidence.mjs, and start from the two flags if that helps. Never merge it. Its gate is wrong as written: it always runs, but it does not check the results of the jobs before it. What matters most: - No test is removed, skipped, renamed or loosened, and the screenshot limit stays at 2%. The speed comes from running each test once, on more runners, and from nothing else. If something can only be made to pass by checking less, stop and ask me. - The check must never say green when a test did not run and pass. Prove it by breaking it: every deliberately broken run the tasks list is run for real on a scratch branch and linked, red at the gate, before the pull request opens. A failure you reasoned about and did not run does not count. - The ruleset on dev is not touched and the check keeps its name. If you come to think a ruleset must change, write me the command and mark the task blocked. - Release 1.11's build comes first. Another session, "safety.viz v1.11.0 demo app", is landing pull requests on dev through this same check, and it has been told this work is coming. Message it before each merge and after it, saying what changed and what it may need to do: after safety.viz#291 the evidence files on its open branches will conflict, and after safety.viz#292 a new test that starts real R carries the tag. If the new check gives a wrong answer once it has landed, a red on good code or a green on bad, revert the workflow change first and diagnose after. Known from the spike, so you do not find them again: - A unit test outside a chart's directory lands in all 14 evidence files. The new tests in safety.viz#291 need the evidence refresh run on that branch before its pull request opens. - The bot cannot start a workflow_dispatch run, and a refresh it starts by repository_dispatch uses the workflow file on dev. Rehearse a changed workflow with a temporary push trigger on a scratch branch, as spike-ci-speed.yml does. - The step that installs the browser sometimes stalls for minutes on a slow package mirror. Re-run it. Do not fix it by dropping --with-deps; the security assessment says why. - The two browser jobs came out at 4 to 6 minutes each in three runs; the real-R job is the longer one and its length varies with R's start-up. Not verified by anyone; check each before building on it, and tell me if one fails: - What the gate sees when only a failed job is re-run. - Whether a unit test can read ci.yml with what the project already depends on, or needs a YAML parser added. Say which you chose and why. - For requirement 411 only: whether a refresh pattern and the tag combine in one --grep. Housekeeping: - Name any scratch branch you push spike/..., open no pull request from it, and list every one at the end for me to approve removing. Leave spike/ci-speed and its worktree as they are. - When requirement 410 closes, give me one message: the lengths of the five most recent checks, the links to the broken runs, and whether you would run requirement 411 now or cut it.
spike/ci-speed, two commits ahead of dev: the temporary workflow and the evidence script change. No pull request. The workflow runs only on pushes to that branch.dev or main changed. Removing the branch and the worktree is yours to approve.