Design spike · safety.viz · 9 October 2026

The 22-minute check can run in about 6

Every pull request to safety.viz waits on one check, “Build, format, and test”. On 9 October it took a median of 22.6 minutes. I timed five other ways of running the same tests on the same runners. Nothing is merged yet.

Decisions

22.6 min
Today. The check runs both test suites, then the evidence guard runs both again.
12.1 min
Run the suites once and let the evidence guard read their results. No test changes.
6.3 min
Also run the browser suite on two runners at once, with the eight real-R tests on their own. No test changes.

Where the 22 minutes go

One recent run on dev, 22 minutes 57 seconds, drawn to scale. The hatched part is work the same run had already done.

Source: CI run on dev, 9 October 04:04 UTC; ci.yml; scripts/evidence.mjs.

Why it doubled in two days

Median length of a successful check, by day, over the last 100 successful runs.

Inside the browser suite

Test time by spec file, from the same run. Purple is the eight real-R tests; blue is the other 398.

What I measured

A temporary workflow on a scratch branch ran each layout on the same kind of runner the check uses, on the same commit of dev. Times are from the first job starting to the last job finishing.

LayoutWall clockResultWhat it showed
Today: one job, suites run twice22.6 minpassesMedian of the 9 successful checks on 9 October.
One job, suites run once, guard reads the results12.1 minpassesThree runs: 12.0 and 12.1 minutes, and 22.0 when the browser install stalled for 9.5 minutes (see the note further down). The guard itself took under a second.
The same, with four browser workers13.1 minpassesNo gain. The browser step took 10.8 minutes, held by the one long file.
The same, four workers, tests inside a file side by side10.6 min1 test failsThe local demo spec starts a server; two workers each start one on the same port and one test cannot connect. Would need test changes.
Four shards, tests inside a file side by side, then a gate8.4 minpassesShards are cut by test count, and the first one gets the whole demo app spec: 7.3 minutes against 0.5 for the second.
Two browser jobs split by real R, a static job, then a gate6.3 minpassesThree runs: 5.3, 6.3 and 6.5 minutes. The spread is the real-R job, whose start-up time varies from run to run.

Sources: round 1, round 2, the temporary workflow, everything the scratch branch changes.

Recommendation

  1. Run the suites once

    22.6 → about 12 minutes
    • The unit and browser steps each write a results file; the evidence guard reads the two files and no longer starts the suites itself.
    • About 30 lines in the evidence script and three changed lines in the workflow. The script change is written and exercised on the scratch branch.
    • Run on its own machine, npm run evidence:check behaves as it does today.
    • No test changes, no new jobs, same check name.
  2. Give the browser suite two runners and a gate

    about 12 → about 6 minutes
    • Four jobs in place of one: static checks with unit tests; browser tests without real R; the eight real-R tests; and a gate that merges the reports and runs the evidence guard.
    • The gate takes the name “Build, format, and test”, so the required check on dev is unchanged and no ruleset edit is needed.
    • The real-R tests are picked by a tag written on each test, so a new one joins its job by carrying the tag. The two browser jobs use one pattern, one of them inverted, so every test runs in exactly one job.
    • The public repository’s runners are free, and the four jobs use about 13 runner-minutes against today’s 22.
  3. Give the evidence refresh run the same split, later

    7 to 11 minutes today; not measured
    • The run that rewrites screenshots and evidence files before a check is a single pass over both suites, so it has no repeat to remove. It ran 52 times in the first nine days of October.
    • Splitting it means collecting rewritten screenshots from two runners before the commit. That is more work than the check’s split and I did not try it.
    • With steps 1 and 2 done, a change that needs new screenshots waits about 16 minutes in all, refresh then check, against about 33 today.

Three things the build has to get right

  • GitHub counts a required job that was skipped as passed. The gate has to run even when a job before it fails, and fail unless all three succeeded. The scratch branch’s gate always ran but did not check the three results; see the security assessment.
  • The evidence files key on each test’s name, so the tag must not change it. I checked on a throwaway test: a tagged test reports the same name, and the two patterns pick it up or leave it out as intended. The scratch branch itself still picks the eight by their requirement IDs.
  • Today’s check uploads one browser report for debugging a failure. The gate should upload the merged one.

Security assessment of the new layout

Added later on 9 October. A second agent, given the scratch branch and the plan but not this page’s conclusions, looked for ways the new check could say green when it should not. It ran the evidence guard against hand-edited results from a real run of the 2,392 unit tests and the real list of 406 browser tests. It pushed nothing, so what it says about GitHub’s own behaviour is read from documentation, not run.

A correction to this page

  • Above, I wrote that a job that drops out cannot pass silently. That holds for a missing test and no further. The gate on the scratch branch ran whenever the jobs before it finished and let the guard alone decide, and the guard looks only at test names and pass or fail.
  • The assessment showed the guard exits clean on a test that failed and then passed on a retry, on unit results marked unsuccessful because a test file failed to load, and on a browser report that carries errors. In today’s single job the test steps’ own exit codes catch all of these.
  • The spike’s timings stand. Its gate was not safe to require as written, and the requirements now say what makes it safe.
FindingSeverityFrom this change?What closes it
The gate can report success while a job before it failedHighYesThe gate’s first step fails unless each of the three jobs ended in success. Nothing in the workflow may continue on error.
GitHub counts a skipped required job as passedHighYes, as a hazard; the spike had this part rightThe gate always runs. A unit test reads the workflow file and fails if the gate’s shape changes.
The guard trusts the results files more than exit codes didMediumMade worseHanded files, the guard also fails on a run marked unsuccessful, on errors, on unexpected or flaky results, and on a test name that appears twice. One test for each.
The ruleset accepts a check of this name from any sourceMediumNo, it was already soA ruleset edit that ties the check to GitHub Actions, which is yours alone. It also means changing the program’s ruleset standard, so it is raised with the sign-off and left out of the build.
Picking tests by tag: a pattern that matches nothing, or a typoLowYesOne variable feeds both browser jobs. The guard already names any test that ran in neither.
Installing the browser without its system packages could shift fonts inside the 2% screenshot limitLowOnly if adoptedNot adopted here. Both workflows keep the packages and name the runner image.
Artifacts downloaded into the checkout could pick up a committed fileLowYesDownload to the runner’s temporary directory, fixed names, and an error when an upload finds no file.
Pull requests from forksFor informationNo new exposureThe check stays on the read-only pull request trigger and never gains write permission.

Looked at and set aside

One thing I saw and did not chase

  • The step that installs the browser also installs system fonts through the runner’s package mirror. It normally takes 21 to 25 seconds. In one of my runs the mirror slowed to 39 kB a second and the step took 9.5 minutes; one of the last 70 real checks shows the same thing at 3.3 minutes.
  • It is rare and it clears on a re-run. The side test above suggests the install can skip the mirror altogether. It is not adopted in this work; the security assessment below says why, and it is kept as a follow-up.

The decision

Decided on 9 October; see Decisions at the top. As first put:

Starting the build

gh issue comment 412 -R jwildfire/obot.roadmap --body "Tree signed off: the two requirements (#410 and #411) and their tasks, as filed on 2026-10-09."
Build the faster required check for safety.viz, in release 1.11. The objective is "A fast check" (jwildfire/obot.roadmap#412). Its tree is filed, and my sign-off comment is on the objective; if you cannot find one from my account, stop and ask.

Run /requirement-session 410. When it closes with its proof, tell me, and ask before starting requirement 411: it is first on the cut line, and whether it runs depends on where release 1.11 stands that day.

Order inside requirement 410:

1. safety.viz#291 first, alone, and merged before anything else starts: the evidence guard reads the results it is handed, and the check runs each suite once. It takes the check from about 22 minutes to about 12 for every other session, so it is worth having on dev today.
2. safety.viz#292: the four jobs behind a gate named "Build, format, and test".
3. obot.roadmap#413 last: the developer guidelines, written from the workflow as merged. Show me the diff.

Read before building:
- The objective's Boundaries, which hold the constraints, the cut line and what is left out.
- The spike and its security assessment: https://jwildfire.github.io/obot.roadmap/reports/sv-ci-speed-spike-2026-10-09/ . Its source is in the obot.roadmap clone at reports/sv-ci-speed-spike-2026-10-09/.
- Requirement 410's Design section. It is the specification, and the part headed "What the security assessment of this layout adds" is not optional.
- The prototype: the branch spike/ci-speed in safety.viz, checked out at safety.viz/.claude/worktrees/spike-ci-speed. Read .github/workflows/spike-ci-speed.yml and its change to scripts/evidence.mjs, and start from the two flags if that helps. Never merge it. Its gate is wrong as written: it always runs, but it does not check the results of the jobs before it.

What matters most:
- No test is removed, skipped, renamed or loosened, and the screenshot limit stays at 2%. The speed comes from running each test once, on more runners, and from nothing else. If something can only be made to pass by checking less, stop and ask me.
- The check must never say green when a test did not run and pass. Prove it by breaking it: every deliberately broken run the tasks list is run for real on a scratch branch and linked, red at the gate, before the pull request opens. A failure you reasoned about and did not run does not count.
- The ruleset on dev is not touched and the check keeps its name. If you come to think a ruleset must change, write me the command and mark the task blocked.
- Release 1.11's build comes first. Another session, "safety.viz v1.11.0 demo app", is landing pull requests on dev through this same check, and it has been told this work is coming. Message it before each merge and after it, saying what changed and what it may need to do: after safety.viz#291 the evidence files on its open branches will conflict, and after safety.viz#292 a new test that starts real R carries the tag. If the new check gives a wrong answer once it has landed, a red on good code or a green on bad, revert the workflow change first and diagnose after.

Known from the spike, so you do not find them again:
- A unit test outside a chart's directory lands in all 14 evidence files. The new tests in safety.viz#291 need the evidence refresh run on that branch before its pull request opens.
- The bot cannot start a workflow_dispatch run, and a refresh it starts by repository_dispatch uses the workflow file on dev. Rehearse a changed workflow with a temporary push trigger on a scratch branch, as spike-ci-speed.yml does.
- The step that installs the browser sometimes stalls for minutes on a slow package mirror. Re-run it. Do not fix it by dropping --with-deps; the security assessment says why.
- The two browser jobs came out at 4 to 6 minutes each in three runs; the real-R job is the longer one and its length varies with R's start-up.

Not verified by anyone; check each before building on it, and tell me if one fails:
- What the gate sees when only a failed job is re-run.
- Whether a unit test can read ci.yml with what the project already depends on, or needs a YAML parser added. Say which you chose and why.
- For requirement 411 only: whether a refresh pattern and the tag combine in one --grep.

Housekeeping:
- Name any scratch branch you push spike/..., open no pull request from it, and list every one at the end for me to approve removing. Leave spike/ci-speed and its worktree as they are.
- When requirement 410 closes, give me one message: the lengths of the five most recent checks, the links to the broken runs, and whether you would run requirement 411 now or cut it.

What this spike left behind