obot release review written 2026-09-12, re-measured 2026-09-14 for tonight prepared for @jwildfire
Two releases are waiting on you. Here is the order, and what to look at in each.
Every release candidate in the portfolio that is open and requesting your review, re-measured from GitHub and the repositories on 14 September: one in open.csr, untouched since 2 September, and one in gsm.safety, the consolidated candidate you asked for on the 12th, which now exists with one notes section and one demo. Each section below says what the release lets someone do, links its pull request, release notes and demo, and lists what to check before you approve.
2 releases to review, 2 pull requestslongest wait 12 days (open.csr)gsm.safety consolidated 12 to 13 Sep, two candidates closedtonight's budget about 1 h 45safety.viz, obot.agent, open.gismo, demo-301: nothing open
Decision · @jwildfire · 2026-09-12, 12:51 UTC, in a comment on this page
"Let's go with option B, but squash it all into a single release: PR/Release notes/Demo for a single v1.2."
So gsm.safety ships once, as v1.2.0, from dev into main: one candidate, one release-notes section, one demo page. Done on the gsm.safety side on 12 and 13 September (gsm.safety#152 and #154) and on the hub on the 14th; the section below says what became of the stack. The plan as first written, with the two routes he chose between, stays there for the record.
Approved · @jwildfire · 2026-09-14, in the prep session
"I approve gsm.safety v1.2. go ahead with the release."
Recorded on the candidate (gsm.safety#88). The merge itself is blocked by the release ruleset on main: it requires an approving review from someone other than the pull request's author, the author is his own account, and the ruleset has no bypass list. It needs one edit from him (a bypass actor for the repository admin role, or the required review count set to 0 for this merge); the rest of the release is prepared and follows the merge. open.csr v0.4.0-RC1 is still waiting.
The report describes one study; five more reference displays cell for cell; the sidebar says where every number came from
main (v0.3.0)
2 Sep · 12 days
review requested
45 min
gsm.safety v1.2.0-RC1gsm.safety#88 · dev → main · 114 commits · retitled 12 Sep, head of 13 Sep
Everything since v1.1.0 as one release: the liver-chart correction, the census rebuilt on metrics, the last two widgets, the FDA thresholds as data
main (v1.1.0)
13 Sep as this candidate; the oldest work in it since 23 Aug
approved 14 Sepmerge blocked by the release ruleset; one ruleset edit needed
done
As first measured on 12 September the queue had four rows: open.csr#75 and three gsm.safety candidates (#68 v1.2.0, #69 v1.3.0, #88 v1.5.0), each containing the one below it. His decision folded the three into one; #68 and #69 were closed as superseded on the 13th. Nothing else is open in safety.viz, obot.agent, open.gismo or demo-301.
first · a decision before any reading
The gsm.safety stack, and what became of it
As it stood on 12 September. gsm.safety's main branch is at v1.1.0, released on 17 August. Three candidates are open into it and each contains the one before: v1.2.0 is three commits on top of main; v1.3.0 is those three plus seven; v1.5.0 is the whole of dev, thirty-five commits, which also carries the v1.4.0 widgets that never got a candidate of their own. Merging any one of them merges everything beneath it. The three pull-request bodies all say the same thing: land them in order, tag each.
The complication
The release workflow on both release branches still points at the organisation the templates moved away from on 4 August, so a tag cut from release/v1.2.0 or release/v1.3.0 as they stand publishes a release with no artifacts, exactly as v1.1.0 did (gsm.safety issue 75). The fix is one commit on dev (gsm.safety pull 73, 27 August) and neither release branch has it. Four other workflows fail at startup on those branches for the same reason, so their pull requests read six green checks while the qualification checks never ran.
Two ways through
A. In order, three tags, the plan of record. Ask a session to cherry-pick the workflow fix onto release/v1.2.0 and merge it forward into release/v1.3.0 first (a few minutes of agent work, one increment each). Then approve v1.2.0 and tag it; v1.3.0 already carries your approval and merges next; v1.5.0 shrinks to the v1.4.0 and v1.5.0 work and is reviewed last. What you keep: a v1.2.0 release whose headline is the liver-chart correction and a v1.3.0 release whose headline is the death count, each readable on its own by the person it affects, and the two demo pages already written for them.
B. One consolidated release. Review gsm.safety pull 88 as the whole diff from v1.1.0 and tag v1.5.0; close pull requests 68 and 69 as superseded. What you gain: one sitting, and the workflow fix comes along because dev has it. What you lose: no v1.2.0 or v1.3.0 GitHub release, so the two clinical corrections become sections inside the v1.5.0 notes rather than releases of their own, which is the argument the prep session made against this on 23 August.
Recommendation, as written this morning
Take A. The cost is one comment asking for the cherry-pick before you start, and it is the only route that keeps the demo pages, your existing approval and the per-release headlines all valid. Take B only if you want gsm.safety finished today in one review.
Either way, v1.4.0 is a loose end: its two widgets are on dev and in the v1.5.0 diff, its release-notes section is written, but no candidate or demo page was ever made and its hub requirement (#165) still reads backlog.
Decided · B, squashed to one v1.2.0
He chose B and went one step further: not v1.5.0 absorbing the rest, but a single v1.2.0, the next version after the one on main, with one pull request, one release-notes section and one demo. That settles the v1.4.0 loose end too, since nothing ships under its own number any more. The workflow fix is on dev, so the artifacts question goes away with the release branches.
What happened since, 12 to 14 September
Retitled and re-versioned. gsm.safety#88 is gsm.safety v1.2.0-RC1, dev into main; DESCRIPTION reads 1.2.0; the body is rewritten for the one release and names the four hub requirements it closes (gsm.safety#151, #152).
One release-notes section. NEWS.md on dev carries a single v1.2.0 (Upcoming) section that leads with the two clinical figures that change, then the four widgets, the census rebuild, FDA phase 0, and an Also in this release list of about twenty hardening fixes from the review rounds. The package dates the census rebuild to v1.2.0 everywhere it speaks of it (gsm.safety#153, #154).
One demo page.gsm.safety v1.2.0, annotated merges the three earlier pages as parts 1, 2 and 4 and adds part 3 for the two widgets. It deploys with this page; until then the candidate's See it move link lands on the earlier v1.2.0 demo, which now carries a banner pointing here.
gsm.safety#68 and #69 closed as superseded, their issues moved onto the v1.2.0 milestone.
Fifteen Copilot review rounds on the candidate, all resolved. Findings were fixed on dev in increments (the input-validation bullets in the notes) or filed for the release after this one: gsm.safety#147, #148 and #155. The last round, on the current head, found nothing and asked for a human review.
Hub bookkeeping. The four requirements it closes (#164, #274, #165, #9) now say they ship in gsm.safety v1.2.0 and point their proof at gsm.safety#88; #165 moved from backlog to review; your decision is quoted on #9.
suggested order · two sittings
Take them in this order
Review open.csr v0.4.0-RC1 nowIndependent of gsm.safety, the biggest clinical surface, and the one with questions only you can answer.
45 min
Review gsm.safety v1.2.0-RC1One diff from v1.1.0, one notes section, one merged demo page. The four parts below are the walk. Read the notes' two clinical corrections first, then the demo part by part, then the Also in this release list once.
60 min
Both can be done from a phone against the demo pages; the diff itself is for the laptop if you want it, and the checklists say where to look in it.
open.csr · v0.4.0-RC1 · open.csr#75
The report describes one study, and the sidebar says where every number came from
v0.3.0 shipped twenty-six displays over two packagings of the CDISC pilot that disagreed about who was on which arm, so section 10 and section 14 of the same document gave different arm sizes. This release reads the pilot's own package everywhere, declares the study once in a model, and fails the build when two displays disagree about it. It carries five more of the 2006 report's displays into the library cell for cell (demographics, exposure, adverse-event and serious-event incidence, subjects by site), draws the disposition flow, derives medications from the study's own SDTM, and fills three in-text tables from the section 14 results. The sidebar work you asked for on 2 September is folded in: Data and Metadata views, every artifact linked both ways, the explorer regrouped as Inputs, Pipeline, Outputs, and a flow diagram on every element.
The one-study claim, with your own eyes. Open the live report at section 10, then section 14. Both should read 86, 84 and 84. That is the release.
Demographics, Table 14-2.01. Open the display and its evidence page. The candidate reads Race (Origin) as a recode of race plus ethnicity, not a data conflict, and says so in the footnote. Agree or not with that reading; it changes a footnote, not a number.
Adverse-event incidence, Tables 14-5.01 and 14-5.02. Open the incidence table. Four Fisher p-values differ from the 2006 report by one thousandth and are held as known differences on the display. Confirm you accept them as such.
The sidebar, against what you asked on 2 September. Open the Data, Metadata and Pipeline tabs, then any display's flow diagram. Every box should be a link and every name real. This is the part of the release you specified in chat, so it is the part most worth a look on a phone.
The medications question the notes put to you. Medications now derive from the vendored SDTM CM domain; the relabelled PHUSE copy stays vendored only so the derivation can be checked against it. The release notes say removing the copy is a question for this review. Answer it in the review.
The Kenward-Roger footnote. Open the repeated-measures display. The spike moved one of the five last-digit differences and not the other four; the footnote now says what was tried. Read it as the wording a reader of the CSR would see.
The review gate. This candidate opened on 2 September, before the ultrareview rule and before you replaced it with a GitHub code review on 11 September (hub #340). No code review is posted on the pull request. Decide whether to require one before approving or to waive it for this candidate and say so in your review.
Scope against the requirement. The six laboratory families (open.csr#64) and Table 12-3 are deferred to v0.4.1 by D0032, so the requirement titled all thirty-one displays (#319) is delivered less those. Decide whether #319 stays open until v0.4.1 or is split now.
Checks. Thirteen gates green on the head per the body (R 166 of 166, vitest 465 of 465). A glance at the Checks tab confirms it on the current commit.
On approval
The tag v0.4.0 publishes with the notes as its body, the three hub requirements move to released, and v0.4.1 (lab families, Table 12-3, the explorer scroll fix) is the next candidate. Requested changes retitle this same pull request to RC2.
opened 2026-09-02refreshed twice the same day as the sidebar work landedunchanged since: same head, same 13 green checks on 14 Sepdev is fully inside the candidate: 0 commits behindcloses 11 open.csr issues
gsm.safety v1.2.0, part 1 · the widget parity work · was gsm.safety#68 · on the merged demo
The liver chart stops reading a participant's baseline as their peak
gsm.safety's composite eDISH view had been treating some participants' own baseline record as an on-treatment peak, because the on-treatment set was everything after study day zero and some participants' first record is not on day zero. This release swaps the chart library under all eleven widgets from safety.viz v1.4.0 to v1.7.0, where the rule was corrected, and adds the two charts R never had: the ALT waterfall and the kidney explorer. A parity guard in CI compares the latest safety.viz release against the package and fails on drift, with an allowlist that must cite a filed requirement for every deferral.
Measured on the example data: 293 of 364 participants plotted instead of 295, nine plotted peaks move and all move down, two participants leave the plot, and nobody changes quadrant. That is the clinical content of the release; everything else is feature or appearance.
The side-by-side on the demo page. Two captures from one report file, differing only in which bundle they load. Read the count line above each chart: 295 of 364 before, 293 after.
The ledger of every value that moves. Nine peaks, two exits, quadrant counts 248 → 246 in Normal and Not-Notable only. Confirm you are content that no participant's classification changes on this data, and that the page says this is a property of the study and not of the rule.
The three participants' own records. The demo prints every ALT record for the participants whose peak moved furthest, with the old and new rule applied line by line. This is the part to read if you want to be sure the correction is right rather than just different.
Try it live. Open hep-explorer on the safety.viz site, choose the composite plot, read the count line and compare with the capture.
The two new widgets. Widget_HepWaterfall and Widget_NepExplorer follow the established pattern; the demo links their safety.viz twins. A glance at each is enough; they carry no correction.
Checks. On the superseded branch six checks passed while four workflows failed at startup, the qualification checks among them. The consolidated candidate is dev, which carries the workflow fix, so every check actually runs there; read them on that pull request, not this one.
on dev since 2026-08-21its own candidate, gsm.safety#68, closed as superseded 13 Sepnow inside gsm.safety#88, reviewed by the Copilot rounds with the rest
gsm.safety v1.2.0, part 2 · the census rebuild · was gsm.safety#69, approved 27 Aug · on the merged demo
The safety census rebuilt on metrics, and the death count corrected
The safety overview reported four deaths on a study with thirteen. The old count matched the text of a discontinuation reason, never opened the death domain, and counted people who were never enrolled. This release replaces every figure on that page with a metric that publishes its own numerator, denominator and provenance, adds the report workflow that reads them and computes nothing of its own, and keeps SafetyCensus() by name with its arithmetic moved out from under it. Five published figures move: deaths 4 → 13, randomised blank → 577, dosed 744 → 762, with a disposition record 100 → 76, and participants with an ECG now says the domain is absent rather than printing a blank. Every figure was measured twice by routes that share no code.
You approved this on 27 August as its own candidate. That approval does not carry to the consolidated pull request, so this part is reviewed again inside it; the content is unchanged and the demo still applies. Two things are new since your approval.
Three census bugs found on 11 September, still open and unmilestoned on the 14th. The Input_* steps fail on an integer-keyed subject domain (gsm.safety#89), Input_Deaths errors on an empty subject domain instead of publishing no figure (#90), and the report can render an NA heading and label the page with a visit instead of the study (#91). None affects the bundled study's figures, and the review rounds since fixed neighbouring defects in the same code (a domain without the ID column, a group column lost on normalisation, empty output names). Decide whether v1.2.0 ships with the three filed for the next release, or whether they land on dev first.
The bundled study moved again under the package. Every census figure was qualified on gsm.core 1.3.1; gsm.core's main, which the package's Remotes track, now carries a different bundled study (758 enrolled, 12 deaths). The pinned assertions skip on that version with both numbers named, so a re-qualification is a deliberate act; the notes say so under Worth knowing about the census. Nothing to decide, worth knowing before you read 13 as a live number.
Re-read the moving figures once. The demo's before-and-after table is the clinical content. If your 27 August approval already covered it, a skim is enough.
Two questions still open from the design, not blocking the tag. Data completeness, the thirteenth metric, is not in this release because the lab domain carries no visit column under the standard mapping (gsm.safety#58). The disagreement metric you approved as D0023.1 was argued as reading one and reads thirteen of thirteen on the current study; the session asked you to re-decide it before it is written. Neither needs an answer to ship, but both are waiting on you.
on dev since 2026-08-22its own candidate, gsm.safety#69, approved 27 Aug and closed as superseded 13 Septhat approval does not carry to gsm.safety#88
gsm.safety v1.2.0, parts 3 and 4 · the last two widgets, and the FDA thresholds · gsm.safety#88, the candidate itself · on the merged demo
The FDA safety guide's thresholds arrive in R as tested data
Phase 0 of the FDA Standard Safety Tables and Figures work. The guide's Appendix Tables 56 to 60 become two package datasets, transcribed row for row with the printed criterion kept beside the parsed number and the guide page on every row: 46 abnormality-level rows and 40 extreme-value parameters in both unit systems. Three functions apply the guide's first rules to a long lab table and hand it back with the answers added: the result as a multiple of its upper limit of normal, the level 1, 2 or 3 abnormality grade with the criterion that fired, and the flag the extreme-value exclusion reads. A requirement matrix keys all 22 of the guide's figures and 16 cross-cutting rules so a static rendering here and its interactive twin in safety.viz cite one ID. An alignment note marks 51 ADaM columns as found, mapped, derived or missing and ends in the column contract phase 1 codes against.
Part 3 rides in the same diff and never had a candidate of its own: Widget_TimeToEvent and Widget_ParticipantProfile, the last two safety.viz renderers without an R binding, the widget contract generalised to more than one data frame so that the next two-frame renderer needs no change to the machinery, an example ADSL, a test that a binding both constructs and feeds its module, and the workflow fix that makes the release publish its artifacts. Its walkthrough is the notes section until the merged demo page exists.
The graded liver panel on the demo page. Every liver result in the example data through the ULN-multiple and grading functions, with the Table 56 cut lines read from the dataset rather than typed. Bilirubin carries 30 level 3 results and alkaline phosphatase 12; these are what the eDISH quadrants in phase 1 will be drawn from.
One row against the PDF. The demo's try-it step: open page 121 of the guide and compare any Table 57 row with the dataset. The transcription is the thing most worth a human spot check, because everything later reads from it.
Where the transcription had to read. Hemoglobin's level 1 criteria are printed as ranges and stored as the upper bound; the platelet unit is kept as printed; three hematology thresholds printed without a comparator are read as greater-than. Each is a Note on its row. Agree with the readings or ask for a change; a correction is one edit in data-raw.
What the grading refuses to do. A result in a unit the guide does not print is left NA and named in a message rather than misgraded; sex-qualified and baseline-qualified rows need columns the functions do not yet take. Read the message block on the demo and decide whether that behaviour is what you want phase 1 to inherit.
The review gate, as it now stands. The twelve-finding local review of 11 September (seven fixed, five filed) and then fifteen Copilot review rounds through the 13th, every thread resolved: about twenty hardening fixes landed on dev as increments, each a bullet under Also in this release (BuildWidgetPayload, the census report and the Derive_* functions refusing malformed settings, empty frames, colliding column names and missing output paths). None moves a figure. This is the gate you set on 11 September in place of ultrareview; skim the list once and confirm it satisfies you.
Three findings filed rather than fixed, for the release after this one. gsm.safety#147 (the provenance report flattens a coverage metric that publishes several groups), #155 (the census qualification script halts with a base error on a newer gsm.core), and #148: the R check workflow exports the App installation token into the environment where a pull request's own code runs. The session says #148 belongs to the upstream gsm.utils template and needs you to carry it there. Decide whether it is acceptable to tag with that open; for a single-owner repository whose pull requests are all yours or your sessions', it is a real finding with a small blast radius.
The definition of done, item by item. Datasets exported with every row of Tables 56 to 60; three functions with worked examples; 22 figure rows in the matrix; every alignment column marked with a decision; the NEWS section; the candidate open with its demo. The closing comment on the requirement (#9) claims each with its proof. The pkgdown reference will show the new functions only after the tag.
The two widgets, part 3. Read part 3 of the merged demo: what each widget takes and draws, the painted-pixel proof, and the try-it steps against the safety.viz twins. Two things to confirm: the parity allowlist is empty, which is the point of the parity requirement, and the time-to-event widget carries the experimental marking you attached to its safety.viz twin when you approved v1.7.0.
Checks. The body reports devtools::check() clean and the four required checks green on every increment; the Checks tab on the head commit is the confirmation, and on dev the workflow fix is present so all workflows actually run.
On approval of the one candidate
dev merges to main, v1.2.0 is tagged with the merged notes as its body, the pkgdown site rebuilds with four new widgets, the census report, the new datasets and functions, and the four requirements (#164, #274, #165, #9) close with their proofs. Phase 1 (#323, four static engines and twelve figures) starts from the alignment note's column contract.
opened as a draft 2026-09-11marked ready the same day after you dropped the ultrareview gateconsolidated and retitled 2026-09-12head 69a142d, 13 Sep, checks green