Decision · D0032 · 2026-09-02 · goal #112

One study, then the whole report

The recommendation in one paragraph: the assessment is right about the one thing that matters — the assembled report describes two different studies — and wrong about nothing that changes the plan. Adopt its build order as four releases. The first, "one study", is shovel-ready: nine issues, about five working sessions, every one of them written out below. Do the first two before the talk whatever September holds, because the public demo currently contradicts itself and that is the kind of thing a questioner finds.

Reviewed: a replication-readiness assessment dated 1 SeptemberChecked against: open.csr dev at b9ecbfa, the pilot's own data on both lanes, the reference report at its pinned SHA-256Response to send: response.md

What he decided

He answered the page in chat the morning it went up, in one line covering all four calls, and asked for the build to start.

@jwildfire · 2026-09-02 · in chat to obot-prime

“I'm good with the 4 calls. go ahead with the v0.4 build now. LMK when the RC (with standard demo demo/news artifacts) are ready for my review.”

Recorded the same morning, out of the session that received it. The four calls he approved are the four published on this page when he read it, each with its recommendation: R1 the four-release roadmap, R2 the data default flipped to the pilot's own package, R3 v0.4.0 re-scoped to one study plus the seven missing displays with the SAP moved to v0.5.0, R4 the build proceeding now. His instruction goes beyond R4's recommendation: the whole first release is to be built, not the first two sessions only, and its release candidate is the thing he reviews.

The report tells two versions of the study — confirmed from the data, not the document

Verified, not relayed

Preparing the subject-level dataset on each of the two lanes open.csr can read gives two answers to "how many people were on each arm". The pilot's own package puts 86, 84 and 84 subjects on placebo, low dose and high dose, by planned and by actual treatment alike. The pharmaverse re-derivation agrees on planned treatment and then moves twelve subjects from high dose to low dose in the actual-treatment column, giving 86, 96 and 72. The assembled report reads both.

DisplaysLaneGrouped byColumn counts
Demographics, disposition, exposure, adverse-event overview, common adverse events, serious-event listing — the six the project started withpharmaverse re-derivationactual treatment86 / 96 / 72
Vital signs, vital-sign change, weight, concomitant medicationspharmaverse re-derivationplanned treatment86 / 84 / 84
Populations, end of studypilot's own packageplanned treatment86 / 84 / 84
The thirteen efficacy displays and the figurepilot's own packageactual treatment86 / 84 / 84 (efficacy set within it)

The project measured the twelve-subject divergence on 27 August and committed it as a fact about the data. What it did not do was follow it into the document. The assessment's real contribution is naming where it bites: not in any one display, but in a report that says one thing in section 10 and another in section 14. Every other finding on the page is downstream of this one, and the plan below starts here. source-agreement.json · design decision D12

What the assessment gets right, and where it needs a correction

Taken finding by finding. Green means adopt as written; amber means the finding stands but the reason or the fix is different from the one it gives; red means it is not so. There is one red, and it is a coding convention, not a defect.

Four releases

The assessment's five phases map onto four releases, each of which ends in something a person can open. The order is the assessment's order and the reason is the assessment's reason: nothing polished gets built on data that disagrees with itself.

v0.4.0 · one study · about five sessions · assessment P0 and P1

Every display reads the pilot's own package, and the build fails if any two disagree

Proves: the report describes one study, and all thirty-one of the reference's post-text displays have a counterpart.

  • A study model — the arms, their order and labels, the analysis sets, the cut-off, the reference profile — as one file the specifications read, replacing the treatment vocabulary hard-coded in the pipeline.
  • The pilot's own package as the default source for every dataset; the pharmaverse re-derivation kept as the measured alternate and nothing else.
  • A consistency gate in the assembler: every placed display's column counts agree with the study model for its analysis set, or the document does not build.
  • The six original displays regenerated on the new lane as recorded iterations; demographics and exposure taken to the reference's shape; the two incidence tables, subjects by site and the six laboratory families added; the in-text tables and the disposition figure as variants of what exists.
  • Medications derived from this study's own SDTM; the Kenward-Roger spike; the stale sentence in section 11.1 corrected under review.
  • Gate for the release: the thirteen existing CI checks, the new consistency gate, and a reference-agreement record for every rebuilt or added display. Identity coverage 31 of 31 is the done line; cell agreement is measured and stated, not promised.

v0.5.0 · the whole text · assessment P2 · needs the review layer first

Every section of the report has approved prose that cites a checked result or a cited source

Proves: the narrative is complete, and every non-numeric claim in it can be traced to a page of a source document.

  • A source-and-fact registry: each input document with checksum, version, rights and the facts taken from it with page references; a stated precedence when documents disagree.
  • The SAP as the library's other face — the shells decision as designed, plus the methods text bound to the plan's sections through the registry. D0030
  • Front matter, ethics and administration, study plan, the efficacy narrative section 11.4 that is declared and empty today, the safety narrative, conclusions and references — generated where a result carries them, drafted for approval where a person must.
  • The eighteen synopsis blocks and five report blocks nobody has read, read. This is the release's real cost and it is his time, which is why the review layer comes first. D0029
  • Gate: no unapproved block in any placed section; every claim resolves to a result or a registry fact; a completeness check against the profile's required sections.

v0.6.0 · the submission document · assessment P3

A paginated, bookmarked PDF with a generated contents page and the appendices it is allowed to carry

Proves: the document model can produce the artefact a reviewer actually opens.

  • The CDISC Pilot profile as a fifth template object: the reference's section order, its 14-x.yy numbering, medications before safety conclusions.
  • Paged rendering in the existing build: page numbers and labels, running headers with study and section, contents and lists of tables and figures generated from the model, long tables split with repeated headers, landscape sections, PDF bookmarks. The display package assembled through r2rtf as the statistician's route.
  • An appendix manifest: each source file registered with rights, sensitivity and order; merged without touching rotation or links; bookmarked; cited rather than embedded where rights do not allow embedding. For this study that means the protocol and case report forms are cited, not copied.
  • Gate: the PDF renders in CI; every internal link and bookmark resolves; every manifest entry is present or explicitly waived; no embedded page lacks a rights record.

v1.0.0 · the replication proven · assessment P4

The acceptance checklist as a CI run anyone can re-execute

Proves: the claim "open.csr reproduces the CDISC pilot report" is a build that passes, not a sentence.

  • Thirty-one of thirty-one displays with agreement records; in-text claims checked against results and registry facts; required sections present and approved; appendices present or waived; one data source throughout; PDF render, navigation, privacy and rights checks all green.
  • A validation dossier generated from the evidence stream — the design's own version-one promise.
  • Version 1.1, named now so nobody mistakes 1.0 for it: the general intake path — a new-study loader, the terminology registry, the deviation log, governance — proved by running a second study, not by describing one.
Assessment phaseReleaseWhat changed in the mapping
P0 — make it tell one studyv0.4.0Unchanged. It is the first gate.
P1 — finish and correct the resultsv0.4.0, laboratory block may slip to v0.4.1Combined with P0, because rebuilding the four displays on the new lane and adding the seven missing ones is the same working session's shape.
P2 — write the full report from traceable factsv0.5.0Sequenced after the review layer; the SAP joins it.
P3 — build the submission documentv0.6.0Unchanged.
P4 — prove the replicationv1.0.0Unchanged; the general intake is explicitly after it.

What fits before the talk

This depends on two answers still open on the September plan: whether the talk is in September or October, and whether the month goes to open.gismo. This page does not re-ask them. It says what open.csr should get under each outcome. D0031

Whatever September holds

Do the first two issues below — the study model with its consistency gate, and the lane flip with the regenerated six — before the talk. Two sessions. Until they land, the public demo the talk will point at contains a section that says 96 people took the low dose and another that says 84. That is not a gap; it is an error a reader can find in a minute, and the fix is cheap.

The first release, as issues

Nine issues in dependency order. Each is written so it can be filed as-is once the milestone is confirmed — nothing starts on an issue without one. Each lands as one pull request to dev on the standard lane; the release candidate is the one thing he reviews. Sizes are working sessions of the kind the last release was built from.

Before anything · housekeeping, not an issue

Fold the 0.3.0 release back into dev

The release branch bumped the version, dropped the "Upcoming" suffix from the release log and regenerated the evidence; main has all of it and dev has none of it — forty-three files, all stamps. A pull request from main into dev, merged on the standard lane. The stacked-release lesson from the R package applies: merge forward the same day, or the next release candidate collides on the version bump.

Size minutes · Needs from him nothing

Issue A · session 1

One study: a study model, and a build that fails when displays disagree about it

Requirement

The study's arms, their order and labels, its analysis sets and its cut-off are declared once, in one file every specification reads; and the assembled document refuses to build when any two placed displays report different subject counts for the same arm and analysis set.

Design

  • library/study.yaml: study id, title, arms with order and labels, analysis sets with the flag that defines each, cut-off, the reference profile id. The treatment vocabulary now hard-coded as trt_levels() reads it. This is the file the SAP decision's second question wanted and did not have. D-SAP2
  • A treatmentConsistency gate in scripts/assemble.mjs beside the existing structure, binding, numeric-fidelity and approval gates: for every placed display, the per-arm N its analysis results dataset reports must equal the N the study model derives for that analysis set on the default lane; and the set of Ns across the document must be one set per analysis set. Any disagreement is a build error naming both displays.
  • The assembled document records which lane every display read, so the answer to "one source?" is on the page, not in a script.

Proof it can fail

  • A test pins one display to the pharmaverse lane and asserts the gate goes red naming it. Run before the gate is trusted; the check that cannot fail is not a check.
  • Requirements STD-MODEL-* and STD-GATE-* in the matrix, verified by tests that read the study model directly.

Files library/study.yaml · pipeline/R/data-prep.R · scripts/assemble.mjs · tests · quality/requirements · Depends on nothing · Size one session · Needs from him R2 answered

Issue B · session 1 to 2

The pilot's own package feeds every display

Requirement

The default source registry serves every dataset from the pilot's own ADaM package; the pharmaverse re-derivation remains readable, measured and named as the alternate, and no committed display reads it.

Design

  • Design decision D12 revised in place, dated: the "every domain pharmaverse already served stays there" rule was a no-regressions rule for six displays, and the document is the thing that regresses under it. ADEX is dropped from the registry; no display needs it.
  • The six original displays regenerate on the new lane as recorded iterations with the change request stated — this is the closed loop the project exists to demonstrate, used on itself.
  • The four vital-sign and medication displays move lanes too. Their baseline and end-of-treatment derivations are the project's own and should be lane-independent, but the two packagings differ on the derived vital-sign records ADVS carries, so the 1,341-statistic three-route agreement is the test, and it is allowed to fail and be investigated.
  • qc/source-agreement.R re-measured and its record re-committed; the assembly's scope note rewritten, since it currently says the safety displays are built from the pharmaverse lane.

Proof it can fail

  • The consistency gate from Issue A, now green on the document; the reference-agreement records for populations, end of study, vital signs and medications still green.
  • The demographics column counts read 86, 84, 84 in the committed output, asserted by test.

Files pipeline/R/data-prep.R · qc/vendor-phuse-data.R · library/tfl/<six>/iterations.yaml · library/templates/ich-e3/assembly.yaml · docs/design/design.md · Depends on A · Size one session · Needs from him R2

Issue C · session 2

Demographics and exposure in the report's own shape

Requirement

The demographics and exposure displays reproduce the reference's Tables 14-2.01 and 14-4.01 cell for cell, with every definition the project supplies stated on the display.

Design

  • Demographics: intent-to-treat set, a Total column, a p-value column (one-way ANOVA for continuous rows, Pearson's chi-square for categorical), the reference's rows — age and its three groups, sex, race-with-origin, MMSE, disease duration, education, and the baseline measures. Race-with-origin is a footnoted recode of race and ethnicity into the 2006 convention. The FDA-standard shape stays as a variant.
  • Exposure: average daily dose and cumulative dose from the subject-level dataset's own columns, for completers at week 24 and the safety population side by side — two column groups from one specification, which the display model supports.
  • Reference-agreement records for both, transcribed from the pinned PDF at pages 46 to 48 and 62, with the transcription self-check the populations record already has.

Proof it can fail

  • Three routes per display, self-tested by perturbation, as for the disposition tables. A recode that produces 217 instead of 218 Caucasian is a build failure.

Files library/tfl/t-demographics · library/tfl/t-exposure · qc/reference-report-agreement.R (extended) · quality/data · Depends on B · Size half a session · Needs from him nothing

Issue D · session 3

Adverse-event and serious-event incidence as the report prints them

Requirement

Two new displays reproduce Tables 14-5.01 and 14-5.02: system organ class and preferred term, subjects with events and a count of events in brackets, Fisher's exact p-values for placebo against each dose.

Design

  • A hierarchy display over system organ class and preferred term, which the shell measurement showed four existing displays already declare; a fisher_pairwise analysis method against a named reference arm; the pilot's own treatment-emergent flag.
  • The reference's incidence table is sixteen pages; agreement uses the mechanical extraction route the efficacy record uses, not hand transcription.
  • The existing overview and common-events tables keep their FDA identifiers and are not placed by the pilot profile.

Proof it can fail

  • All-body-system row: 65, 77, 76 subjects with 281, 412, 433 events, p-values 0.007 and 0.014 — the first cells the record checks.

Files library/tfl/t-ae-incidence (new) · library/tfl/t-sae-incidence (new) · pipeline/R/methods · qc/efficacy-reference.R pattern reused · Depends on B · Size one session · Needs from him nothing

Issue E · session 3

Subjects by site, the in-text tables, and the disposition figure

Requirement

Table 14-1.03 exists; the five in-text tables and the disposition figure the reference's body carries are variants of displays the library already has, placed in the sections that discuss them.

Design

  • Subjects by site: intent-to-treat, efficacy and completer counts by pooled and actual site, from the subject-level dataset's site and pooled-site columns. Small.
  • In-text Tables 11-1, 12-1, 12-3 and 12-4 are in-text variants of demographics, incidence, vital-sign change and weight — one analysis results dataset, two renderings, by design decision D6. Table 12-2, the special-interest category, is a small filter variant of incidence.
  • Figure 10-1, the disposition flow, is a new figure type from the subject-level dataset: randomised, treated, completed, discontinued by reason. Figure 9-1, the study schema, is a drawing and belongs to the template as an asset, in the document release.

Files library/tfl/t-subjects-by-site (new) · library/tfl/f-disposition (new) · display.yaml variants · library/templates/ich-e3/assembly.yaml · Depends on C, D · Size half a session · Needs from him nothing

Issue F · sessions 4 to 5 · may slip to v0.4.1

The six laboratory families

Requirement

Tables 14-6.01 to 14-6.06 exist and agree with the reference: continuous summaries by parameter and visit; beyond-range frequencies; clinically-significant-change frequencies; shifts by visit; shifts overall; Hy's law shifts.

Design

  • Chemistry and haematology from the two vendored laboratory datasets; Hy's law from the third. Two new analysis methods — a shift table over baseline-by-post categories, and an abnormality frequency with a chi-square column — and one new display shape, the parameter-by-visit block, which the shift-by-visit table needs at thirty-three reference pages.
  • Order of work: continuous summaries and the two frequency tables first, because they share a parameter loop; the shift family second; Hy's law last, because it is one page.
  • Agreement by mechanical extraction; the reference's layout breaks parameter names across lines, so the extractor needs the same care the incidence table needed.

Why it may slip

It is the largest item in the release by a wide margin and the only one that does not bear on "one study". If the release is cut at four sessions, this ships as v0.4.1 with the same gate.

Files library/tfl/t-lb-* (six new) · pipeline/R/methods · qc · Depends on B · Size two sessions · Needs from him nothing

Issue G · session 5

Concomitant medications from this study's own SDTM

Requirement

The medication analysis dataset is derived from the pilot's own SDTM medications domain, with the derivation on record, and the current remapped dataset is retired only after the derived one reproduces every cell of the medication table.

Design

  • Vendor cm.xpt from the same repository at the same pinned commit, with the blob-verified provenance every vendored file already carries; a derive_adcm() in the tested data-preparation layer producing the columns the display reads; a derivation note in the style of an analysis data reviewer's guide.
  • Switch-over rule: the derived dataset and the remapped one must agree on every publishable statistic in the medication table's analysis results dataset before the registry points at the new one. Then the remapped file and its reversal code are removed — with his approval, since nothing is deleted without it.

Files qc/vendor-phuse-data.R · pipeline/R/derive-adcm.R (new) · pipeline/inst/extdata · Depends on B · Size half a session · Needs from him approval to remove the remapped file

Issue H · half a session · bounded spike

The five last digits: Kenward-Roger

Requirement

The repeated-measures display reproduces all 769 published figures, or the residual is measured and stated.

Design

  • Refit with the mmrm package using Kenward-Roger degrees of freedom and its adjusted covariance, keeping the identified fit — the same observations, the six covariance parameters to six significant figures, the REML criterion — as the first assertion.
  • Pass mark: 769 of 769. On failure, the record states which cells still differ and the display's footnote stays. The package adds a compiled dependency; the CI install cost is part of what the spike measures.

Files pipeline/R/mmrm.R · pipeline/DESCRIPTION · quality/data/efficacy-reference.json · Depends on nothing · Size half a session · Needs from him nothing

Issue I · his ten minutes

Section 11.1 stops saying there is no efficacy data

The approved block for section 11.1 still reads "no efficacy analysis set is defined for this report" and explains that the data source contains no efficacy dataset. It has been false since 27 August, and the release notes say so twice. It is a parameterised block under approval, so the change is his to approve, not the pipeline's to make: a one-paragraph rewrite, reviewed in the Text view, approved in the frontmatter. The assembly's scope note changes with it in Issue B.

Files library/text/TXT-E3-1101.md · Depends on B, so the new sentence is true of the new lane · Size minutes · Needs from him the approval

How it ships

  • Milestone v0.4.0 carries all nine; the SAP issue moves to v0.5.0 if R3 is yes. A hub requirement for the release, in the usual three parts, with the nine as its sub-issues.
  • Each issue is one pull request to dev on the standard lane. The release candidate release/v0.4.0 is the single thing put in front of him, with the demo tier frozen at the tag as the last release was.
  • The release note leads with the sentence the assessment could not write: the report describes one study, and all thirty-one displays have a counterpart.

The four calls

R1

Adopt the four-release roadmap?

v0.4.0 one study; v0.5.0 the whole text, after the review layer; v0.6.0 the submission document; v1.0.0 the replication proven, with the general intake named as 1.1.

Recommendation: yes. It is the assessment's own order with one change — the SAP moves to the text release — and every release ends in something he can open.

R2

Flip the data default so the pilot's own package feeds every display?

Revises design decision D12. The pharmaverse re-derivation stays readable and measured, as the alternate. No display reads it.

Recommendation: yes. The rule that kept the six original displays where they were was protecting them from change; it turned out to be protecting the document from agreeing with itself. The last technical reason — exposure needing ADEX — does not hold.

R3

Re-scope v0.4.0 to one study plus the seven missing displays, and move the SAP to v0.5.0?

The milestone today holds the SAP shells and a scroll bug. This puts the nine issues above in it and moves the SAP to the release where the fact registry gives it something to bind to.

Recommendation: yes. The SAP page's five questions are still open, so nothing already in motion is displaced; and a plan-of-record built on a report that disagrees with itself would be the polished thing on conflicting data the assessment warns about.

R4

Before the talk: the two sessions, or the September week?

Issues A and B alone, two sessions, so the public demo stops contradicting itself; or the whole first release at about five, which needs September's answer to go to open.csr.

Recommendation: the two sessions regardless, and the week only if S2 sends it here. The September plan's argument for open.gismo still stands. The two sessions are not a bid for the month; they are the cost of not being wrong in public.

Measured for this page, not relayed