What he decided
He answered the page in chat the morning it went up, in one line covering all four calls, and asked for the build to start.
“I'm good with the 4 calls. go ahead with the v0.4 build now. LMK when the RC (with standard demo demo/news artifacts) are ready for my review.”
Recorded the same morning, out of the session that received it. The four calls he approved are the four published on this page when he read it, each with its recommendation: R1 the four-release roadmap, R2 the data default flipped to the pilot's own package, R3 v0.4.0 re-scoped to one study plus the seven missing displays with the SAP moved to v0.5.0, R4 the build proceeding now. His instruction goes beyond R4's recommendation: the whole first release is to be built, not the first two sessions only, and its release candidate is the thing he reviews.
The report tells two versions of the study — confirmed from the data, not the document
Verified, not relayed
Preparing the subject-level dataset on each of the two lanes open.csr can read gives two answers to "how many people were on each arm". The pilot's own package puts 86, 84 and 84 subjects on placebo, low dose and high dose, by planned and by actual treatment alike. The pharmaverse re-derivation agrees on planned treatment and then moves twelve subjects from high dose to low dose in the actual-treatment column, giving 86, 96 and 72. The assembled report reads both.
| Displays | Lane | Grouped by | Column counts |
|---|---|---|---|
| Demographics, disposition, exposure, adverse-event overview, common adverse events, serious-event listing — the six the project started with | pharmaverse re-derivation | actual treatment | 86 / 96 / 72 |
| Vital signs, vital-sign change, weight, concomitant medications | pharmaverse re-derivation | planned treatment | 86 / 84 / 84 |
| Populations, end of study | pilot's own package | planned treatment | 86 / 84 / 84 |
| The thirteen efficacy displays and the figure | pilot's own package | actual treatment | 86 / 84 / 84 (efficacy set within it) |
The project measured the twelve-subject divergence on 27 August and committed it as a fact about the data. What it did not do was follow it into the document. The assessment's real contribution is naming where it bites: not in any one display, but in a report that says one thing in section 10 and another in section 14. Every other finding on the page is downstream of this one, and the plan below starts here. source-agreement.json · design decision D12
What the assessment gets right, and where it needs a correction
Taken finding by finding. Green means adopt as written; amber means the finding stands but the reason or the fix is different from the one it gives; red means it is not so. There is one red, and it is a coding convention, not a defect.
- It reviewed version 0.2.0. The current release is 0.3.0, from 27 August. The counts in the assessment — twenty-six displays, thirteen efficacy tables, the figure — are the 0.3.0 contents, so its substance is current. The stamp is stale because the release's version bump was never folded back into
dev, which a checkout ofdevstill reports as 0.2.0. That is the project's housekeeping fault and it is the first item below.Effect on the plan: none, beyond folding the release forward. - Two versions of the study in one report. Confirmed above from the data. The assessment's fix — one source, one assignment, a build that fails on disagreement — is the right one, and it is the whole of the first release.Effect on the plan: it is the plan's first gate.
- Race is not a data conflict. The assessment reads "218 Caucasian plus 12 Hispanic" against "230 White" as the two data copies disagreeing. Both copies say the same thing: 230 White, 23 Black, 1 American Indian, and separately 12 Hispanic by ethnicity. The 2006 report printed a combined "Race (Origin)" column with Hispanic as a category, so 218 plus 12 is 230. It is a documented recode of two variables into one, footnoted on the display, in the way the project already handles the Complete Study row. No source-priority rule is needed.Effect on the plan: a footnote in the demographics rebuild, not a decision.
- Exposure needs no exposure dataset. The pilot's package has no ADEX, and design decision D12 kept the pharmaverse lane partly because it was "the only source of ADEX". The reference's exposure table summarises average daily dose and cumulative dose, both of which the pilot's own subject-level dataset carries as columns. Exposure moves lanes with no derivation at all.Effect on the plan: the last reason to keep the pharmaverse lane as a default goes away.
- "Rebuild four" is really "move two, add two". Demographics and exposure exist in nearly the reference's shape and need the lane flip plus small additions — the reference's demographics is on the intent-to-treat set with a p-value column and rows for MMSE, education and disease duration. The two adverse-event tables are a different shape entirely: system organ class and preferred term, subject counts with event counts in brackets, Fisher's exact p-values against placebo. Those are new displays. The existing FDA-standard adverse-event tables stay in the library as variants; the pilot profile simply does not place them.Effect on the plan: two issues instead of one, and nothing is thrown away.
- The seven missing displays are the seven it names. Subjects by site, and six laboratory families: continuous summaries, beyond-range frequencies, clinically-significant-change frequencies, shifts by visit, shifts overall, and Hy's law shifts. All three laboratory datasets are already vendored. The by-visit shift table is thirty-three pages in the reference and is the largest single display in the whole set.Effect on the plan: the laboratory block is the one item in the first release allowed to slip to a point release.
- The narrative gap is a review gap, not a writing gap. Ten of fifteen report blocks are approved; the other five have been agent-drafted since July and nobody has read them. The synopsis has eighteen more in the same state. Writing the missing sections is cheap; the constraint is one person's reading time, and the review layer the programme designed on 27 August is the dependency the assessment does not name.Effect on the plan: the text release is scheduled after the review layer, not before.
- The SAP is both an input and an output, and the library is the join. The assessment wants the SAP ingested and every method bound to a section of it. The v0.4.0 plan on the books makes the SAP an output — shells rendered from the same display specifications the report renders from. These are the two ends of one model: the specification library is the machine-readable plan; for a study open.csr runs from the start, the SAP document is rendered from it; for a study that arrives with a SAP, that document is ingested and reconciled against it. Neither replaces the other. The pilot's own SAP is not vendored because its redistribution rights are unestablished, which the SAP decision already says. D0030Effect on the plan: the SAP moves to the text release, where the fact registry gives it something to bind to.
- The five last-digit differences are what it says they are. Kenward-Roger degrees of freedom and the adjusted standard errors are not implemented, by a recorded scope decision. The route to close them is the
mmrmpackage, which implements both, not recovered SAS output. It is a bounded spike with a numerical pass mark.Effect on the plan: half a session, with permission to fail. - The concomitant-medication lineage is exactly as described. The dataset in use is a remapped copy from a differently named folder, with half its subjects relabelled into a synthetic second study and the relabelling reversed. The study's own SDTM medication domain is in the same public repository at the same pinned commit. Derive from it, with the derivation on record.Effect on the plan: one issue, with a cell-for-cell agreement check against the current table before the switch.
- Rights and privacy before any appendix is copied. Agreed, and the project already refuses harder than the assessment asks: pages 154 to 409 of the reference carry a third party's copyright, and the agreement record declines to vendor the PDF at all for that reason. An appendix builder for this study registers and cites those pages; it embeds only what is the project's own or openly licensed.Effect on the plan: the appendix manifest is rights-first by construction.
- A study-specific report profile is one file, not a feature. The assessment asks for a CDISC Pilot profile with the reference's section order and table numbers. The template library is already plural and numbering is assigned at build time by design decision D6, so the profile is a fifth template object declaring the reference's structure — no engine change. Its ordering of medications before safety conclusions is a declaration in that file.Effect on the plan: a small item in the document release.
- HTML is not a submission document. Agreed. The route proposed below is paged rendering of the existing document model with page numbers, running headers, generated contents and lists, landscape sections and bookmarks — in the same build the site already runs — plus the display package assembled through
r2rtffor the statistician's review. Word output is a second target, not the first.Effect on the plan: the document release's core. - Pixel recreation is optional. Stronger: it is not a target. The agreement checks compare cell strings, which is the right unit; the 2006 listing layout is not something to reproduce.Effect on the plan: none.
- The study-team intake table is right, and most of it is after version one. The division of responsibility — software calculates, organises, traces, drafts and checks; people supply the truth about the study and approve the interpretation — is the project's own position. Three of its rows are near: a study model in the first release, a source-and-fact registry in the second, an appendix and rights manifest in the third. The rest is proved by running a second study, and nothing before that should claim it.Effect on the plan: the intake path is version 1.1, stated as such.
Four releases
The assessment's five phases map onto four releases, each of which ends in something a person can open. The order is the assessment's order and the reason is the assessment's reason: nothing polished gets built on data that disagrees with itself.
v0.4.0 · one study · about five sessions · assessment P0 and P1
Every display reads the pilot's own package, and the build fails if any two disagree
Proves: the report describes one study, and all thirty-one of the reference's post-text displays have a counterpart.
- A study model — the arms, their order and labels, the analysis sets, the cut-off, the reference profile — as one file the specifications read, replacing the treatment vocabulary hard-coded in the pipeline.
- The pilot's own package as the default source for every dataset; the pharmaverse re-derivation kept as the measured alternate and nothing else.
- A consistency gate in the assembler: every placed display's column counts agree with the study model for its analysis set, or the document does not build.
- The six original displays regenerated on the new lane as recorded iterations; demographics and exposure taken to the reference's shape; the two incidence tables, subjects by site and the six laboratory families added; the in-text tables and the disposition figure as variants of what exists.
- Medications derived from this study's own SDTM; the Kenward-Roger spike; the stale sentence in section 11.1 corrected under review.
- Gate for the release: the thirteen existing CI checks, the new consistency gate, and a reference-agreement record for every rebuilt or added display. Identity coverage 31 of 31 is the done line; cell agreement is measured and stated, not promised.
v0.5.0 · the whole text · assessment P2 · needs the review layer first
Every section of the report has approved prose that cites a checked result or a cited source
Proves: the narrative is complete, and every non-numeric claim in it can be traced to a page of a source document.
- A source-and-fact registry: each input document with checksum, version, rights and the facts taken from it with page references; a stated precedence when documents disagree.
- The SAP as the library's other face — the shells decision as designed, plus the methods text bound to the plan's sections through the registry. D0030
- Front matter, ethics and administration, study plan, the efficacy narrative section 11.4 that is declared and empty today, the safety narrative, conclusions and references — generated where a result carries them, drafted for approval where a person must.
- The eighteen synopsis blocks and five report blocks nobody has read, read. This is the release's real cost and it is his time, which is why the review layer comes first. D0029
- Gate: no unapproved block in any placed section; every claim resolves to a result or a registry fact; a completeness check against the profile's required sections.
v0.6.0 · the submission document · assessment P3
A paginated, bookmarked PDF with a generated contents page and the appendices it is allowed to carry
Proves: the document model can produce the artefact a reviewer actually opens.
- The CDISC Pilot profile as a fifth template object: the reference's section order, its 14-x.yy numbering, medications before safety conclusions.
- Paged rendering in the existing build: page numbers and labels, running headers with study and section, contents and lists of tables and figures generated from the model, long tables split with repeated headers, landscape sections, PDF bookmarks. The display package assembled through
r2rtfas the statistician's route. - An appendix manifest: each source file registered with rights, sensitivity and order; merged without touching rotation or links; bookmarked; cited rather than embedded where rights do not allow embedding. For this study that means the protocol and case report forms are cited, not copied.
- Gate: the PDF renders in CI; every internal link and bookmark resolves; every manifest entry is present or explicitly waived; no embedded page lacks a rights record.
v1.0.0 · the replication proven · assessment P4
The acceptance checklist as a CI run anyone can re-execute
Proves: the claim "open.csr reproduces the CDISC pilot report" is a build that passes, not a sentence.
- Thirty-one of thirty-one displays with agreement records; in-text claims checked against results and registry facts; required sections present and approved; appendices present or waived; one data source throughout; PDF render, navigation, privacy and rights checks all green.
- A validation dossier generated from the evidence stream — the design's own version-one promise.
- Version 1.1, named now so nobody mistakes 1.0 for it: the general intake path — a new-study loader, the terminology registry, the deviation log, governance — proved by running a second study, not by describing one.
| Assessment phase | Release | What changed in the mapping |
|---|---|---|
| P0 — make it tell one study | v0.4.0 | Unchanged. It is the first gate. |
| P1 — finish and correct the results | v0.4.0, laboratory block may slip to v0.4.1 | Combined with P0, because rebuilding the four displays on the new lane and adding the seven missing ones is the same working session's shape. |
| P2 — write the full report from traceable facts | v0.5.0 | Sequenced after the review layer; the SAP joins it. |
| P3 — build the submission document | v0.6.0 | Unchanged. |
| P4 — prove the replication | v1.0.0 | Unchanged; the general intake is explicitly after it. |
What fits before the talk
This depends on two answers still open on the September plan: whether the talk is in September or October, and whether the month goes to open.gismo. This page does not re-ask them. It says what open.csr should get under each outcome. D0031
Whatever September holds
Do the first two issues below — the study model with its consistency gate, and the lane flip with the regenerated six — before the talk. Two sessions. Until they land, the public demo the talk will point at contains a section that says 96 people took the low dose and another that says 84. That is not a gap; it is an error a reader can find in a minute, and the fix is cheap.
- If open.gismo gets the month, as the September plan recommends: those two sessions and nothing else. Everything from the third issue on is the first post-talk release.
- If open.csr gets a week: the whole of v0.4.0 fits, laboratory block included, at five sessions of the kind that produced v0.3.0.
- Under no outcome does anything from the text or document releases start before the talk. Both need his reading time, which the rehearsal week is for.
The first release, as issues
Nine issues in dependency order. Each is written so it can be filed as-is once the milestone is confirmed — nothing starts on an issue without one. Each lands as one pull request to dev on the standard lane; the release candidate is the one thing he reviews. Sizes are working sessions of the kind the last release was built from.
Before anything · housekeeping, not an issue
Fold the 0.3.0 release back into dev
The release branch bumped the version, dropped the "Upcoming" suffix from the release log and regenerated the evidence; main has all of it and dev has none of it — forty-three files, all stamps. A pull request from main into dev, merged on the standard lane. The stacked-release lesson from the R package applies: merge forward the same day, or the next release candidate collides on the version bump.
Issue A · session 1
One study: a study model, and a build that fails when displays disagree about it
Requirement
The study's arms, their order and labels, its analysis sets and its cut-off are declared once, in one file every specification reads; and the assembled document refuses to build when any two placed displays report different subject counts for the same arm and analysis set.
Design
library/study.yaml: study id, title, arms with order and labels, analysis sets with the flag that defines each, cut-off, the reference profile id. The treatment vocabulary now hard-coded astrt_levels()reads it. This is the file the SAP decision's second question wanted and did not have. D-SAP2- A
treatmentConsistencygate inscripts/assemble.mjsbeside the existing structure, binding, numeric-fidelity and approval gates: for every placed display, the per-arm N its analysis results dataset reports must equal the N the study model derives for that analysis set on the default lane; and the set of Ns across the document must be one set per analysis set. Any disagreement is a build error naming both displays. - The assembled document records which lane every display read, so the answer to "one source?" is on the page, not in a script.
Proof it can fail
- A test pins one display to the pharmaverse lane and asserts the gate goes red naming it. Run before the gate is trusted; the check that cannot fail is not a check.
- Requirements
STD-MODEL-*andSTD-GATE-*in the matrix, verified by tests that read the study model directly.
Issue B · session 1 to 2
The pilot's own package feeds every display
Requirement
The default source registry serves every dataset from the pilot's own ADaM package; the pharmaverse re-derivation remains readable, measured and named as the alternate, and no committed display reads it.
Design
- Design decision D12 revised in place, dated: the "every domain pharmaverse already served stays there" rule was a no-regressions rule for six displays, and the document is the thing that regresses under it. ADEX is dropped from the registry; no display needs it.
- The six original displays regenerate on the new lane as recorded iterations with the change request stated — this is the closed loop the project exists to demonstrate, used on itself.
- The four vital-sign and medication displays move lanes too. Their baseline and end-of-treatment derivations are the project's own and should be lane-independent, but the two packagings differ on the derived vital-sign records ADVS carries, so the 1,341-statistic three-route agreement is the test, and it is allowed to fail and be investigated.
qc/source-agreement.Rre-measured and its record re-committed; the assembly's scope note rewritten, since it currently says the safety displays are built from the pharmaverse lane.
Proof it can fail
- The consistency gate from Issue A, now green on the document; the reference-agreement records for populations, end of study, vital signs and medications still green.
- The demographics column counts read 86, 84, 84 in the committed output, asserted by test.
Issue C · session 2
Demographics and exposure in the report's own shape
Requirement
The demographics and exposure displays reproduce the reference's Tables 14-2.01 and 14-4.01 cell for cell, with every definition the project supplies stated on the display.
Design
- Demographics: intent-to-treat set, a Total column, a p-value column (one-way ANOVA for continuous rows, Pearson's chi-square for categorical), the reference's rows — age and its three groups, sex, race-with-origin, MMSE, disease duration, education, and the baseline measures. Race-with-origin is a footnoted recode of race and ethnicity into the 2006 convention. The FDA-standard shape stays as a variant.
- Exposure: average daily dose and cumulative dose from the subject-level dataset's own columns, for completers at week 24 and the safety population side by side — two column groups from one specification, which the display model supports.
- Reference-agreement records for both, transcribed from the pinned PDF at pages 46 to 48 and 62, with the transcription self-check the populations record already has.
Proof it can fail
- Three routes per display, self-tested by perturbation, as for the disposition tables. A recode that produces 217 instead of 218 Caucasian is a build failure.
Issue D · session 3
Adverse-event and serious-event incidence as the report prints them
Requirement
Two new displays reproduce Tables 14-5.01 and 14-5.02: system organ class and preferred term, subjects with events and a count of events in brackets, Fisher's exact p-values for placebo against each dose.
Design
- A hierarchy display over system organ class and preferred term, which the shell measurement showed four existing displays already declare; a
fisher_pairwiseanalysis method against a named reference arm; the pilot's own treatment-emergent flag. - The reference's incidence table is sixteen pages; agreement uses the mechanical extraction route the efficacy record uses, not hand transcription.
- The existing overview and common-events tables keep their FDA identifiers and are not placed by the pilot profile.
Proof it can fail
- All-body-system row: 65, 77, 76 subjects with 281, 412, 433 events, p-values 0.007 and 0.014 — the first cells the record checks.
Issue E · session 3
Subjects by site, the in-text tables, and the disposition figure
Requirement
Table 14-1.03 exists; the five in-text tables and the disposition figure the reference's body carries are variants of displays the library already has, placed in the sections that discuss them.
Design
- Subjects by site: intent-to-treat, efficacy and completer counts by pooled and actual site, from the subject-level dataset's site and pooled-site columns. Small.
- In-text Tables 11-1, 12-1, 12-3 and 12-4 are in-text variants of demographics, incidence, vital-sign change and weight — one analysis results dataset, two renderings, by design decision D6. Table 12-2, the special-interest category, is a small filter variant of incidence.
- Figure 10-1, the disposition flow, is a new figure type from the subject-level dataset: randomised, treated, completed, discontinued by reason. Figure 9-1, the study schema, is a drawing and belongs to the template as an asset, in the document release.
Issue F · sessions 4 to 5 · may slip to v0.4.1
The six laboratory families
Requirement
Tables 14-6.01 to 14-6.06 exist and agree with the reference: continuous summaries by parameter and visit; beyond-range frequencies; clinically-significant-change frequencies; shifts by visit; shifts overall; Hy's law shifts.
Design
- Chemistry and haematology from the two vendored laboratory datasets; Hy's law from the third. Two new analysis methods — a shift table over baseline-by-post categories, and an abnormality frequency with a chi-square column — and one new display shape, the parameter-by-visit block, which the shift-by-visit table needs at thirty-three reference pages.
- Order of work: continuous summaries and the two frequency tables first, because they share a parameter loop; the shift family second; Hy's law last, because it is one page.
- Agreement by mechanical extraction; the reference's layout breaks parameter names across lines, so the extractor needs the same care the incidence table needed.
Why it may slip
It is the largest item in the release by a wide margin and the only one that does not bear on "one study". If the release is cut at four sessions, this ships as v0.4.1 with the same gate.
Issue G · session 5
Concomitant medications from this study's own SDTM
Requirement
The medication analysis dataset is derived from the pilot's own SDTM medications domain, with the derivation on record, and the current remapped dataset is retired only after the derived one reproduces every cell of the medication table.
Design
- Vendor
cm.xptfrom the same repository at the same pinned commit, with the blob-verified provenance every vendored file already carries; aderive_adcm()in the tested data-preparation layer producing the columns the display reads; a derivation note in the style of an analysis data reviewer's guide. - Switch-over rule: the derived dataset and the remapped one must agree on every publishable statistic in the medication table's analysis results dataset before the registry points at the new one. Then the remapped file and its reversal code are removed — with his approval, since nothing is deleted without it.
Issue H · half a session · bounded spike
The five last digits: Kenward-Roger
Requirement
The repeated-measures display reproduces all 769 published figures, or the residual is measured and stated.
Design
- Refit with the
mmrmpackage using Kenward-Roger degrees of freedom and its adjusted covariance, keeping the identified fit — the same observations, the six covariance parameters to six significant figures, the REML criterion — as the first assertion. - Pass mark: 769 of 769. On failure, the record states which cells still differ and the display's footnote stays. The package adds a compiled dependency; the CI install cost is part of what the spike measures.
Issue I · his ten minutes
Section 11.1 stops saying there is no efficacy data
The approved block for section 11.1 still reads "no efficacy analysis set is defined for this report" and explains that the data source contains no efficacy dataset. It has been false since 27 August, and the release notes say so twice. It is a parameterised block under approval, so the change is his to approve, not the pipeline's to make: a one-paragraph rewrite, reviewed in the Text view, approved in the frontmatter. The assembly's scope note changes with it in Issue B.
How it ships
- Milestone v0.4.0 carries all nine; the SAP issue moves to v0.5.0 if R3 is yes. A hub requirement for the release, in the usual three parts, with the nine as its sub-issues.
- Each issue is one pull request to
devon the standard lane. The release candidaterelease/v0.4.0is the single thing put in front of him, with the demo tier frozen at the tag as the last release was. - The release note leads with the sentence the assessment could not write: the report describes one study, and all thirty-one displays have a counterpart.
The four calls
R1
Adopt the four-release roadmap?
v0.4.0 one study; v0.5.0 the whole text, after the review layer; v0.6.0 the submission document; v1.0.0 the replication proven, with the general intake named as 1.1.
Recommendation: yes. It is the assessment's own order with one change — the SAP moves to the text release — and every release ends in something he can open.
R2
Flip the data default so the pilot's own package feeds every display?
Revises design decision D12. The pharmaverse re-derivation stays readable and measured, as the alternate. No display reads it.
Recommendation: yes. The rule that kept the six original displays where they were was protecting them from change; it turned out to be protecting the document from agreeing with itself. The last technical reason — exposure needing ADEX — does not hold.
R3
Re-scope v0.4.0 to one study plus the seven missing displays, and move the SAP to v0.5.0?
The milestone today holds the SAP shells and a scroll bug. This puts the nine issues above in it and moves the SAP to the release where the fact registry gives it something to bind to.
Recommendation: yes. The SAP page's five questions are still open, so nothing already in motion is displaced; and a plan-of-record built on a report that disagrees with itself would be the polished thing on conflicting data the assessment warns about.
R4
Before the talk: the two sessions, or the September week?
Issues A and B alone, two sessions, so the public demo stops contradicting itself; or the whole first release at about five, which needs September's answer to go to open.csr.
Recommendation: the two sessions regardless, and the week only if S2 sends it here. The September plan's argument for open.gismo still stands. The two sessions are not a bid for the month; they are the cost of not being wrong in public.
Measured for this page, not relayed
- The two lanes were prepared through the project's own data-preparation function and tabulated by planned and actual treatment. The twelve-subject movement is in the actual-treatment column of the pharmaverse lane and nowhere else; race, sex, ethnicity, site and planned arm agree on every subject.
- The reference report was fetched at the SHA-256 the agreement record pins and its text extracted. The post-text set is Tables 14-1.01 to 14-7.04 — thirty — and Figure 14-1; the body carries Tables 11-1 and 12-1 to 12-4 and Figures 9-1 and 10-1. Titles and page numbers for the twelve displays the plan adds were read from their title pages, not from the assessment.
- The pilot's ADaM package on GitHub has no exposure dataset, and its subject-level dataset carries average daily dose and cumulative dose as columns. Its SDTM folder carries the medications domain. Both checked by listing the repository.
- Fifteen report text blocks: ten approved on 25 July, five agent-drafted and unread; the section 11.1 block still contains the stale sentence.
mainis not an ancestor ofdev: the release stamp in forty-three files has not been folded forward.devhas no commitsmainlacks, so the fold is a fast-forward.- The repeated-measures fit uses
nlme; themmrmpackage is not installed on the machine that builds this. The install cost is part of the spike. - The colleague's reference paths and the assessment's HTML are not reproduced here; the response file names its findings by content.