Decision · D0030 · 2026-08-27 · open.csr #57 · milestone v0.4.0

The SAP, and what a display looks like with no numbers in it

Recommendation in one paragraph: build the SAP as a fourth template over the libraries that already exist, and make a shell a state of a display rather than a separate artifact — because a shell turns out to be derivable from the spec alone, so nothing needs authoring twice and the two documents cannot disagree about what will be produced. Give the SAP its own text blocks: twenty-four of the thirty-three the report uses are written in the past tense about what happened, and only nine could be shared. Do not vendor the pilot's own SAP.

What is already here

Most of the statistical analysis plan is in this repository, dispersed. Seventeen of the twenty-six displays cite the CDISCPILOT01 statistical analysis plan by section, at real precision — §9.7.1.2 for the disposition test, §10.1.1 for the ADAS-Cog model, §10.2.1 for the NPI-X windowing, Appendix 1 §14.1.2 for the CIBIC+ scale — spanning its sections 8 through 14.

What does not exist is the document that assembles them. To learn what this study planned, a reader opens twenty-six display specifications and reads their footnotes.

site/config.json has declared a Statistical Analysis Plan since v0.1, as status: planned, with the note "Needs a SAP template in library/templates/ before it can assemble." This page is about what that template would have to be able to do.

The claim worth making is the one the framework was built for: the SAP and the report are the same objects in different states. One says what will be produced; the other reports what was. Assembled from one library they cannot disagree about the first — which is drift a real programme spends continuous effort policing.

How this was decided

By building the thing, not by reasoning about it. The load-bearing question is whether a display can render as a shell — its own rows and columns with no numbers in them — from its specification alone, before any study has run.

So a shell builder was written that reads only analysis.yaml and display.yaml, with no access to any ARD or outputs/ directory, and run against all twenty-six displays on dev at b9ecbfa. Everything below is what it produced or what the specs measured; nothing here is a projection.

What the shell builder produced

Three shapes of display exist in the library, and all three shell cleanly. These are real output, reformatted for the page.

t-populations — every row named in the spec21 of the 26 displays are this shape
PlaceboXanomeline Low DoseXanomeline High DoseTotal
Intent-To-Treat (ITT)————
Safety————
Efficacy————
Complete Week 24————
Complete Study————
t-ae-common — a hierarchy rather than rows4 of the 26: t-ae-common, t-cibic-categorical, t-conmeds, t-demographics
PlaceboXanomeline Low DoseXanomeline High DoseTotal
one row per system organ class, then per preferred term————
l-ae-serious — a listing, columns declared outright1 of the 26; it has no rows block at all, which is correct for a listing
USUBJIDTRT01AAGESEXAEBODSYSAEDECODAESEVAERELASTDYAENDYAEOUT
one row per serious adverse event

The hierarchical case is not a limitation, it is the convention. A real SAP shell says "adverse events by system organ class and preferred term" and does not list which terms, because nobody knows them yet. The four displays whose rows come from the data are the four whose shells are most obviously right.

Where the columns come from, which was the one open gap

Columns are derived from the data today — render.R takes them from the distinct group1_level values present in the ARD. A shell has no data, so this had to be answered before anything else. Three sources were checked against all twenty-six displays:

Candidate sourceWhat it actually yieldsVerdict
The spec's group: keyNames a variable, not levels — and three different ones are in use: TRT01A on 10 displays, TRTP on 9, TRT01P on 6, and empty on the listing.Not sufficient alone
A declared treatment vocabularyAll 25 grouped displays yield exactly Placebo, Xanomeline Low Dose, Xanomeline High Dose — identical across every one, and already a function in the pipeline (trt_levels()).Sufficient for the treatment columns
The spec's total: key for the extra columnPredicts whether a Total column appears on 26 of 26 displays. No exceptions.Sufficient for the rest

So the column set is fully derivable: the treatment vocabulary, plus total:, plus — on the single display that carries one — a p-value column that exists because the spec declares a p-value analysis. The listing declares its columns outright and needs none of this.

Five, each with a recommendation

D-SAP1 · state or artifact

Is a shell a state of a display, or a second thing that can be wrong?

This is the fork the rest follows from, and it is cheap now and expensive in a year.

  • A state. One display object, rendered with numbers for the report and without them for the plan. Nothing is authored twice, so the two documents cannot disagree about structure — the disagreement is not caught, it is impossible.
  • An artifact. A shell file per display, reviewable and signable as its own thing. Costs the guarantee: the moment a shell is a separate file it can fall behind the display, and the programme acquires a drift check it did not need.
Recommendation: a state. The build above is the argument — a shell came out of every one of the twenty-six specs with no ARD in reach, so the second object would carry no information the first does not already hold. A signable artifact can still be produced from the state later, if a sponsor needs one to sign; that is a rendering question and it stays open.
D-SAP2 · columns without data

Where do a shell's columns come from when no study has run?

  • From a declared treatment vocabulary plus the spec's own keys. Measured above: the vocabulary covers every grouped display, total: covers the extra column on 26 of 26, and listings declare their columns outright.
  • From the study model in site/config.json. Equivalent in effect, and puts the vocabulary where a second study would have to change it anyway.
  • Leave columns out of the shell. Honest, and useless — a shell with no columns cannot show what the table will look like, which is the only reason to draw one.
Recommendation: the declared vocabulary, sourced from the study model rather than hard-coded. trt_levels() already exists in the pipeline and is already the single source for the report's own column order, so this is a move rather than a new concept. The only real work is teaching the renderer to take levels from the model when there is no ARD — one branch, in one function.
D-SAP3 · what it carries

Does the SAP carry every display, or only the ones its analyses name?

The existing templates differ widely, so there is no house answer to inherit: the report and the display package carry 26, the abbreviated report 22, the synopsis 6.

  • Every display in the library. Simple, and wrong the first time a display is built that the plan never promised.
  • Only what its assembly names, like every other template. Consistent with how the other four work, and it makes "planned but not produced" and "produced but not planned" both visible as a difference between two assemblies.
Recommendation: only what its assembly names. And add the check that falls out of it for free — a display in the report's assembly and absent from the SAP's is either an unplanned analysis or a gap in the plan, and both are worth a build failure rather than a silence.
D-SAP4 · the words

Can the SAP share the report's text blocks?

Measured rather than assumed, across all 33 blocks in the text library: 24 are past-tense only — was, were, showed, occurred, discontinued — 9 are tense-neutral, and none uses will or shall. The library was written to report a study that had finished.

  • Share the library, add a tense variant per block. Keeps one home for the words. Costs a variant mechanism on every block to serve nine that need none.
  • A separate SAP text library, sharing only where a block is genuinely tenseless. Two homes, and the nine shared blocks are shared honestly rather than by a switch.
Recommendation: a separate library, sharing the nine. The measurement says the overlap is the exception. A tense switch on twenty-four blocks would be a mechanism built to disguise the fact that these are different sentences about different things — and a plan that reads as though the study already happened is a real review finding, not a formatting quibble.
D-SAP5 · the reference

Is the pilot's own SAP a reference to reproduce, the way the 2006 report is?

Established: no SAP is vendored. The eleven files in pipeline/inst/extdata/phuse-cdiscpilot01/ are the transport datasets and their provenance record, and none is the plan. Its content reaches the repository only as the section citations in seventeen displays' footnotes.

Not established, and not assumed here: whether the document is redistributable at all. The data is MIT via phuse-org/phuse-scripts; that says nothing about the study documentation, and the 2006 report was handled by extracting figures mechanically rather than by vendoring it.

Recommendation: no. The displays already cite it by section, which is the useful part and costs no licence question. Reproducing it would mean claiming our assembled plan matches a document the repository does not hold and may not be able to hold — and that is a qualification claim with nothing behind it.

What it costs, and what it does not

The honest limit. A shell shows what a table will look like. It cannot show that the analysis behind it is the right one — that judgement is what the plan's prose is for, and none of it is derivable from a spec. This work makes the structure of a plan free; it makes none of the statistics free.

Where each number came from

ClaimHow it was measured
17 of 26 displays cite the SAP by sectiongit grep for "statistical analysis plan" across library/tfl/, counted by display directory
21 name every row, 4 declare a hierarchy, 1 is a listingThe rows: block of every display.yaml, classified by whether it carries label: or levels:
Three grouping variables, one set of levelsgroup: in each analysis.yaml against the distinct group1_level values in each committed ARD
total: predicts the Total column on 26 of 26Each spec's total: compared with whether Total appears in that display's ARD levels
24 past-tense, 9 neutral, 0 future, of 33Every library/text/*.md body scanned for past-tense verbs and for will/shall
Templates carry 26, 26, 22 and 6loadAssembly() on each of the four template models, displays counted across slots and post-text
No SAP is vendoredThe 11 entries in PROVENANCE.json under pipeline/inst/extdata/phuse-cdiscpilot01/

Measured against jwildfire/open.csr dev at b9ecbfa on 2026-08-27. Nothing was filed, built or changed in open.csr for this page; the shell builder was a throwaway script in a scratch worktree and is not committed.