open.csr · after the user meeting · 2026-09-02

How every item in the sidebar is made, and what changes next

The demo's sidebar lists Documents, Displays, Text and Values; the top bar adds Templates. This page answers, for each one, the questions the meeting raised: what goes in, what the pipeline does with it, what comes out, where you can see it, and what it links to today. It then documents the refactor the questions point at: two new sidebar sections, Data and Metadata, linked explicitly from every artifact that used them, and the boundary decision that makes the ARD an input rather than something this system computes. State described is the v0.4.0 candidate.

The short version

Everything in the sidebar is one of three kinds of thing. It is declared (a spec or a model that a person or an agent writes, and the pipeline reads), computed (the pipeline makes it from declared things and data, and nobody edits it), or written (prose, which a person approves). The sidebar today shows the computed and the written things, and one kind of declared thing. Its two missing sections are the two kinds of input that have no page of their own: the Data that was measured, and the Metadata that was declared about the study and its documents.

Documents are computed from templates, text, displays and values. Displays are computed from a display specification and prepared data, through an ARD. Text is written, and binds numbers by address instead of typing them. Values are computed from named ARD addresses. Templates are declared. Data and Metadata are the inputs all of those read, and they are visible today only as labels.

The one rule that makes the sidebar honest: a computed thing is never edited; you change what it was computed from and rebuild. That rule is why every item below can state its inputs exactly.

The picture

Solid boxes are what exists today; dashed boxes are outputs; the two purple boxes are the sections the refactor adds. Every arrow in this diagram is a link the refactor makes explicit in the app.

data · proposed section
Study data
13 vendored transport files (10 ADaM, SDTM cm and dm, one relabelled copy), provenance to the upstream commit, two source lanes, documented derivations
metadata · proposed section
Study model · document models · specifications
study.yaml, the four template objects, each display's analysis.yaml and display.yaml, value declarations, text frontmatter, environment records
pipeline · R
regenerate(slug)
prepared data + specification → ARD → rendered table and RTF, one iteration folder per rebuild
displays · 30
ard.json, table.html, table.rtf, manifest.json
one row per statistic, each with an address
pipeline · R
regenerate_values()
named addresses and declared arithmetic → a store
values · 15
values.json
value, formatted value, the ARD row and hash it came from
text · 33 blocks
library/text/TXT-*.md
written prose; numbers arrive as bindings to ARD rows or to values
documents · 4 assembled
assemble.mjs → csr, synopsis, abbreviated, display package
seven gates: structure, bindings, digits, cross-references, values, treatment consistency, coverage

Read left to right, top to bottom: data and metadata feed the pipeline; the pipeline makes displays; displays feed values and text; templates assemble all three into documents. The documents' own provenance appendix (section 16.1.9) is where the whole chain is currently written down, once, at the end.

Sidebar item by item

Each card states the item's inputs and outputs as file paths, the steps that make it, where it shows, what it links to now, a short walkthrough script against the live candidate, and the refactor. Counts are the v0.4.0 candidate's.

Documents

sidebar · 5 listed · 4 assembled · 1 plannedcomputed

A document is the assembled report: one template object filled with this study's text, displays and values, then gated. The four that exist are the full ICH E3 clinical study report, the E3 synopsis, the abbreviated report and the post-text display package; the statistical analysis plan is listed as planned and has no assembler input yet.

Inputs
  • A template object: library/templates/<id>/sections.yaml (what this kind of document is) and assembly.yaml (what this study puts in each section)
  • Text blocks named in the assembly slots, from library/text/; generated-tier blocks only when approved
  • Each placed display's current iteration: its ARD and its pre-rendered variant (outputs/<slug>/current.json)
  • The values store outputs/values/values.json, and the study model library/study.yaml for the treatment-consistency gate
Outputs
  • docs/assembled/<basename>.json: the document as data, with every gate's result
  • docs/assembled/<basename>.html: the rendered document, including the provenance appendix at 16.1.9
  • The Reader pane, which renders the same document from library source, resolving bindings against the committed ARDs, so the app and CI agree

How it is made

  1. The R pipeline has already run: every display's current.json names its latest iteration.
  2. node scripts/assemble.mjs --all discovers every directory under library/templates/ with a sections.yaml.
  3. For each template it walks assembly.yaml: a slot's text blocks are rendered with their bindings resolved, in-text displays are drawn from the display's ARD, post-text displays are numbered from their order (the slug is the identity; the 14.x number is derived here).
  4. Seven gates run over the document being built: structure, binding resolution, numeric fidelity (no digit without a binding), cross-references, values (each re-derived from the ARDs), treatment consistency (every display's population against the study model and each other), display coverage.
  5. The JSON and HTML are written; the site build then wraps them for the Reader.

Where it shows

  • Demo → Documents pane: the clinical study report, and the other three from the sidebar. Each section shows its text blocks with an Edit control and its in-text tables; section 14 lists the post-text displays; 16.1.9 is the provenance appendix.

What it links to today

  • Every placed display links to its Displays entry; every text block shows its id chip and opens the editor in place.
  • Cross-references ({{xref:display:…}}, {{xref:section:…}}) resolve to numbered links inside the document.
  • 16.1.9 lists each display's ARD file, hash and environment, but as text: nothing in it links to a dataset, to the study model, or to the template that placed the display.

Walkthrough script

  1. Show the CSR at section 10. Say: nothing on this page was typed. The headings come from the document model, the prose from text blocks, the table from a display, and every number in the prose from an address into that display's results.
  2. Show the outlined numbers in 10.1. Say: the outline marks a binding. Hover one: it names the display, the analysis and the statistic it came from.
  3. Show section 11.1. Say: this block was rewritten last week and approved on 2 September; the approval is in the block's own header, and the report would not include it as generated prose until it was.
  4. Show the synopsis from the sidebar. Say: same displays, same text blocks, a different document model, so the demographics table is 14.1.2 in one and 13.2 in the other, and nobody renumbered anything.
  5. Show 16.1.9. Say: this is the chain of custody, written by the assembler: which ARD, which hash, which R version. The refactor turns each of those lines into a link.

Refactor needed

  • An inputs panel per document, above the text: the template object (Metadata), the study model (Metadata), the placed displays, the text blocks and their approval state, the values cited, and the datasets reached through the displays (Data). Every entry a link.
  • 16.1.9 becomes a generated view of the Metadata and Data pages rather than a flat list: each ARD line links its display iteration, its dataset entries and its environment record.
  • The Documents sidebar group shows the gate state per document (green, or the failing gate's name), since a document that failed a gate is not a document.
  • Known defect carried to v0.4.1: selecting a section in the sidebar updates the hash but does not scroll to it (open.csr #38).

Displays

sidebar · 30 · 27 tables, 1 listing, 2 figurescomputed

A display is a table, listing or figure the pipeline produced from a specification and prepared data. Its identity is the slug; its regulatory id, type, E3 position and datasets are declared; its numbers are computed into an ARD and rendered from it.

Inputs
  • library/tfl/<slug>/analysis.yaml: what to compute (dataset, analysis set, grouping column, the analyses and their methods, the denominator)
  • library/tfl/<slug>/display.yaml: how to show it (columns, rows, formats, footnotes, variants, figure block)
  • custom.R where a built-in method does not fit: 19 displays carry one, one borrows another's through custom_from
  • Prepared datasets from prepare_data(): the vendored transport files read through the source registry, with the documented derivations applied; adsl as the denominator
  • library/study.yaml: the arms in print order and the flag each analysis set means
Outputs
  • outputs/<slug>/vNNN/ard.json: one row per statistic, plus a provenance envelope (spec and display hashes, dataset hashes and row counts, R and package versions, the population record)
  • table.html and table.rtf per variant (table-in-text.* for the in-text redraw), from the same rendered cells
  • manifest.json: who regenerated it, why (the change request), every hash, and each variant's file; current.json points at the latest iteration
  • Snapshots of both specs beside the ARD, so an iteration is reproducible from its own folder

How it is made

  1. Rscript -e 'pkgload::load_all("pipeline"); regenerate("t-demographics")'.
  2. prepare_data() resolves each dataset the spec names to a source lane (the pilot's own package by default), reads the vendored file, applies the derivations the contract documents (population flags, age groups, the race recode, derived medications, baseline measures), and records a hash and row count per dataset.
  3. build_ard() runs each analysis in analysis.yaml, built-in or custom, and stacks the results into the ARD; ard_population() counts the subjects per arm the ARD summarises.
  4. render_display() draws each variant from the ARD and display.yaml; the RTF writer emits the same cells as RTF.
  5. The iteration folder is written with its manifest; current.json advances. Nothing under outputs/ is ever edited by hand, and a guard test re-hashes every committed RTF to prove it.

Where it shows

  • Demo → Displays pane, for example demographics: a header (slug, regulatory id, type, E3 position, status, current iteration, datasets, evidence, requirements, used in) and four tabs: Rendered display, ARD, Specs, Iteration timeline.
  • The gallery page for the same display outside the app, and its evidence page under Quality.

What it links to today

  • Used in → the documents that place it, with the number each gives it.
  • Evidence → the display's evidence page; requirements → the matrix rows.
  • Datasets is a label (adsl), not a link. The ARD tab shows the dataset hashes and the environment as text. The Specs tab shows the two YAML files but not the study model they depend on. Nothing links a display to the data it read or to the metadata it was built against.

Walkthrough script

  1. Show demographics, the header. Say: the slug is the name; the number 14.1.2 is assigned by whichever document places it. It reads one dataset, and it is at iteration v007.
  2. Show the Specs tab. Say: two files. The first says what to compute: intent-to-treat set, grouped by planned treatment, these analyses. The second says how to print it: column order, row labels, one decimal for a mean. Change either, rebuild, and you get v008.
  3. Show the ARD tab. Say: 548 rows, one per statistic. Row: analysis age, group Placebo, stat mean, 75.2. That row has an address, t-demographics:age:mean;group=Placebo, and any sentence in the report that needs 75.2 cites the address.
  4. Show the provenance at the top of the ARD tab. Say: the hash of the input dataset, the R and package versions, and the number of subjects per arm this table summarised. The build fails if that last number disagrees with any other table's.
  5. Show the Iteration timeline. Say: every rebuild kept, each with its change request. v001 to v007 is the history of this table since July.
  6. Show the RTF link on the Rendered display tab. Say: the submission-format file, from the same cells, hash-recorded so a hand-edited copy cannot pass as the pipeline's.

Refactor needed

  • The datasets label becomes links into the Data section, one per dataset, carrying the hash and row count the ARD recorded, so the reader lands on the exact input.
  • The header gains a Metadata line: the study model (arms, analysis set as this display resolved it), the template positions that place it, and the environment record of the current iteration.
  • The ARD tab states the ARD's origin: computed here from the specification (today's case) or received from a producer (the boundary decided below), with the producer named and the envelope's hashes either way.
  • The Specs tab links each spec to its Metadata page, where the same specification is shown with its version history and the displays that share custom code.

Text

sidebar · 33 blocks · 19 boilerplate, 9 parameterized, 5 generatedwritten10 approved · 23 draft

A text block is prose for one ICH E3 section, written by a person or drafted by a model, with numbers that arrive as bindings. Its tier says how it was made; its approval says whether the report may carry it; its frontmatter lists the displays it binds.

Inputs
  • library/text/<ID>.md: YAML frontmatter (id, E3 section, tier, version, displays, approval, provenance model and prompt, requirements, allowed digits) and the prose
  • ARD rows, by address: {{ard:<display>:<analysis>:<stat>;group=…}}, resolved against each display's current iteration (6 displays are bound by prose today)
  • Values, by name: {{value:<id>}} (3 blocks use one)
  • The assembler, for cross-reference tokens that resolve to numbers it owns
Outputs
  • No file is generated from a block; it is rendered wherever it is placed. Rendering resolves the bindings, tracks each substitution's span, and runs the two text gates: every binding resolves to exactly one row, and no digit in the rendered prose lacks a binding
  • In the assembled document, the block's rendered prose and its gate result
  • Editing in the app produces a unified diff of the source file and a patch; nothing is written from the browser

How it is made

  1. A person writes the file, or an agent drafts it with the model and prompt recorded in the frontmatter (generated tier).
  2. Numbers are not typed. The author looks up the address in the display's ARD tab or the value's name in the Values page and cites it.
  3. The assembler and the site render the block with bindings resolved; the fidelity gate fails the build on any stray digit.
  4. A reviewer sets approval in the frontmatter. Until then a generated block is held out of the report and shows as draft prose.

Where it shows

  • Demo → Text pane: one card per block with its id, tier, approval, bindings and an editor; the same editor opens in a drawer under the block in the Reader.
  • In every document that places the block.

What it links to today

  • Binding chips link to the display and focus the bound row.
  • The block's id anchors it from the Reader; the E3 section is a label from the frontmatter.
  • The displays list in the frontmatter is maintained by hand and can drift from the bindings actually used; the values a block cites are not listed anywhere; the section a block belongs to is not linked to the document model that defines it.

Walkthrough script

  1. Show the Text pane, block TXT-E3-1101. Say: the header says how this prose was made and whether it is approved. The chips are the numbers it uses, each an address.
  2. Show Edit, then type a digit into the prose. Say: the gate fails before you finish the sentence: a number with no address. Delete it and the gate passes. This is the same code CI runs.
  3. Show Copy patch. Say: an edit becomes a source change for review; the browser never writes to the repository.
  4. Show a generated-tier block marked draft. Say: the model and the prompt are on record; the report leaves it out until someone approves it in the file.

Refactor needed

  • Each block lists its inputs, derived at build rather than declared by hand: the displays and the iteration each binding resolved against, the values cited, and the document model section it belongs to (Metadata). The hand-maintained displays: list is retired or checked against the derived one.
  • The approval record links to the Metadata page that shows every block's approval trail, so a reviewer sees the whole report's state in one place.
  • Binding chips carry the ARD hash they resolved against, so a block rendered against a newer iteration says so.

Values

sidebar · 15 · 13 from an address, 2 by arithmeticcomputed

A value is a number the report reuses by name: randomised N, median age. It is either one ARD address named once, or a declared arithmetic over other values. The pipeline resolves it and the assembler re-derives it on every build.

Inputs
  • library/values/values.yaml: id, label, either source: (an ARD address, display included) or derived: (sum, difference, ratio, percent over other ids), a presentation format, notes
  • The ARD row each source address resolves to, in the display's current iteration
Outputs
  • outputs/values/values.json: per value the unscaled number, the formatted string, and the source (display, analysis, iteration, ARD file, ARD hash); the store's own provenance (source file hash, commit, environment)
  • In every document, the values gate's result: a value fails when its address resolves to zero or several rows, or the ARD no longer agrees with the store

How it is made

  1. Someone adds a value to the YAML with a label and an address.
  2. Rscript -e 'pkgload::load_all("pipeline"); regenerate_values()' resolves every address against the committed ARDs, evaluates the derived ones, and writes the store.
  3. Prose cites {{value:randomised-n}}; the assembler substitutes the formatted string and records the substitution like any binding.
  4. At assembly the gate re-derives every value from the same ARDs the report is built from, in JavaScript, and agrees with the R builder by construction: the derivation vocabulary is closed so both can evaluate it.

Where it shows

  • Demo → Values pane: one row per value with its label, number, an ARD badge, the provenance path with its hash, and what cites it.

What it links to today

  • The provenance path links to the display and its iteration; requirements link to the matrix.
  • Cited by lists text blocks. The datasets behind the ARD, and the study model that defines the arm named in the address, are two steps away and not linked.

Walkthrough script

  1. Show the Values pane, randomised-n. Say: 254 has a name. The name points at one row of the disposition table's results, iteration v003, hash recorded.
  2. Show a derived value. Say: the arithmetic is declared, not typed: the sum of two named values. R computes it for the store; the build recomputes it in JavaScript and they must agree.
  3. Show Cited by. Say: the sentences that use the name. Change the display, rebuild, and every one of them changes together, or the build fails and says which value moved.

Refactor needed

  • Each value links through its ARD to the Data entries the ARD was computed from, and to the study model's definition of the arm its address names (Metadata).
  • Values become citable from displays as well as text: a display footnote that states a denominator cites the value rather than a digit.

Templates

top bar, not the sidebar · 4 template objectsdeclared

A template object says what a kind of document is (sections.yaml, numbered sections and what each may contain) and what this study puts into it (assembly.yaml, which blocks and displays go in which slot). The full ICH E3 model has 119 sections; the synopsis, the abbreviated report and the display package are restrictions of it.

Inputs
  • Written by hand: library/templates/<id>/sections.yaml and assembly.yaml; the study block in the assembly must agree with library/study.yaml (a gate checks)
Outputs
  • None of its own. It is read by the assembler to make Documents, and rendered in the Templates pane with the display numbers the assembler would derive, so a reordering is visible before a build

Where it shows and what it links to today

  • Demo → Templates pane: every section of the model, marked populated or not, with the placed displays and their derived numbers linking to the Displays pane. Documents do not link back to the template that made them.

Refactor needed

  • Templates move into the Metadata section as its Document models entry, and every document links to the template object it was assembled from; the top-bar link stays as a shortcut.

Data

proposed sidebar sectioninput

What was measured. Today it is the best-recorded and least-visible thing in the system: every dataset's upstream path, commit and hash are in the repository, every ARD names the datasets it read with their hashes and row counts, and none of it has a page.

What exists today
  • pipeline/inst/extdata/phuse-cdiscpilot01/: 13 gzipped transport files vendored from phuse-org/phuse-scripts (MIT): 10 ADaM datasets, the SDTM cm and dm domains, and the relabelled medications copy kept for the derivation check
  • PROVENANCE.json: upstream repository, pinned commit and date, licence, and per file the upstream path, blob SHA-1, SHA-256 and byte counts; qc/vendor-phuse-data.R --check re-verifies
  • The source registry data_sources(): two lanes, the pilot's own package by default and the pharmaverse re-derivation as the measured alternate; their divergences recorded in quality/data/source-agreement.json
  • The preparation layer's derivations, documented in the contract: population flags, age groups, the race recode, baseline measures, derived medications from SDTM CM, the screened population from DM
  • In every ARD and manifest: dataset, hash, row count, source package and version
  • The extdata README, in the repository only
Where it is visible today
  • Nowhere in the app. A display's header says adsl as a label; the ARD tab shows a hash; the evidence page lists dataset names under traceability; 16.1.9 repeats the hashes as text

The section, as proposed

  1. One page per dataset: name and domain, lane, upstream path and pinned commit, file hashes, the prepared frame's rows and columns and its hash, the derivations applied to it (each linking to the contract paragraph and the code), and the displays that read it with the iteration and hash each recorded. The reverse index is the point: from a dataset to every number it produced.
  2. One page for the lanes: which dataset comes from where by default, the alternate lane, and the measured divergences between them.
  3. One page for the study's data package as a whole: the provenance record rendered, the licence, the verification command and its last result.
  4. Under the boundary decided below, one page per ARD received: producer, envelope hashes, the datasets the producer names, and the displays rendered from it.

Linked from

  • Every display header (datasets), every ARD tab (the hashes), every value (through its ARD), every document's inputs panel and 16.1.9, and the Metadata study page (the source the study model declares).

Metadata

proposed sidebar sectiondeclared

What was declared, as opposed to measured or written. It is scattered across five kinds of file today and shown in three different places, or not at all.

What exists today
  • The study model library/study.yaml: id, title, phase, cut-off, arms in print order, the columns that carry an arm label, the analysis sets with the flag defining each and the subjects each holds per arm, the data source. Read by the pipeline's treatment vocabulary and analysis-set registry and by the assembler's gate. Not shown anywhere.
  • The document models: the four template objects. Shown in the Templates pane.
  • Display specifications: analysis.yaml and display.yaml per display, snapshotted per iteration with their hashes, plus the regulatory id, type and E3 position the registry carries. Shown in each display's Specs tab and header.
  • Value declarations values.yaml, and text frontmatter (tier, version, approval, model and prompt, requirements). Shown per item.
  • Environment records: R, operating system and package versions in every manifest and ARD envelope; the commit when the tree was clean. Shown as text in the ARD tab and 16.1.9.
  • The registry site/config.json and the requirement matrices under quality/requirements/. Shown on the Quality pages.
  • The pilot's own analysis results metadata (its define.xml) is the reference the efficacy specifications were written from; it lives outside the repository and is cited in footnotes.
Where it is visible today
  • The Templates pane (document models), each display's Specs tab (its two files), the Quality pages (requirements). The study model, the environment records and the approval trail have no page.

The section, as proposed

  1. Study: the study model rendered: arms, group variables, analysis sets with their counts and the flag behind each, cut-off, source (linking to Data). Every display and value that names an arm or a set links here.
  2. Document models: the Templates pane, moved. Each document links to the model it was assembled from.
  3. Specifications: per display, the two specs with their version history and hashes and the custom code shared between displays; per value, its declaration; per text block, its frontmatter. Each artifact links to its own entry.
  4. Approvals: the approval state of every text block in one list, with who and when, so the report's readiness is one page.
  5. Environments: every distinct environment record across iterations, with the iterations built in each; 16.1.9 links here.
  6. Requirements: the matrices, already built, linked from here as well as from Quality.

Linked from

  • Every display (study model, specs, environment), every document (document model, study model, approvals), every text block (its section, its approval), every value (its declaration, the arm it names), and Data (the source the study declares).

The links, as one table

Rows are where a reader starts; columns are where a link should take them. Each cell says whether the link exists in the candidate and what the refactor makes it.

Fromto Datato Metadatato Displaysto Textto Valuesto Documents
Documentno · hashes as text in 16.1.9 → dataset pages via the placed displaysno → the template object, the study model, the approvals listyesyescited values shown in prose → the values panel per documentsynopsis and report do not link each other → same display, other document
Displaylabel only → one link per dataset with the recorded hashspecs shown, not linked → study model, spec history, environment, template positionsshared custom code not shown → displays sharing codeno → blocks bound to this displayno → values sourced from this ARDyes (Used in)
Text blockno → through its displaysno → its section in the document model; its approval in the trailyes (binding chips)no needno → values citedanchored from the Reader → documents placing it
Valueno → through its ARDno → its declaration; the arm namedyes (provenance path)yes (Cited by)derived inputs named → linkedno → documents where it is cited
Dataset (new)lanes, provenancestudy model sourceevery display that read it, with hashthrough displaysthrough displaysthrough displays
Metadata (new)study source → Datacross-links between model, specs, approvalsdisplays using an arm, a set, a spec, an environmentblocks per section; approvalsdeclarationsdocuments per model

The refactor, as increments

Ordered so each one is useful on its own and the earlier ones do not need re-doing when the later ones land. The first three are what the meeting asked for; the fourth is the boundary question raised on 2 September; the fifth is what the fourth makes possible.

R1 · Data section
A page per dataset, a page for the lanes, a page for the package; the reverse index from dataset to displays
Built by site.mjs from what already exists: PROVENANCE.json, the source registry, the prepared-data manifests in every ARD. The demo's tab list grows from five to seven (TAB_IDS in core.js), the sidebar gets a Data group, and the registry in site/config.json gains a data entry per dataset with its status. No pipeline change.
R2 · Metadata section
Study, Document models, Specifications, Approvals, Environments, Requirements
Templates moves in as Document models. The study page renders study.yaml; the specifications page renders each display's two YAML files with iteration history from the manifests; approvals and environments are collected from frontmatter and manifests. No pipeline change.
R3 · Explicit links
Every cell in the table above that says "no" becomes a link, both ways
Display headers link datasets and the study model; document inputs panels; text blocks list derived inputs (the hand-kept displays: list retired); values link through their ARD; 16.1.9 becomes links. Mostly renderer work in site.mjs and the demo client; one contract change, the derived inputs record per text block, written by the assembler into the document JSON.
R4 · The ARD is an input (design decision D13)
The system boundary moves to the ARD: open.csr renders, gates and assembles; it does not compute statistics
Everything downstream of the ARD already reads only the ARD. The change: regenerate() accepts an ARD file and skips the build; the analysis layer (analysis.yaml, custom.R, prepare_data(), the source registry, the vendored data) moves to a reference-producer package or the study repository; the treatment-consistency gate derives the population from the ARD rows when the envelope lacks it. The Data section gains its ARDs-received page. D9 is revised from "the pipeline regenerates" to "the pipeline renders". The current 30 committed ARDs keep the demo working as inputs from day one.
R5 · Producer-independent Data
The Data section describes whatever produced the ARDs, not only this repository's pipeline
Once R4 lands, a dataset page and an ARD page are populated from the producer's envelope: a SAS producer, a pharmaverse producer and this repository's reference producer all appear the same way. The three-route qualification becomes ARD against rendered against the reference report; recomputation from ADaM belongs to the producer.

What does not change

Recommendation. File one requirement for R1 to R3 as v0.5.0, "the sidebar says where every number came from", and one decision artifact for R4 with the three code changes named above, so the Data section is designed for producers from the start rather than retrofitted.