The demo's sidebar lists Documents, Displays, Text and Values; the top bar adds Templates. This page answers, for each one, the questions the meeting raised: what goes in, what the pipeline does with it, what comes out, where you can see it, and what it links to today. It then documents the refactor the questions point at: two new sidebar sections, Data and Metadata, linked explicitly from every artifact that used them, and the boundary decision that makes the ARD an input rather than something this system computes. State described is the v0.4.0 candidate.
Everything in the sidebar is one of three kinds of thing. It is declared (a spec or a model that a person or an agent writes, and the pipeline reads), computed (the pipeline makes it from declared things and data, and nobody edits it), or written (prose, which a person approves). The sidebar today shows the computed and the written things, and one kind of declared thing. Its two missing sections are the two kinds of input that have no page of their own: the Data that was measured, and the Metadata that was declared about the study and its documents.
Documents are computed from templates, text, displays and values. Displays are computed from a display specification and prepared data, through an ARD. Text is written, and binds numbers by address instead of typing them. Values are computed from named ARD addresses. Templates are declared. Data and Metadata are the inputs all of those read, and they are visible today only as labels.
The one rule that makes the sidebar honest: a computed thing is never edited; you change what it was computed from and rebuild. That rule is why every item below can state its inputs exactly.
Solid boxes are what exists today; dashed boxes are outputs; the two purple boxes are the sections the refactor adds. Every arrow in this diagram is a link the refactor makes explicit in the app.
study.yaml, the four template objects, each display's analysis.yaml and display.yaml, value declarations, text frontmatter, environment recordsRead left to right, top to bottom: data and metadata feed the pipeline; the pipeline makes displays; displays feed values and text; templates assemble all three into documents. The documents' own provenance appendix (section 16.1.9) is where the whole chain is currently written down, once, at the end.
Each card states the item's inputs and outputs as file paths, the steps that make it, where it shows, what it links to now, a short walkthrough script against the live candidate, and the refactor. Counts are the v0.4.0 candidate's.
A document is the assembled report: one template object filled with this study's text, displays and values, then gated. The four that exist are the full ICH E3 clinical study report, the E3 synopsis, the abbreviated report and the post-text display package; the statistical analysis plan is listed as planned and has no assembler input yet.
library/templates/<id>/sections.yaml (what this kind of document is) and assembly.yaml (what this study puts in each section)library/text/; generated-tier blocks only when approvedoutputs/<slug>/current.json)outputs/values/values.json, and the study model library/study.yaml for the treatment-consistency gatedocs/assembled/<basename>.json: the document as data, with every gate's resultdocs/assembled/<basename>.html: the rendered document, including the provenance appendix at 16.1.9current.json names its latest iteration.node scripts/assemble.mjs --all discovers every directory under library/templates/ with a sections.yaml.assembly.yaml: a slot's text blocks are rendered with their bindings resolved, in-text displays are drawn from the display's ARD, post-text displays are numbered from their order (the slug is the identity; the 14.x number is derived here).{{xref:display:…}}, {{xref:section:…}}) resolve to numbered links inside the document.A display is a table, listing or figure the pipeline produced from a specification and prepared data. Its identity is the slug; its regulatory id, type, E3 position and datasets are declared; its numbers are computed into an ARD and rendered from it.
library/tfl/<slug>/analysis.yaml: what to compute (dataset, analysis set, grouping column, the analyses and their methods, the denominator)library/tfl/<slug>/display.yaml: how to show it (columns, rows, formats, footnotes, variants, figure block)custom.R where a built-in method does not fit: 19 displays carry one, one borrows another's through custom_fromprepare_data(): the vendored transport files read through the source registry, with the documented derivations applied; adsl as the denominatorlibrary/study.yaml: the arms in print order and the flag each analysis set meansoutputs/<slug>/vNNN/ard.json: one row per statistic, plus a provenance envelope (spec and display hashes, dataset hashes and row counts, R and package versions, the population record)table.html and table.rtf per variant (table-in-text.* for the in-text redraw), from the same rendered cellsmanifest.json: who regenerated it, why (the change request), every hash, and each variant's file; current.json points at the latest iterationRscript -e 'pkgload::load_all("pipeline"); regenerate("t-demographics")'.prepare_data() resolves each dataset the spec names to a source lane (the pilot's own package by default), reads the vendored file, applies the derivations the contract documents (population flags, age groups, the race recode, derived medications, baseline measures), and records a hash and row count per dataset.build_ard() runs each analysis in analysis.yaml, built-in or custom, and stacks the results into the ARD; ard_population() counts the subjects per arm the ARD summarises.render_display() draws each variant from the ARD and display.yaml; the RTF writer emits the same cells as RTF.current.json advances. Nothing under outputs/ is ever edited by hand, and a guard test re-hashes every committed RTF to prove it.adsl), not a link. The ARD tab shows the dataset hashes and the environment as text. The Specs tab shows the two YAML files but not the study model they depend on. Nothing links a display to the data it read or to the metadata it was built against.t-demographics:age:mean;group=Placebo, and any sentence in the report that needs 75.2 cites the address.A text block is prose for one ICH E3 section, written by a person or drafted by a model, with numbers that arrive as bindings. Its tier says how it was made; its approval says whether the report may carry it; its frontmatter lists the displays it binds.
library/text/<ID>.md: YAML frontmatter (id, E3 section, tier, version, displays, approval, provenance model and prompt, requirements, allowed digits) and the prose{{ard:<display>:<analysis>:<stat>;group=…}}, resolved against each display's current iteration (6 displays are bound by prose today){{value:<id>}} (3 blocks use one)approval in the frontmatter. Until then a generated block is held out of the report and shows as draft prose.displays: list is retired or checked against the derived one.A value is a number the report reuses by name: randomised N, median age. It is either one ARD address named once, or a declared arithmetic over other values. The pipeline resolves it and the assembler re-derives it on every build.
library/values/values.yaml: id, label, either source: (an ARD address, display included) or derived: (sum, difference, ratio, percent over other ids), a presentation format, notesoutputs/values/values.json: per value the unscaled number, the formatted string, and the source (display, analysis, iteration, ARD file, ARD hash); the store's own provenance (source file hash, commit, environment)Rscript -e 'pkgload::load_all("pipeline"); regenerate_values()' resolves every address against the committed ARDs, evaluates the derived ones, and writes the store.{{value:randomised-n}}; the assembler substitutes the formatted string and records the substitution like any binding.A template object says what a kind of document is (sections.yaml, numbered sections and what each may contain) and what this study puts into it (assembly.yaml, which blocks and displays go in which slot). The full ICH E3 model has 119 sections; the synopsis, the abbreviated report and the display package are restrictions of it.
library/templates/<id>/sections.yaml and assembly.yaml; the study block in the assembly must agree with library/study.yaml (a gate checks)What was measured. Today it is the best-recorded and least-visible thing in the system: every dataset's upstream path, commit and hash are in the repository, every ARD names the datasets it read with their hashes and row counts, and none of it has a page.
pipeline/inst/extdata/phuse-cdiscpilot01/: 13 gzipped transport files vendored from phuse-org/phuse-scripts (MIT): 10 ADaM datasets, the SDTM cm and dm domains, and the relabelled medications copy kept for the derivation checkPROVENANCE.json: upstream repository, pinned commit and date, licence, and per file the upstream path, blob SHA-1, SHA-256 and byte counts; qc/vendor-phuse-data.R --check re-verifiesdata_sources(): two lanes, the pilot's own package by default and the pharmaverse re-derivation as the measured alternate; their divergences recorded in quality/data/source-agreement.jsonadsl as a label; the ARD tab shows a hash; the evidence page lists dataset names under traceability; 16.1.9 repeats the hashes as textWhat was declared, as opposed to measured or written. It is scattered across five kinds of file today and shown in three different places, or not at all.
library/study.yaml: id, title, phase, cut-off, arms in print order, the columns that carry an arm label, the analysis sets with the flag defining each and the subjects each holds per arm, the data source. Read by the pipeline's treatment vocabulary and analysis-set registry and by the assembler's gate. Not shown anywhere.analysis.yaml and display.yaml per display, snapshotted per iteration with their hashes, plus the regulatory id, type and E3 position the registry carries. Shown in each display's Specs tab and header.values.yaml, and text frontmatter (tier, version, approval, model and prompt, requirements). Shown per item.site/config.json and the requirement matrices under quality/requirements/. Shown on the Quality pages.Rows are where a reader starts; columns are where a link should take them. Each cell says whether the link exists in the candidate and what the refactor makes it.
| From | to Data | to Metadata | to Displays | to Text | to Values | to Documents |
|---|---|---|---|---|---|---|
| Document | no · hashes as text in 16.1.9 → dataset pages via the placed displays | no → the template object, the study model, the approvals list | yes | yes | cited values shown in prose → the values panel per document | synopsis and report do not link each other → same display, other document |
| Display | label only → one link per dataset with the recorded hash | specs shown, not linked → study model, spec history, environment, template positions | shared custom code not shown → displays sharing code | no → blocks bound to this display | no → values sourced from this ARD | yes (Used in) |
| Text block | no → through its displays | no → its section in the document model; its approval in the trail | yes (binding chips) | no need | no → values cited | anchored from the Reader → documents placing it |
| Value | no → through its ARD | no → its declaration; the arm named | yes (provenance path) | yes (Cited by) | derived inputs named → linked | no → documents where it is cited |
| Dataset (new) | lanes, provenance | study model source | every display that read it, with hash | through displays | through displays | through displays |
| Metadata (new) | study source → Data | cross-links between model, specs, approvals | displays using an arm, a set, a spec, an environment | blocks per section; approvals | declarations | documents per model |
Ordered so each one is useful on its own and the earlier ones do not need re-doing when the later ones land. The first three are what the meeting asked for; the fourth is the boundary question raised on 2 September; the fifth is what the fourth makes possible.
site.mjs from what already exists: PROVENANCE.json, the source registry, the prepared-data manifests in every ARD. The demo's tab list grows from five to seven (TAB_IDS in core.js), the sidebar gets a Data group, and the registry in site/config.json gains a data entry per dataset with its status. No pipeline change.study.yaml; the specifications page renders each display's two YAML files with iteration history from the manifests; approvals and environments are collected from frontmatter and manifests. No pipeline change.displays: list retired); values link through their ARD; 16.1.9 becomes links. Mostly renderer work in site.mjs and the demo client; one contract change, the derived inputs record per text block, written by the assembler into the document JSON.regenerate() accepts an ARD file and skips the build; the analysis layer (analysis.yaml, custom.R, prepare_data(), the source registry, the vendored data) moves to a reference-producer package or the study repository; the treatment-consistency gate derives the population from the ARD rows when the envelope lacks it. The Data section gains its ARDs-received page. D9 is revised from "the pipeline regenerates" to "the pipeline renders". The current 30 committed ARDs keep the demo working as inputs from day one.Recommendation. File one requirement for R1 to R3 as v0.5.0, "the sidebar says where every number came from", and one decision artifact for R4 with the three code changes named above, so the Data section is designed for producers from the start rather than retrofitted.