Every CSR-automation product surveyed consumes tables as frozen inputs. Every review tool closes a comment without knowing whether anything changed. Between them sits the request — untracked, re-keyed by hand, and adjudicated by email. This part maps who does what, what is genuinely new since the July 25 kickoff survey, and where the seam still is.
The 2025–26 commercial wave is large, funded, and one-directional: data → tables → prose, optimized for first-draft speed. Meanwhile the head of regulatory medical writing at J&J puts the bottleneck elsewhere: “the majority of a 12-week CSR cycle is spent in cross-functional review.” No product surveyed treats a change request to a table as a first-class object; no product links a regenerated table to the sentences quoting it. One new competitor — Clymb Clinical's Clymbr Hub (February 2026) — has assembled the ingredients on the TLF side and is the one to watch. The reverse edge, a reviewer's comment becoming a spec edit that regenerates the display and the prose in one transaction, remains unclaimed by anyone. And since the kickoff survey, the substrate moved in open.csr's favour: {cards} 0.8.0 shipped a purpose-built ARD diff, and CDISC ARS was verified to have no lifecycle model at all — the versioning layer has to be built, and nobody has built it.
The five incumbents from the kickoff survey, re-examined specifically for what happens when a reviewer wants a table changed. The pattern is uniform: the TLF package is an input, never an artifact the tool owns.
Word add-in; proprietary RAG restricted to allowed sources; 275+ eCTD templates; claims >90% TLF summarization accuracy. The closest it comes to the loop: teams can “refresh results at any point” — a re-import of a replaced TLF package, with no published evidence the refresh diffs numbers or flags affected prose. Certara's own webinar calls in-text tables “a significant productivity killer” — naming the pain, not solving it (product, webinar). Review and e-signature are delegated to Veeva Vault via the AI Partner Program (Oct 2025). Two structural notes: sibling product Phoenix TFL Studio ships an in-product TLF review surface (comments, audit logs) — but in the PK/PD line, disconnected from CoAuthor's CSR authoring; and Certara agreed in April 2026 to sell its regulatory-writing services arm to Veristat for up to $135M, keeping the software (press release).
Word plug-in + Vault integration; neuro-symbolic generation; the strongest traceability language in the market: “every sentence is linked to its source”, “update once, propagate everywhere.” Examined closely, the propagation is approved-content reuse across documents — one approved paragraph flowing into the CSR and the 2.7.4 — not numbers flowing from a regenerated table into dependent sentences. “Validate against source changes” is listed as a capability with no published mechanism (unverified). Best public benchmark: <7% of content requiring re-authoring on a GSK CSR (yseop.com). €10M growth funding Sept 2025; TIME Best Inventions 2025; still independent.
The cleanest determinism story surveyed: a rules-based layer transcribes every number (“no probabilistic inference”), the language model “never touches a number,” and output passes back through verification. But its part-owner describes the architecture plainly in AMWA Journal: sources are uploaded, the draft is generated, and “nothing is stored in the system… it functions much like a photocopier” (AMWA Journal 2024). A stateless tool cannot have a change-request concept: a revised TLF package means full regeneration and manual reconciliation against the writer's refined draft. Parent acquired by Indegene (2024); 2.7.4 and IB modules shipping mid-2026.
Web platform (not Word); TLFs and ADaM land in a knowledge graph; click a number in the narrative to jump to its source cell. That trace is a read-time audit affordance — no claim of staleness detection when the number underneath changes; Narrativa's own article on cross-referencing recommends manual periodic validation. And its TLF-automation page states the current loop plainly: writers “request changes in the output, which requires additional time to rewrite the code that produced the TLF documents” — their automation “shortens response times to medical writer change requests.” The request stays an out-of-band message to programming; the tool speeds the reply (narrativa.com).
CSR module in a full eClinical stack; auto-populates up to 70% of an ICH E3 template from protocol, SAP and TLFs; per-section status indicators, revision history, and manual per-section regeneration. Note for the record: the widely-repeated claim that Clinion cross-checks narrative values against tables traces to an AI-generated aggregator site, not to Clinion — their own pages never claim it (misattributed; treat as unverified).
The wave is growing fast — and every newcomer lands on the same side of the seam.
| Product | What it does | Seam verdict |
|---|---|---|
| Saama TLF Analyzer Oct 2025 | Multi-modal GenAI over TLFs: “Figures-to-Text” narrating KM/forest/waterfall plots, NL search across the package, protocol+SAP grounding, cited summaries (launch) | Reads tables beautifully; no change-request object, no regeneration linkage |
| Weave Bio $36M | IND/NDA authoring; Parexel partnership; the much-cited AutoIND benchmark (~97% faster) is vendor-authored, covers IND nonclinical prose only — no TLFs, no CSR, quality scores 69.6–77.9% (arXiv) | Adjacent scope; same one-directional shape |
| AlphaLife AuroraPrime | Word 365 + Veeva RIM; ingests TFLs from RTF/Excel; “one-click batch updates” when sources change; GenAI QC vs golden standards; Part 11 claim | Document-level refresh, not a change-request lifecycle |
| Assyro | The only vendor asserting live binding: citations are “a live link, not a static footnote… when source data updates, the citation updates with it” | Scoped to Module 2.5/2.7 evidence citations, not TLF cells; no verifiable footprint |
| ReadoutAI | Expert system picks the statistic, LLM phrases it, a fact-checking layer validates prose against computed statistics — architecturally the closest to number-guarded generation after TriloDocs | No change-request or versioning story published |
| TrialAssure LINK AI, Telperian, ZYLiQ, Instem | GenAI TLF processing and CSR drafting in Word; no-code TFL generation with Part 11-ready audit trails; post-text→in-text conversion; “80% needing only human QA” | First-draft acceleration across the board; none touch the loop |
Veeva's Falcon AI agents (announced May 2026) target TMF intake, safety cases and HA interactions — no CSR, no tables; agentic authoring inside Vault + Word is scheduled late 2027. Medidata Plus (July 2026) is a data-layer play with nothing on regulatory writing. IQVIA.ai (March 2026) names CSR generation as a target use case — extract, author by section, separate verifier agent — with no published numbers. Merck's internal iRAP (PHUSE US 2026) drafts CSRs from protocol, templates and “immutable TLFs” — the adjective is the whole finding: even the most sophisticated internal build treats displays as frozen inputs (partially unverified — paper PDF not directly retrievable).
One company has assembled nearly the whole pipeline on the TLF side, and its announcements bracket open.csr's territory from both ends.
Clymbr Hub (launched 17 Feb 2026) bundles SAP Builder → TFL Designer → TFL Code Generators → TFL Viewer → CSR Builder (press release). Three pieces matter:
ARS-native shell authoring producing human-readable shells and machine-readable CDISC ARS v1.0 JSON simultaneously; the {siera} R package meta-programs one script per output from that JSON. The community tier is free and COSA-approved — but ships one-time JSON/Excel export with explicitly no version control, no audit trail, no change management; the Enterprise tier adds exactly those, plus “review, approval and governance workflows” and Part 11 claims (clinstandards.org). The public GitHub repo contains a 2022 design-thinking workshop and no source. Shell versioning is precisely the thing behind the paywall — which is a market read worth taking seriously.
In-platform commenting on outputs, review status, approval workflow, version tracking, e-signature, role-based access — spanning medical writing and biostatistics (launched May 2025). Its PHUSE paper states the as-is problem verbatim: “scattered RTF and PDF files, manual feedback tracking via emails and Excel spreadsheets” (PHUSE SD09). Claims 70% reduction in review time (vendor figure, unverified).
Their ISS case study: 223 shells built in ~12 hours, 84% of programs auto-generated, and — the sentence that matters — “global changes (e.g., stat updates, additional treatment arm, precision) were applied across shells in minutes” (case study, unaudited). Supporting research: their PharmaSUG 2025 paper finds GenAI CSR drafting from ARD beats drafting from RTF/PDF by 3–5× in processing efficiency (AI-349) — independent confirmation of the ARD-first bet.
What Clymb has not shipped: the propagation runs forward (spec → shells → code → outputs). Nothing describes the reverse edge — a TFL Viewer comment becoming a spec edit — nor CSR Builder re-flagging prose when a regenerated table's numbers move. The “auditable link between narratives, source TFLs, and SAP specifications” is asserted with no published mechanism. The closed loop remains unoccupied — but this is the neighbour most likely to occupy it first.
Where CSR review actually happens today — and why none of it can route a table comment anywhere useful.
Claims 85% of the top-25 pharma and four of the top-five CROs; marketed for CSRs by name. Mechanics (from the v6.4 guide): click text → propose a change, comment, and category; every contribution gets a numeric comment ID and becomes a color-coded discussion thread; participants exit with a status; the owner downloads a reconciliation report — participants, statuses, comment counts — as the audit trail (quick guide). A table is reviewable — but only as text in a document: no cell-level anchoring, no typed change-request object, no linkage to the program or dataset that produced the table (inferred from silence across seven Ideagen sources). The lifecycle is open → reply → accept/reject → closed. There is no “regenerated” and no “re-verified.”
Annotations attach to a specific document version, on the PDF rendition — not the source. Reviewers cast verdicts (Ready for Approval / Requires Changes); “Requires Changes” routes the envelope to a revise step where a human uploads a new version. e-Signature captures name, time and meaning at document approval. Two telling details: resolved annotations are not brought forward to the next version, and nothing connects an annotation on a table to the program behind it — the handoff to statistical programming is entirely out-of-band (annotations, workflows). Vault's own co-authoring is weak enough (<100 pages, Office only, per-seat licensing) that pharma buys PleaseReview to run the review and checks the result back into Vault.
The default process underneath everything: tracked changes, comment logs, and CRMs where writers deliberately bring only “critical and at times conflicting comments” (Trilogy Writing). The same source documents review role-drift — statisticians rewording prose to taste, clinicians checking the abbreviation list — and every published mitigation is social, not technical: fewer reviewers, prioritized comments, someone with authority in the room.
The structured-content platform proves the review economics open.csr wants: components carry metadata, source references, approval status and lineage; “review once, reuse everywhere”; updates cascade. It is the closest precedent for approving a fragment rather than a document — and every fragment is narrative text. No Docuvera component's provenance is a statistical program.
Reconstructed from PHUSE/PharmaSUG papers, practitioner accounts and product documentation. This is the process the framework in Part 2 replaces.
| # | Step | Who / artifact |
|---|---|---|
| 1 | Displays produced and QC'd (double programming; statistician reviews against SAP/shells) | Programmer, validator, statistician · programming tracking sheet |
| 2 | Package handed over as hundreds of bookmarked RTF/PDF files | Statistician → medical writer |
| 3 | Reviewer forms an objection — wrong population, missing stratum, unusable footnote — and writes it wherever they happen to be: a PleaseReview thread, a Vault annotation, a Word comment, an Excel row, an email | Writer / clinician / statistician |
| 4 | Writer triages; only critical or conflicting comments reach the Comment Resolution Meeting | Medical writer · CRM agenda |
| 5 | Manual translation of the prose comment into something a programmer can act on — a tracker row, a spec edit, an email. No tool carries the comment to the code | Writer or lead statistician |
| 6 | Scope adjudication: analysis change → signed SAP amendment; shell/layout change → informal (shells are typically unsigned) | Biostatistician |
| 7 | Program edited; double programmer re-runs and re-QCs; Round Number increments in the tracking sheet | Programmers · tracking sheet (the industry's change counter) |
| 8 | Statistician re-reviews the regenerated output against SAP and shells | Biostatistician · QC checklist (✓ / NA / P) |
| 9 | New RTF/PDF issued; writer re-copies affected numbers into in-text tables and re-checks every downstream sentence | Medical writer · 100%-check QC convention |
| 10 | Original comment closed in the reconciliation report / Vault annotation resolved (and not carried forward) | Review owner |
| 11 | Document approval through Vault lifecycle states with Part 11 e-signature; E3 §16.1.5 carries the required signatures; the biostatistician appears as documented author | Sponsor signatories |
Bounding cycle times from the sources: topline tables ~1 week from lock; a full TLF (re)production round 2 weeks–1 month; a whole CSR 3–6 months, with most of a 12-week cycle spent in cross-functional review (Neil Garrett, Global Head of Regulatory Medical Writing, J&J, AMWA 2025). No published per-change-request cycle time exists — the request is not an object anyone measures. And the sharpest practitioner statement of the problem, from a statistician's CSR-review paper: “It is cumbersome to revise the final TLFs, if medical writer provides the comments on TLFs after the database lock” (PharmaSUG IB09).
What actually gets signed — the governance fact the framework builds on: the biostatistician's formal instrument is the SAP signature — “by signing the SAP, signatories agree to the statistical analyses… and to the basic format of the TFLs”; analysis changes thereafter require a signed amendment, while “the TLF shells… usually do not get signed, and thus can be changed without any formal amendment” (Being a Clinical Biostatistician). Output-level approval today lives in QC checklists and tracker columns, not signatures. The asymmetry — formal ceremony for analysis semantics, low ceremony for presentation — is exactly the two-tier review model Part 2 makes mechanical.
Verified against source — model YAML, NAMESPACEs, release tags — not documentation prose. Three findings change the design space.
No release since v1.0.0 (April 2024); the repo is in maintenance. The model's only versioning construct is an optional integer ordinal on four classes — zero occurrences of status, lifecycle, amendment, approval, author, date or reason anywhere in ars_ldm.yaml; even ReferenceDocument (the SAP link) carries no version or date. The gap is a known, three-year-old open issue — #206 asks for exactly the timestamped result-set concept open.csr's iteration ledger implements. Conclusion: adopt ARS vocabulary for analysis semantics, and build the versioning layer above it without waiting. One construct is worth adopting: AnalysisReasonEnum — SPECIFIED IN PROTOCOL / SPECIFIED IN SAP / DATA DRIVEN / REQUESTED BY REGULATORY AGENCY — the standard-sanctioned vocabulary for why a display exists.
{cards} 0.8.0 (after the kickoff survey) shipped compare_ard() / is_ard_equal(): key-matched, tolerance-aware comparison returning rows-added, rows-removed, and per-statistic differences. This is the engine a display change request needs, purpose-built, maintained upstream — experimental, so wrap and pin it. Equally important: cards::mock_*() creates empty ARDs in the results schema for table shells — so a shell and a result are the same artifact type, and one diff engine serves shell review, first-run review, and re-run review. Still no ARD serializer upstream (verified against NAMESPACE) — open.csr's owned ard.json schema remains necessary, and must carry the version stamps ARS and tfrmt JSON lack.
Exhaustive search: no open-source TLF tracker or shell change-control tool exists — not in pharmaverse (100 repos enumerated), not at Atorus (48 repos), not anywhere on GitHub. What exists is internal and proprietary, described in papers but never released: Keymed's ATTS platform (tracker + shell library + auto-QC), Merck's A&R Grid, CDARS's shell-metadata database — the last naming the unsolved problem: when programmers edit code directly, “the challenge is how to identify and capture changes, sync back and update the metadata database.” {tfrmt} remains the only open tool where the shell and the production spec are the same versionable artifact (JSON round-trip, mock-before-data), though its JSON carries no schema version. And {siera} 0.5.6 quietly shipped the best precedent in the ecosystem: an inst/method-library/ self-described as “the reviewable, diff-able source of truth” with stable method IDs referenceable from ARS files — plus a caution: validating against the CDISC eTFL reference package, siera found the published AE-overview values were computed from only the first AE record per subject. Published references are not yet trustworthy oracles; committed, tested ARDs are.
pharmaverse/cardinal cut its first release (v0.0.1, 23 Jul 2026): 26 FDA Standard Safety Tables + 4 figures — as imperative R in Quarto with a committed PNG, no declarative spec, no versioned display metadata. FDA meanwhile operationalized ST&Fs in review via MAPP 6025.9 (effective Jun 2025) — they are reviewer tools, so a sponsor shipping them pre-empts the reviewer's own view. On the agent side: no published “agent edits spec, human reviews diff” pattern exists in clinical reporting — the industry vocabulary is whole-artifact human-in-the-loop sign-off. The nearest precedents: Merck's PharmaSUG 2026 paper where the LLM emits an intermediate design document and the human reviews that; ClinAgent's automated diff against prior released specifications; and the R Consortium's pharma-skills repo (89★) with a four-phase human-reviewed lifecycle for agent skills — including a statistical-reviewer skill. The strongest measured result in the space says specification quality, not model sophistication, governs safety: behavior-driven specs moved AI-generated SDTM correctness from 20% to 84–90% (PharmaSUG 2026, Amazon).
Regulatory posture, verified primary-source: FDA's AI credibility guidance is still draft (Federal Register API sweep; no final notice exists) and its scope excludes “operational efficiencies” that don't affect the reliability of study results — deterministic rendering of validated numbers sits outside the framework; generating evidence sits inside. EMA's reflection paper (final, Sept 2024) supplies two directly usable rules: §2.3.5 — AI-drafted regulatory text needs “close human supervision” and “quality review mechanisms” — and §2.3.3.3 — after database lock, “any non-prespecified modifications to data processing or models implies that analysis results are considered post hoc,” with model changes requiring a SAP amendment via regulatory interaction. The EU AI Act does not class CSR drafting as high-risk, and Art. 50(4) expressly treats human editorial review as the discharging control. Part 2 encodes all three as mechanical properties of the loop.
Cross-cutting findings this pass establishes, each load-bearing for the framework in Part 2.
| Finding | Evidence |
|---|---|
| No product has a tracked TLF change-request object. | Uniform across 18 products; Merck's iRAP consumes “immutable TLFs”; TriloDocs is stateless “like a photocopier” by its owner's description |
| No product links a regenerated table to the prose quoting it. | Narrativa click-to-trace and Yseop sentence-links are read-time audit affordances; neither claims staleness detection; the reverse edge is unclaimed by every vendor |
| Vendors themselves describe the fallback. | Narrativa: automation “shortens response times to medical writer change requests” — the request is still an out-of-band message to programming |
| The review loop, not the first draft, is the bottleneck. | J&J: majority of a 12-week cycle is cross-functional review; the entire 2025–26 product wave optimizes drafting speed |
| Comment lifecycles terminate at closed/resolved. | PleaseReview reconciliation reports; Vault resolved annotations (not carried forward); no tool has a “regenerated / re-verified” state |
| Change control is process, not product. | Post-lock changes ride QMS paper — email, Excel trackers, Jira; the tracking sheet's Round Number column is the closest thing to a request counter |
| The two-tier governance already exists in practice. | Signed SAP for analysis semantics vs unsigned shells for presentation — the industry's own asymmetry, waiting to be made mechanical |
| The versioning layer must be built; nobody has. | ARS: no lifecycle fields (verified absence, issue #206 open 3 years); no open-source shell tracker anywhere; TFL Designer paywalls exactly this |
| The diff engine now exists upstream. | cards::compare_ard() (0.8.0, experimental) + mock_*() shells in the results schema — one differ for shells and results |
| ARD-first drafting is independently validated. | Clymb's PharmaSUG 2025 paper: GenAI drafting from ARD beats RTF/PDF parsing 3–5× on processing efficiency |
| Number-guarded generation exists; number-change handling does not. | TriloDocs Layer-1 verification, ReadoutAI fact-checking, IQVIA verifier agent — all prevent invented numbers; none handles a number changing |
| Clymb is the neighbour to watch. | ARS-native specs + codegen + TLF review with e-signature + CSR builder announced; comment→spec edit and prose re-flagging unpublished |
The structural conclusion is unchanged from the kickoff survey but now much better evidenced: the change request → spec edit → regenerated ARD → updated sentence transaction is unoccupied territory, the industry's own artifacts (Round Numbers, SAP amendments, 100%-check QC) prove the demand, and the standard capable of carrying display identity across that loop (CDISC ARS) has been adopted by exactly zero CSR-authoring vendors — only by the TLF-side tools. Part 2 proposes how open.csr occupies it.