open.csr roadmap research · 2026-07-27 · Part 1

The landscape: who owns the change request?

Every CSR-automation product surveyed consumes tables as frozen inputs. Every review tool closes a comment without knowing whether anything changed. Between them sits the request — untracked, re-keyed by hand, and adjudicated by email. This part maps who does what, what is genuinely new since the July 25 kickoff survey, and where the seam still is.

The short version

The 2025–26 commercial wave is large, funded, and one-directional: data → tables → prose, optimized for first-draft speed. Meanwhile the head of regulatory medical writing at J&J puts the bottleneck elsewhere: “the majority of a 12-week CSR cycle is spent in cross-functional review.” No product surveyed treats a change request to a table as a first-class object; no product links a regenerated table to the sentences quoting it. One new competitor — Clymb Clinical's Clymbr Hub (February 2026) — has assembled the ingredients on the TLF side and is the one to watch. The reverse edge, a reviewer's comment becoming a spec edit that regenerates the display and the prose in one transaction, remains unclaimed by anyone. And since the kickoff survey, the substrate moved in open.csr's favour: {cards} 0.8.0 shipped a purpose-built ARD diff, and CDISC ARS was verified to have no lifecycle model at all — the versioning layer has to be built, and nobody has built it.

18
products & platforms examined in this pass
0
with a tracked TLF change-request object
0
that re-flag prose when a table's numbers change
1
near-neighbour assembling the pieces (Clymb)
12 wk
typical CSR cycle — most of it review, not drafting

1 · The authoring wave: core functionality, and where each tool stops

The five incumbents from the kickoff survey, re-examined specifically for what happens when a reviewer wants a table changed. The pattern is uniform: the TLF package is an input, never an artifact the tool owns.

Certara CoAuthor — refresh, not request

Word add-in; proprietary RAG restricted to allowed sources; 275+ eCTD templates; claims >90% TLF summarization accuracy. The closest it comes to the loop: teams can “refresh results at any point” — a re-import of a replaced TLF package, with no published evidence the refresh diffs numbers or flags affected prose. Certara's own webinar calls in-text tables “a significant productivity killer” — naming the pain, not solving it (product, webinar). Review and e-signature are delegated to Veeva Vault via the AI Partner Program (Oct 2025). Two structural notes: sibling product Phoenix TFL Studio ships an in-product TLF review surface (comments, audit logs) — but in the PK/PD line, disconnected from CoAuthor's CSR authoring; and Certara agreed in April 2026 to sell its regulatory-writing services arm to Veristat for up to $135M, keeping the software (press release).

Yseop Copilot — reuse propagation, not number propagation

Word plug-in + Vault integration; neuro-symbolic generation; the strongest traceability language in the market: “every sentence is linked to its source”, “update once, propagate everywhere.” Examined closely, the propagation is approved-content reuse across documents — one approved paragraph flowing into the CSR and the 2.7.4 — not numbers flowing from a regenerated table into dependent sentences. “Validate against source changes” is listed as a capability with no published mechanism (unverified). Best public benchmark: <7% of content requiring re-authoring on a GSK CSR (yseop.com). €10M growth funding Sept 2025; TIME Best Inventions 2025; still independent.

TriloDocs — the photocopier, by design

The cleanest determinism story surveyed: a rules-based layer transcribes every number (“no probabilistic inference”), the language model “never touches a number,” and output passes back through verification. But its part-owner describes the architecture plainly in AMWA Journal: sources are uploaded, the draft is generated, and “nothing is stored in the system… it functions much like a photocopier” (AMWA Journal 2024). A stateless tool cannot have a change-request concept: a revised TLF package means full regeneration and manual reconciliation against the writer's refined draft. Parent acquired by Indegene (2024); 2.7.4 and IB modules shipping mid-2026.

Narrativa — the vendor that concedes the seam

Web platform (not Word); TLFs and ADaM land in a knowledge graph; click a number in the narrative to jump to its source cell. That trace is a read-time audit affordance — no claim of staleness detection when the number underneath changes; Narrativa's own article on cross-referencing recommends manual periodic validation. And its TLF-automation page states the current loop plainly: writers “request changes in the output, which requires additional time to rewrite the code that produced the TLF documents” — their automation “shortens response times to medical writer change requests.” The request stays an out-of-band message to programming; the tool speeds the reply (narrativa.com).

Clinion — section regeneration inside an EDC suite

CSR module in a full eClinical stack; auto-populates up to 70% of an ICH E3 template from protocol, SAP and TLFs; per-section status indicators, revision history, and manual per-section regeneration. Note for the record: the widely-repeated claim that Clinion cross-checks narrative values against tables traces to an AI-generated aggregator site, not to Clinion — their own pages never claim it (misattributed; treat as unverified).

New entrants, 2025–26

The wave is growing fast — and every newcomer lands on the same side of the seam.

ProductWhat it doesSeam verdict
Saama TLF Analyzer Oct 2025Multi-modal GenAI over TLFs: “Figures-to-Text” narrating KM/forest/waterfall plots, NL search across the package, protocol+SAP grounding, cited summaries (launch)Reads tables beautifully; no change-request object, no regeneration linkage
Weave Bio $36MIND/NDA authoring; Parexel partnership; the much-cited AutoIND benchmark (~97% faster) is vendor-authored, covers IND nonclinical prose only — no TLFs, no CSR, quality scores 69.6–77.9% (arXiv)Adjacent scope; same one-directional shape
AlphaLife AuroraPrimeWord 365 + Veeva RIM; ingests TFLs from RTF/Excel; “one-click batch updates” when sources change; GenAI QC vs golden standards; Part 11 claimDocument-level refresh, not a change-request lifecycle
AssyroThe only vendor asserting live binding: citations are “a live link, not a static footnote… when source data updates, the citation updates with it”Scoped to Module 2.5/2.7 evidence citations, not TLF cells; no verifiable footprint
ReadoutAIExpert system picks the statistic, LLM phrases it, a fact-checking layer validates prose against computed statistics — architecturally the closest to number-guarded generation after TriloDocsNo change-request or versioning story published
TrialAssure LINK AI, Telperian, ZYLiQ, InstemGenAI TLF processing and CSR drafting in Word; no-code TFL generation with Part 11-ready audit trails; post-text→in-text conversion; “80% needing only human QA”First-draft acceleration across the board; none touch the loop

The platform incumbents are elsewhere

Veeva's Falcon AI agents (announced May 2026) target TMF intake, safety cases and HA interactions — no CSR, no tables; agentic authoring inside Vault + Word is scheduled late 2027. Medidata Plus (July 2026) is a data-layer play with nothing on regulatory writing. IQVIA.ai (March 2026) names CSR generation as a target use case — extract, author by section, separate verifier agent — with no published numbers. Merck's internal iRAP (PHUSE US 2026) drafts CSRs from protocol, templates and “immutable TLFs” — the adjective is the whole finding: even the most sophisticated internal build treats displays as frozen inputs (partially unverified — paper PDF not directly retrievable).

2 · Clymb Clinical — the competitor to watch

One company has assembled nearly the whole pipeline on the TLF side, and its announcements bracket open.csr's territory from both ends.

Clymbr Hub (launched 17 Feb 2026) bundles SAP Builder → TFL Designer → TFL Code Generators → TFL Viewer → CSR Builder (press release). Three pieces matter:

TFL Designer: the spec is the artifact — and versioning is the paywall

ARS-native shell authoring producing human-readable shells and machine-readable CDISC ARS v1.0 JSON simultaneously; the {siera} R package meta-programs one script per output from that JSON. The community tier is free and COSA-approved — but ships one-time JSON/Excel export with explicitly no version control, no audit trail, no change management; the Enterprise tier adds exactly those, plus “review, approval and governance workflows” and Part 11 claims (clinstandards.org). The public GitHub repo contains a 2022 design-thinking workshop and no source. Shell versioning is precisely the thing behind the paywall — which is a market read worth taking seriously.

TFL Viewer: the only dedicated TLF review product found

In-platform commenting on outputs, review status, approval workflow, version tracking, e-signature, role-based access — spanning medical writing and biostatistics (launched May 2025). Its PHUSE paper states the as-is problem verbatim: “scattered RTF and PDF files, manual feedback tracking via emails and Excel spreadsheets” (PHUSE SD09). Claims 70% reduction in review time (vendor figure, unverified).

The one published claim of change propagation as a mechanism

Their ISS case study: 223 shells built in ~12 hours, 84% of programs auto-generated, and — the sentence that matters — “global changes (e.g., stat updates, additional treatment arm, precision) were applied across shells in minutes” (case study, unaudited). Supporting research: their PharmaSUG 2025 paper finds GenAI CSR drafting from ARD beats drafting from RTF/PDF by 3–5× in processing efficiency (AI-349) — independent confirmation of the ARD-first bet.

What Clymb has not shipped: the propagation runs forward (spec → shells → code → outputs). Nothing describes the reverse edge — a TFL Viewer comment becoming a spec edit — nor CSR Builder re-flagging prose when a regenerated table's numbers move. The “auditable link between narratives, source TFLs, and SAP specifications” is asserted with no published mechanism. The closed loop remains unoccupied — but this is the neighbour most likely to occupy it first.

3 · The review layer: comments that die at “resolved”

Where CSR review actually happens today — and why none of it can route a table comment anywhere useful.

PleaseReview (Ideagen) — the incumbent, and its ceiling

Claims 85% of the top-25 pharma and four of the top-five CROs; marketed for CSRs by name. Mechanics (from the v6.4 guide): click text → propose a change, comment, and category; every contribution gets a numeric comment ID and becomes a color-coded discussion thread; participants exit with a status; the owner downloads a reconciliation report — participants, statuses, comment counts — as the audit trail (quick guide). A table is reviewable — but only as text in a document: no cell-level anchoring, no typed change-request object, no linkage to the program or dataset that produced the table (inferred from silence across seven Ideagen sources). The lifecycle is open → reply → accept/reject → closed. There is no “regenerated” and no “re-verified.”

Veeva Vault — document-scoped revision, annotations left behind

Annotations attach to a specific document version, on the PDF rendition — not the source. Reviewers cast verdicts (Ready for Approval / Requires Changes); “Requires Changes” routes the envelope to a revise step where a human uploads a new version. e-Signature captures name, time and meaning at document approval. Two telling details: resolved annotations are not brought forward to the next version, and nothing connects an annotation on a table to the program behind it — the handoff to statistical programming is entirely out-of-band (annotations, workflows). Vault's own co-authoring is weak enough (<100 pages, Office only, per-seat licensing) that pharma buys PleaseReview to run the review and checks the result back into Vault.

Word, SharePoint, and the Comment Resolution Meeting

The default process underneath everything: tracked changes, comment logs, and CRMs where writers deliberately bring only “critical and at times conflicting comments” (Trilogy Writing). The same source documents review role-drift — statisticians rewording prose to taste, clinicians checking the abbreviation list — and every published mitigation is social, not technical: fewer reviewers, prioritized comments, someone with authority in the room.

Docuvera — fragment-level approval, prose only

The structured-content platform proves the review economics open.csr wants: components carry metadata, source references, approval status and lineage; “review once, reuse everywhere”; updates cascade. It is the closest precedent for approving a fragment rather than a document — and every fragment is narrative text. No Docuvera component's provenance is a statistical program.

4 · The as-is loop: eleven steps, three artifacts, no owner

Reconstructed from PHUSE/PharmaSUG papers, practitioner accounts and product documentation. This is the process the framework in Part 2 replaces.

#StepWho / artifact
1Displays produced and QC'd (double programming; statistician reviews against SAP/shells)Programmer, validator, statistician · programming tracking sheet
2Package handed over as hundreds of bookmarked RTF/PDF filesStatistician → medical writer
3Reviewer forms an objection — wrong population, missing stratum, unusable footnote — and writes it wherever they happen to be: a PleaseReview thread, a Vault annotation, a Word comment, an Excel row, an emailWriter / clinician / statistician
4Writer triages; only critical or conflicting comments reach the Comment Resolution MeetingMedical writer · CRM agenda
5Manual translation of the prose comment into something a programmer can act on — a tracker row, a spec edit, an email. No tool carries the comment to the codeWriter or lead statistician
6Scope adjudication: analysis change → signed SAP amendment; shell/layout change → informal (shells are typically unsigned)Biostatistician
7Program edited; double programmer re-runs and re-QCs; Round Number increments in the tracking sheetProgrammers · tracking sheet (the industry's change counter)
8Statistician re-reviews the regenerated output against SAP and shellsBiostatistician · QC checklist (✓ / NA / P)
9New RTF/PDF issued; writer re-copies affected numbers into in-text tables and re-checks every downstream sentenceMedical writer · 100%-check QC convention
10Original comment closed in the reconciliation report / Vault annotation resolved (and not carried forward)Review owner
11Document approval through Vault lifecycle states with Part 11 e-signature; E3 §16.1.5 carries the required signatures; the biostatistician appears as documented authorSponsor signatories

Bounding cycle times from the sources: topline tables ~1 week from lock; a full TLF (re)production round 2 weeks–1 month; a whole CSR 3–6 months, with most of a 12-week cycle spent in cross-functional review (Neil Garrett, Global Head of Regulatory Medical Writing, J&J, AMWA 2025). No published per-change-request cycle time exists — the request is not an object anyone measures. And the sharpest practitioner statement of the problem, from a statistician's CSR-review paper: “It is cumbersome to revise the final TLFs, if medical writer provides the comments on TLFs after the database lock” (PharmaSUG IB09).

What actually gets signed — the governance fact the framework builds on: the biostatistician's formal instrument is the SAP signature — “by signing the SAP, signatories agree to the statistical analyses… and to the basic format of the TFLs”; analysis changes thereafter require a signed amendment, while “the TLF shells… usually do not get signed, and thus can be changed without any formal amendment” (Being a Clinical Biostatistician). Output-level approval today lives in QC checklists and tracker columns, not signatures. The asymmetry — formal ceremony for analysis semantics, low ceremony for presentation — is exactly the two-tier review model Part 2 makes mechanical.

5 · Standards and open source: the substrate moved under us

Verified against source — model YAML, NAMESPACEs, release tags — not documentation prose. Three findings change the design space.

ARS has no lifecycle model — verified, and known upstream

No release since v1.0.0 (April 2024); the repo is in maintenance. The model's only versioning construct is an optional integer ordinal on four classes — zero occurrences of status, lifecycle, amendment, approval, author, date or reason anywhere in ars_ldm.yaml; even ReferenceDocument (the SAP link) carries no version or date. The gap is a known, three-year-old open issue — #206 asks for exactly the timestamped result-set concept open.csr's iteration ledger implements. Conclusion: adopt ARS vocabulary for analysis semantics, and build the versioning layer above it without waiting. One construct is worth adopting: AnalysisReasonEnumSPECIFIED IN PROTOCOL / SPECIFIED IN SAP / DATA DRIVEN / REQUESTED BY REGULATORY AGENCY — the standard-sanctioned vocabulary for why a display exists.

The diff primitive now exists upstream: cards::compare_ard()

{cards} 0.8.0 (after the kickoff survey) shipped compare_ard() / is_ard_equal(): key-matched, tolerance-aware comparison returning rows-added, rows-removed, and per-statistic differences. This is the engine a display change request needs, purpose-built, maintained upstream — experimental, so wrap and pin it. Equally important: cards::mock_*() creates empty ARDs in the results schema for table shells — so a shell and a result are the same artifact type, and one diff engine serves shell review, first-run review, and re-run review. Still no ARD serializer upstream (verified against NAMESPACE) — open.csr's owned ard.json schema remains necessary, and must carry the version stamps ARS and tfrmt JSON lack.

Nobody ships open-source shell versioning — and one vendor charges for exactly it

Exhaustive search: no open-source TLF tracker or shell change-control tool exists — not in pharmaverse (100 repos enumerated), not at Atorus (48 repos), not anywhere on GitHub. What exists is internal and proprietary, described in papers but never released: Keymed's ATTS platform (tracker + shell library + auto-QC), Merck's A&R Grid, CDARS's shell-metadata database — the last naming the unsolved problem: when programmers edit code directly, “the challenge is how to identify and capture changes, sync back and update the metadata database.” {tfrmt} remains the only open tool where the shell and the production spec are the same versionable artifact (JSON round-trip, mock-before-data), though its JSON carries no schema version. And {siera} 0.5.6 quietly shipped the best precedent in the ecosystem: an inst/method-library/ self-described as “the reviewable, diff-able source of truth” with stable method IDs referenceable from ARS files — plus a caution: validating against the CDISC eTFL reference package, siera found the published AE-overview values were computed from only the first AE record per subject. Published references are not yet trustworthy oracles; committed, tested ARDs are.

{cardinal} shipped, FDA hardened the target — and the agentic pattern is unclaimed

pharmaverse/cardinal cut its first release (v0.0.1, 23 Jul 2026): 26 FDA Standard Safety Tables + 4 figures — as imperative R in Quarto with a committed PNG, no declarative spec, no versioned display metadata. FDA meanwhile operationalized ST&Fs in review via MAPP 6025.9 (effective Jun 2025) — they are reviewer tools, so a sponsor shipping them pre-empts the reviewer's own view. On the agent side: no published “agent edits spec, human reviews diff” pattern exists in clinical reporting — the industry vocabulary is whole-artifact human-in-the-loop sign-off. The nearest precedents: Merck's PharmaSUG 2026 paper where the LLM emits an intermediate design document and the human reviews that; ClinAgent's automated diff against prior released specifications; and the R Consortium's pharma-skills repo (89★) with a four-phase human-reviewed lifecycle for agent skills — including a statistical-reviewer skill. The strongest measured result in the space says specification quality, not model sophistication, governs safety: behavior-driven specs moved AI-generated SDTM correctness from 20% to 84–90% (PharmaSUG 2026, Amazon).

Regulatory posture, verified primary-source: FDA's AI credibility guidance is still draft (Federal Register API sweep; no final notice exists) and its scope excludes “operational efficiencies” that don't affect the reliability of study results — deterministic rendering of validated numbers sits outside the framework; generating evidence sits inside. EMA's reflection paper (final, Sept 2024) supplies two directly usable rules: §2.3.5 — AI-drafted regulatory text needs “close human supervision” and “quality review mechanisms” — and §2.3.3.3 — after database lock, “any non-prespecified modifications to data processing or models implies that analysis results are considered post hoc,” with model changes requiring a SAP amendment via regulatory interaction. The EU AI Act does not class CSR drafting as high-risk, and Art. 50(4) expressly treats human editorial review as the discharging control. Part 2 encodes all three as mechanical properties of the loop.

6 · The seam, verified

Cross-cutting findings this pass establishes, each load-bearing for the framework in Part 2.

FindingEvidence
No product has a tracked TLF change-request object.Uniform across 18 products; Merck's iRAP consumes “immutable TLFs”; TriloDocs is stateless “like a photocopier” by its owner's description
No product links a regenerated table to the prose quoting it.Narrativa click-to-trace and Yseop sentence-links are read-time audit affordances; neither claims staleness detection; the reverse edge is unclaimed by every vendor
Vendors themselves describe the fallback.Narrativa: automation “shortens response times to medical writer change requests” — the request is still an out-of-band message to programming
The review loop, not the first draft, is the bottleneck.J&J: majority of a 12-week cycle is cross-functional review; the entire 2025–26 product wave optimizes drafting speed
Comment lifecycles terminate at closed/resolved.PleaseReview reconciliation reports; Vault resolved annotations (not carried forward); no tool has a “regenerated / re-verified” state
Change control is process, not product.Post-lock changes ride QMS paper — email, Excel trackers, Jira; the tracking sheet's Round Number column is the closest thing to a request counter
The two-tier governance already exists in practice.Signed SAP for analysis semantics vs unsigned shells for presentation — the industry's own asymmetry, waiting to be made mechanical
The versioning layer must be built; nobody has.ARS: no lifecycle fields (verified absence, issue #206 open 3 years); no open-source shell tracker anywhere; TFL Designer paywalls exactly this
The diff engine now exists upstream.cards::compare_ard() (0.8.0, experimental) + mock_*() shells in the results schema — one differ for shells and results
ARD-first drafting is independently validated.Clymb's PharmaSUG 2025 paper: GenAI drafting from ARD beats RTF/PDF parsing 3–5× on processing efficiency
Number-guarded generation exists; number-change handling does not.TriloDocs Layer-1 verification, ReadoutAI fact-checking, IQVIA verifier agent — all prevent invented numbers; none handles a number changing
Clymb is the neighbour to watch.ARS-native specs + codegen + TLF review with e-signature + CSR builder announced; comment→spec edit and prose re-flagging unpublished

The structural conclusion is unchanged from the kickoff survey but now much better evidenced: the change request → spec edit → regenerated ARD → updated sentence transaction is unoccupied territory, the industry's own artifacts (Round Numbers, SAP amendments, 100%-check QC) prove the demand, and the standard capable of carrying display identity across that loop (CDISC ARS) has been adopted by exactly zero CSR-authoring vendors — only by the TLF-side tools. Part 2 proposes how open.csr occupies it.