goal #79 · dashboard design session · 2026-07-28

Overview and domain dashboards: what the study is doing, and what just changed

Mockups and open questions for the four top-level pages. Grounded in two research passes: a survey of ~20 commercial and open-source clinical dashboards plus the primary regulatory texts (ICH E6(R3), E8(R1), E9, TransCelerate QTL, FDA RBM Q&A, DMC, DILI and Standard Safety Tables & Figures), and an inventory of what demo-301's snapshots can actually draw today. Every number in these mockups is real data from the live demo. Where a panel needs pipeline work that isn't built, it says so on the panel.

The thesis

"What changed since last review" is the product

Across roughly twenty products surveyed, almost nobody ships a since-last-review frame. The market leads with "real-time, updates automatically" — which structurally precludes the question, because a live view has no stored point to compare against. The exceptions are narrow and instructive: gsm.kri's own flag-transition list, EMA's eRMR (a change column with exactly three values), and Oracle Empirica's since-last-refresh counts.

We publish immutable snapshots. That is not a limitation to apologize for — it is the one architecture that answers the question a periodic reviewer actually asks. So the change frame should be the spine of the Study Overview, not a section inside it, and a review-period control should sit in the global chrome and re-scope every page beneath it. The closest working model outside clinical software is SonarQube's "new code period": an explicit, user-selectable reference point, with the quality gate scoped to the delta.

Second finding worth naming: funnel plots — the standard statistical answer to "is this site really an outlier" — appear in zero commercial RBQM products. They are purely academic today. That is a cheap, citable differentiator, and it happens to solve our worst data problem (see RBQM below).

Page 1

Study Overview

Two jobs, in this order: what changed since the snapshot you last reviewed, then is the study broadly on track. The review-period control at the top is the page's organizing device — it names the comparison in words, and everything below is scoped to it.

open.gismo Study OverviewRBQMSafetyData Explorer DEMO-301

DEMO-301 — Safety Review

Phase 2 · 765/1000 participants · 148/150 sites Comparing ps-001 → ps-002 ▾ ✓ 9c41f2a
Since ps-001 — 11 changes across 2 domains
New — needs review (4)
SITE4323AE Rate · Smith, Newtown Square▲ on track → amberreview
SITE7983Query Age▲ amber → redreview
SITE8091Risk Score 21.1 → 28.9▲ worsereview
Serious AEs 500 → 506within expectationno action
Persisting — already acknowledged (5)
SITE7838Lab Grade 3+ Rate→ red, unchangedack. 07-21 JW
Resolved (2)
SITE3311SAE Rate · amber → on track▼ betterclosed
Study health at ps-002
2Red sites= 2
39Amber▲ +3
109On track

617 site × metric cells below accrual threshold — not evaluated.

Enrolment vs plan — real curve, drawable today
target 1000 765 Jan 2012 Mar 29
Acceptable ranges (trial level)
needs pipeline work →
Screen-failure rate

24.0% within range

limit 30% · secondary 26%

Withdrawal of consent

2.1% within range

limit 5% · secondary 4%

Important deviations

7.8% trending to limit

limit 10% · secondary 8%

What's real: the flag tiles, the eleven change rows, and the enrolment curve (765 participants over 87 days against the 1000 target) are all computed from the live snapshots today. What isn't: the acceptable-ranges strip is drawn from ICH E6(R3)'s three states (within range / trending toward limit / breached) but no QTL workflow runs in the demo — see the RBQM section. The review-state column ("review", "ack. 07-21 JW", "closed") is the single biggest gap in the open-source ecosystem and the hardest question in this document — Q3.

Page 2

RBQM

You asked for the riskiest sites by SRS plus QTL charts. Both are here — but each comes with a finding you should see before we build it.

A ranked "riskiest sites" list is the one thing the statistics literature explicitly warns against. Spiegelhalter's 2005 Statistics in Medicine paper proposes funnel plots specifically to "avoid spurious ranking of institutions into league tables"; Goldstein and Spiegelhalter showed that once uncertainty is accounted for, ranks largely dissolve. Our data makes the case concretely: 110 of 143 sites score exactly zero, the top-five includes sites with 2 and 3 enrolled participants, and there is a four-way tie at rank 3. A league table would put a 2-subject site at the top of the page.

So the mockup shows both: the ranked table you asked for, with the denominator beside every score, and a funnel plot that puts the same sites against their precision. Which one leads is Q4.

open.gismo Study OverviewRBQMSafetyData Explorer DEMO-301

RBQM

12 site KRIs · 143 sites scored · gsm.kri vs ps-001 ▾ ✓ ps-002
2Red sites
39Amber
109On track
7Score movers
Sites needing attention — risk score, with denominator
SiteNScoreSince ps-001Drivers
SITE80912828.9▲ +7.9 IPD AE
SITE79833128.1▲ +14.0 QRY
SITE78382428.1= 0.0 LB
SITE5444214.9— low precision AE
SITE0531314.9— low precision SDSC

Sites below the precision floor are dimmed, not ranked. Score is a weighted sum ÷ maximum ×100 — hover any score to see the contributing flags and their weights.

Same sites, against precision — funnel plot
study mean SITE5444 · n=2 · inside SITE8091 · n=28 SITE7983 · n=31 participants enrolled →

The 2-participant site sits inside the wide end of the funnel — high score, no evidence. Only out-of-funnel sites are labelled.

Quality tolerance limits
not built — needs gsm.qtl

A trial-level acceptable-range chart (metric over time with the tolerance band and a secondary limit) would go here. gsm.qtl is not installed anywhere in the workspace and demo-301 runs zero QTL workflows; the QTLs also need an eligibility domain the demo never extracted. Scope: install gsm.qtl, copy its metric workflows, add an IE mapping + input extract, re-run. Roughly a four-repo change.

Available charts
all →
Site KRI report

gsm.kri · HTML

Country KRI report

gsm.kri · HTML

KRI scatter + bounds

gsm.viz · 22k bounds rows

Static metric charts

25 PNG · incl. srs0001

A real defect in the risk score, found while inventorying: the SRS denominator in our data is 114, not the 178 you'd get from summing all twelve site-KRI maximum weights. Two KRIs — study discontinuation and treatment discontinuation, both weighted 32, the joint-heaviest in the model — are all-NA in this dataset, so their rows are dropped by the weight join and they vanish from the denominator entirely. Two of the three heaviest KRIs contribute nothing to the published score, and any tooltip claiming "out of 178" would be wrong. Q6 asks whether we fix the data or expose the effective denominator in the UI.

Page 3 · your open question

Safety — three shapes, and a recommendation

You said you weren't sure about a safety overview. The research turned up an unusually clean answer to "what belongs here," and one firm answer to "what doesn't."

There is a regulator-authored information architecture available for free. FDA's Standard Safety Tables and Figures: Integrated Guide (60 tables, 22 figures) is organized as: trials analyzed and exposure → AE overview → serious AEs → AEs leading to discontinuation → AEs of special interest → subgroups → data availability → laboratory → DILI screening → vital signs. Two things in that order we would otherwise get wrong: missing-data analysis comes before the lab results (differential data availability is a safety display, not a data-quality footnote — a low event rate at a site with 40% missing labs is not reassurance), and every screening plot ships paired with the listing of subjects in the quadrant of interest.

And the firm answer: there should be no safety composite score, ever. Nothing in any regulatory text supports collapsing heterogeneous safety signals into one number, and FDA says it will inspect the sponsor's reasoning — which is exactly what a weighted composite launders. RBQM composites work because KRIs proxy process quality; safety signals proxy causation, where the weighting is the scientific claim.

Option A

Exposure & census

Lead with denominators: randomized, on-treatment, cumulative person-time, disposition and discontinuations, follow-up completeness, lab coverage. Charts below.

Safest and fully drawable today. Honest but undramatic — it answers "can I trust the rates" rather than "is anything wrong."

Drawable now
Option B · recommended

Observed vs expected

Headline row of prespecified events — serious AEs, deaths, AESIs, possible Hy's Law cases — each showing the observed rate against a predicted rate with uncertainty, plus a review-worthiness queue of participants needing case review.

This is FDA's explicitly recommended pattern for blinded aggregate safety review, and crossing a trigger produces a handoff, not an in-app analysis.

Mostly drawable; needs a predicted-rate source and a Hy's Law summary emitter
Option C

Chart index

The nine safety renderers as a gallery with a thin summary strip — essentially today's page, tidied.

Cheapest. But FDA's DMC guidance is blunt that listings without context "are rarely useful," and safetyGraphics already does this — every chart a domain deep dive, no synthesis layer. Matching the incumbent isn't the goal.

Drawable now
open.gismo Study OverviewRBQMSafetyData Explorer DEMO-301

Safety

765 participants · 21,840 participant-days on treatment vs ps-001 ▾ ✓ ps-002
Prespecified events — observed vs expected — pooled across arms (blinded view)
EventObservedRateExpectedSince ps-001
Serious adverse events50619.6%18–24%within▲ +6 events
Grade 3+ laboratory8071.4%0.8–1.2%above▲ +235
Possible Hy's Law cases222.9%<1%above▲ +3
Deaths10.1%count only= 0
Discontinuations for AE192.5%context= 0

Expected ranges are illustrative here — a real deployment sources them from placebo databases, historical data or registries, and FDA asks that the comparison account for uncertainty rather than compare two point values.

Needs case review — ordered by review-worthiness, not risk
ParticipantFindingWhy now
S4323-011ALT 8.2× · TB 2.4×Hy's Law quadrant, newopen
S7983-004ALT 5.1× rising2 consecutive visitsopen
S7838-002Grade 4 neutropeniaseverity + seriousack 07-21

Ordered by severity, recency, spread across sites, data completeness and narrative availability — the vigiRank pattern. It claims only "look at this next," never "this is how dangerous the drug is."

Who is under-observed — FDA puts this before the lab results
36% Baseline Week 12

Lab coverage by visit, from the real mapped data (12,160 baseline records falling to 272 at week 12). Reassurance from a late visit is worth what its denominator is worth.

Available charts
all 9 →
Hepatic Explorer

eDISH · 765 participants

AE Explorer

10 SOCs · 2,583 events

Safety Histogram

16 analytes

QT Explorer

ECG · 8,130 records

Every number above is real except the expected-rate column: 506 serious AEs, 807 grade 3+ labs, 22 participants in the Hy's Law quadrant, 1 death, and the lab-coverage curve all come from the unified data that went live this evening. One caveat the mockup embodies: the Hy's Law count currently exists only inside a 10MB rendered HTML file — surfacing it as a tile needs gsm.safety to emit a summary alongside the chart.

Components

Three conventions to settle once

ComponentProposalWhy
Change chip Three values — new worsened resolved — plus a neutral no meaningful change. Arrow glyph and a word, never colour alone. Direction of "bad" comes from metric metadata, not the sign of the delta. EMA's eRMR uses exactly three change values. IBCS (now an ISO standard) classifies variances as good and bad, not positive and negative — for many KRIs neither direction is inherently good, only "outside expectation". Colour-only status fails WCAG 1.4.1, and red/green is the worst possible pair for colour vision deficiency.
Suppression rule Compute every delta; render a directional chip only when the change clears a noise threshold, otherwise a neutral glyph with the number still on hover. Drop uninteresting transitions entirely (gsm already drops NA → green). This is what keeps the change list short enough to read. Clinical decision support gives us the cautionary base rate: 49–96% of alerts are overridden, and acceptance drops ~30% for each additional repeated alert. The empirical RBQM flag rate is ~2.4% of sites — treat materially more as a defect.
Review state Per row: unreviewed / acknowledged (with reason and reviewer) / resolved. Resolved is demoted but retrievable. A row that changes reverts to unreviewed. The largest gap in the open-source ecosystem — only ClinSight has it. Empirica's headline tile is literally "percentage of tracked alerts reviewed". The most-reported RBQM failure in practice is dashboards whose findings never become actions. But this breaks static-snapshot purity — see Q3.

Deliberately not used, with reasons: gauges and traffic lights (angle and area are the least accurate encodings, and NHS England ran a national programme to replace RAG reporting with something better); dual axes; bare percentages without denominators; and sparklines as the primary trend read — with two snapshots a sparkline is a two-point line, and process-behaviour charts need roughly 12–25 points before their limits mean anything.

Decisions

Questions for you

IDQuestionMy recommendation
Q1Safety overview shape — A exposure & census, B observed-vs-expected with a review queue, or C chart index?B, with A's denominators as its second row. It is the only shape FDA actually prescribes for blinded aggregate review, and it gives the page a job beyond hosting charts.
Q2Treatment arm. The unified demo data now carries arm (Placebo / 40mg / 80mg) and the safety charts can split by it. FDA has stated twice that reviewing data by treatment group — including coded A/B/C labels — counts as unblinded. Does the study-team view ever split by arm?No arm anywhere in the study-team view — not a toggle, not a colour, not hidden behind a role, and ideally absent from the payload. Declare the product ICH E9 Type-1 monitoring. If a DMC mode is ever wanted, make it a separate deployment, not a role flag. ICH E9 §4.5 goes further than data: staff must stay blind "because of the possibility that their attitudes to the trial will be modified" — so a view must not telegraph a direction even without showing the numbers. This is the most consequential question here.
Q3Review state — do we ship per-row "reviewed / acknowledged / resolved" that persists across snapshots? It is the biggest gap in the ecosystem and the one thing that turns a viewer into a workflow. But review state is mutable and lives outside the immutable snapshot.Yes, in the study repo — a side-channel JSON committed by a PR, so the audit trail is the git history and the snapshot stays pure. That also makes review a reviewable artifact, which suits a GxP setting. Alternative: stay a stateless viewer and let review live in the meeting, as some vendors deliberately choose.
Q4Riskiest sites — ranked table (your ask) or funnel plot as the primary widget?Ranked table leads, funnel plot beside it, with denominators always visible and low-precision sites dimmed rather than ranked. You get the scannability you asked for; the funnel keeps us honest and is a genuine differentiator.
Q5QTL / acceptable ranges — build now or defer? It needs gsm.qtl installed, its workflows copied, an eligibility domain added and a re-run.Build it, but as its own requirement. ICH E6(R3) makes pre-specified acceptable ranges a sponsor obligation and gsm.kri has no QTL surface at all — it is our clearest regulatory gap. Keep it strictly separate from site KRIs; conflating levels is the named failure mode in TransCelerate's guidance.
Q6The SRS denominator defect — two 32-weight KRIs are all-NA and silently drop out, so scores are out of 114 rather than 178. Fix the data, or expose the effective denominator?Both. Populate the discontinuation flags in the demo data so the score means what it says, and show "33 of 114 possible" in the tooltip regardless — a score whose denominator moves silently is not auditable.
Q7Trend axis. Snapshot-over-snapshot trends are blocked by a bug — og_run stamps every snapshot with today's date, so both cuts share a date and the longitudinal accumulation path is fed NULL. Fix now?Fix it — it is small and it unblocks everything. Until then, trends on these pages are calendar-time curves (enrolment, AE onset, lab coverage), which are real and drawable today. Index the snapshot axis by review cycle rather than date, as one vendor does; with irregular cadence it matches how a reviewer experiences the study.
Q8Audience — is Study Overview for the study team, or a portfolio reader one level up? We already have a cross-study risk score widget.Study team for v1. A portfolio grid is a different product; the cross-study widget can seed it later.

One more thing

safetyGraphics and safetyCharts were archived on CRAN on 2026-03-25 — the reference implementation is no longer installable. It is the cleanest possible statement of why this project exists, and it is a verifiable fact rather than a competitive claim. It also names a project you co-authored, so how loudly we say it is your call.

Research notes: ~20 products surveyed (four major vendors' marketing sites blocked automated fetch — those claims rest on vendor self-description and are flagged as such in the underlying report). Primary regulatory texts were read in full. The one gap worth knowing: no guidance sets a numeric minimum cell size for site-level cross-tabs, so any k-anonymity floor we apply is engineering judgment without a citation behind it.