API reference

The library core

What a page or a widget can rely on before it asks for a chart: where the bundles are, what they define and the version they report; and the two steps every chart shares before anything is drawn, exported as BioViz.core: a variable named one way, and named variables resolved to one row per participant. The charts and the connection to R each have a reference of their own.

Loading the library

The bundles are committed, so a page needs no build step and no package manager. Copy the folder dist/bio.viz-0.2.0/ and load the script-tag bundle; it defines one global, BioViz:

<script src="dist/bio.viz-0.2.0/bio.viz.js"></script>
<script>
  console.log(BioViz.version); // "0.2.0"
</script>

An ES module bundle with the same exports sits beside it:

import { version, core, r } from './dist/bio.viz-0.2.0/bio.viz.esm.js';
FileWhat it is
dist/bio.viz-0.2.0/bio.viz.jsThe script-tag bundle. Defines the global BioViz.
dist/bio.viz-0.2.0/bio.viz.esm.jsThe ES module bundle. The same exports, as named exports.
*.mapA source map for each, so a debugger shows the source.

Nothing else is bundled into either file. safety.viz and R are loaded beside bio.viz on a page: safety.viz with its own script tag, and R the first time a statistic is asked for.

version

A string: the version of the library, 0.2.0. It equals the version field of package.json and is fixed when the bundle is built, so it says which build a page loaded. The folder the bundle sits in carries the same number.

A variable

A variable is what goes on an axis, makes a group, a colour or a panel. It is written one way wherever a chart takes one, and it is one of two things:

// a biomarker at a visit, with a value type
const y = { measure: 'IL-6', visit: 'Week 4', value: 'change' };

// a column
const x = { col: 'ARM' };

The biomarker and the visit are written as the results table writes them (IL-6, TNF-alpha, Week 4); nothing is assumed about what a name may contain. Which column holds the biomarker's name, the visit and the result is a setting.

variable(spec)

Checks a variable and returns it in full, frozen. A chart calls it where the variable is written, so a mistake is found there and not in an empty chart.

KeyForMeaning
measurea biomarkerThe biomarker's name, as the results table writes it.
visita biomarkerThe visit's name, as the results table writes it. Required, except with the value type baseline, which takes none.
valuea biomarkerThe value type. raw when not given.
cola columnThe column's name.
typea column'number' to read the column as a number. Without it the value is passed on as the table holds it.
cuteitherTo cut the number into groups: 'median', 'tertiles', 'quartiles', or the cut points as a list, ascending. A column cut must be read as a number (type: 'number'). See the cut rule.

It returns { kind: 'measure', measure, visit, value } or { kind: 'column', col, type }, with visit and type null where they do not apply, and cut when the variable has one (typed points as a frozen copy). A variable it returned can be handed back to it.

A malformed variable is refused: variable throws a TypeError whose message begins bio.viz:, quotes the variable as written and names what is wrong. It refuses a variable that:

The cut rule

One rule cuts a number into groups, the same in every chart that makes groups from one and the same in R. A variable carries the cut it asks for:

{ measure: 'CRP', visit: 'Baseline', cut: 'median' }   // two groups
{ measure: 'CRP', visit: 'Baseline', cut: 'tertiles' } // three
{ measure: 'CRP', value: 'baseline', cut: 'quartiles' } // four
{ col: 'AGE', type: 'number', cut: [40, 60] }           // typed points: three groups

The same groups in R, for gsm.bio or anyone checking a chart (tools/r-cut.R writes the expected results the unit tests hold this library to, with exactly these lines):

CUT_PROBS <- list(median = 0.5, tertiles = c(1, 2) / 3, quartiles = c(1, 2, 3) / 4)

cut_bound <- function(p) format(signif(p, 4), scientific = FALSE, trim = TRUE)

bound_labels <- function(points) {
  k <- length(points)
  if (k == 0) return(character(0))
  bounds <- vapply(points, cut_bound, character(1))
  middle <- if (k > 1) paste0("> ", bounds[-k], ", ≤ ", bounds[-1]) else character(0)
  c(paste0("≤ ", bounds[1]), middle, paste0("> ", bounds[k]))
}

cut_labels <- function(points) unique(bound_labels(points))

cut_points <- function(x, cut) {
  asked <- if (is.character(cut)) {
    stats::quantile(x, CUT_PROBS[[cut]], type = 7, na.rm = TRUE, names = FALSE)
  } else {
    cut
  }
  list(asked = asked, points = unique(asked))
}

cut_groups <- function(x, points) {
  as.character(cut(x, breaks = c(-Inf, points, Inf), right = TRUE, labels = bound_labels(points)))
}

A label holds the sign ≤ (U+2264), so R must run in a UTF-8 locale. Where a chart writes a cut variable into the identity of the rows it hands R, it writes it as the settings do, { measure, visit, value, cut } with the visit left out for a baseline value, or { col, type: 'number', cut }; typed points are always a list, so in R write them as one (I(c(2, 5)) or list(2, 5) for jsonlite), even a single point. jsonlite writes a number to four decimal places unless told otherwise, so write the identity with digits = NA too, which writes 15 significant digits, enough for a point typed with 15 or fewer: with the default, a typed point of 0.000012345 is written 0, and the key does not match.

CUTS

The cuts a variable may name, as a list: median, tertiles, quartiles. Typed points are a list of numbers instead.

cutPoints(values, cut)

The cut points of a variable's values, and the groups they make. values holds one value per participant; one that is not a finite number is missing and is left out. cut is one of CUTS or the typed points. It returns:

FieldWhat it is
cutThe cut, as given.
nHow many values it was worked out on.
askedThe points asked for: R's quantile() of the values, or the typed points.
pointsThe points used: asked with a repeated point once.
repeatedWhether a point repeated and collapsed.
mergedWhether points written alike gave groups the same label, which merged.
labelsThe label of each group, low to high: one more than there are points, fewer where groups merged.

With no value there is nothing to cut at: no points and no groups.

cutGroup(value, points)

The group a value falls in, counted from 0, low to high, as R's cut(right = TRUE) places it, so a value equal to a cut point is in the group below it, and its place among the labels of the cut, so groups that merged are one. Null for a missing value. points are a cut's points.

cutLabels(points)

The labels of the groups a cut's points make, low to high, as above: cutLabels([2, 5]) is ['≤ 2', '> 2, ≤ 5', '> 5']. A label that repeats is given once: cutLabels([2.7928, 2.793, 2.7932]) is ['≤ 2.793', '> 2.793, ≤ 2.793', '> 2.793'].

cutWords(cut)

A cut in words, for a label or a legend: cut at the median, cut at the tertiles, cut at 2 and 5. label adds it after the variable: CRP at Baseline, cut at the median.

VALUE_TYPES

The five value types, as a list: raw, baseline, change, fold_change, percent_change. Each is worked out for one participant from that participant's own results for the biomarker.

Value typeWhat it isNot worked out when
rawThe result at the visit.There is no usable result at the visit.
baselineThe baseline: the result at the baseline visits.There is no usable result at any baseline visit.
changeThe result at the visit minus the baseline.Either is missing.
fold_changeThe result at the visit divided by the baseline.Either is missing, or the baseline is zero or negative.
percent_change100 × (result at the visit − baseline) ÷ baseline.Either is missing, or the baseline is zero or negative.

The terms, exactly:

This follows safety.viz's shift plot, which names its baseline visits the same way and brings several to one value with the same statistics. It differs in three places: the shift plot has no fold change; it computes a percent change from a negative baseline, where this does not; and its first is the first row in the table, where here it is the first baseline visit as named in settings.

frame(tables, variables, settings)

Resolves named variables to one row per participant.

const { data, dropped } = BioViz.core.frame(
  { results, participants },
  {
    y: { measure: 'IL-6', visit: 'Week 4', value: 'change' },
    x: { col: 'ARM' }
  },
  { baseline_visits: ['Baseline'] }
);
// data:    [{ USUBJID: 'BIO-001', y: -1.027, x: 'Placebo' }, …]   186 records
// dropped: [{ reason: 'No result at the visit', variable: 'y', n: 13 },
//           { reason: 'Result at the visit is missing or not a number', variable: 'y', n: 1 }]
ArgumentMeaning
tables{ results, participants }. Each is an array of records, one object per row. results is required; participants may be left out.
variablesThe variables, each under the name its field is to have. The name is the caller's: y, x, panel, or anything else that is not the id column.
settingsColumn names and how the baseline is found: see DEFAULT_SETTINGS. May be left out.

It returns:

MemberMeaning
dataOne record per participant: the participant's id under the name of the id column, and one field per variable. Nothing else.
id_colThe name of the id field in data.
variablesThe variables in full, by name, as variable returns them.
participantsHow many participants were seen. It equals the records in data plus the counts in dropped.
droppedParticipants who are not in data, counted: a list of { reason, variable, n }. See DROPPED.
unusedRows that were read and not used, counted: a list of { reason, table, n }. See UNUSED.
baseline_visitsThe baseline visits that were used, or null when no variable needed a baseline.

The tables

Only the results table is required. It has one record per participant, biomarker and visit, and needs the id column always, and the biomarker, visit and result columns when a variable is a biomarker.

A column variable is read from the participant table when that table is given and has the column. Otherwise it is read from the results rows, where a participant's rows must agree: the column holds one value for the participant, and an empty cell beside a filled one counts as that one value. A participant whose rows hold more than one value is dropped and counted. A column that neither table has is refused.

A wide table, with one row per participant and a column per variable, needs no special handling. Give it as results, or as participants beside a long results table, and name its columns as column variables:

BioViz.core.frame(
  { results: wide },
  { y: { col: 'IL6_WEEK4', type: 'number' }, x: { col: 'ARM' } }
);

Who is in the frame

With a participant table given, it says who the participants are, and in what order: one record per participant in it. A participant who has results and is not in the participant table is left out and counted. A participant who is in it and has no results is dropped, and counted, for any biomarker variable. A second row for the same participant in the participant table is not used, and a row of either table with no id is not used; both are counted in unused.

With no participant table, the participants are everyone with a row of results, in the order first seen.

A participant is dropped when a variable cannot be worked out for them, and is counted once, under the first such variable in the order the variables were given. So data.length plus the counts in dropped is always participants.

By default every variable is required. The setting required names the ones that are; a variable not named there is left as null in a participant's record, which R reads as missing, instead of dropping the participant. required: [] keeps everyone. This is for a chart that compares many variables in pairs and must keep a participant who lacks one of them.

What it refuses

frame throws a TypeError whose message begins bio.viz: when the call cannot be made: tables that are not arrays of records, a table other than the two, variables not given by name, a variable named as the id column is, a malformed variable, a setting that is not known or has a value it cannot take, a column a setting names that the table does not have, or a column variable whose column is in neither table.

It does not refuse on what the tables hold. A biomarker or a visit that no row has gives an empty data and a count of everyone under No result at the visit.

It changes nothing it is given.

visits(results, settings)

The visits of a results table, in visit order, as a list of names. Only visits with at least one usable result are listed. A chart offers these in its visit control, and the first of them is the baseline visit when baseline_visits names none.

Visit order is one order, whatever order the rows come in:

So visits Screening (−1), Baseline (0), Day 1 (no number) and Week 4 (4) are in the order Screening, Baseline, Week 4, Day 1.

BioViz.core.visits(results); // ['Baseline', 'Week 2', 'Week 4', 'Week 8', 'Week 12']

results is the results table and settings the same settings frame takes; only the column names are read.

DEFAULT_SETTINGS

The settings and their defaults. The names are safety.viz's, so one column mapping drives both libraries, and the defaults are the columns of the synthetic study.

SettingDefaultMeaning
id_col'USUBJID'The participant's id, in the results table. Also the name of the id field in the frame.
measure_col'TEST'The biomarker's name.
value_col'STRESN'The result.
visit_col'VISIT'The visit's name.
visit_order_col'VISITNUM'A number that orders the visits (visit order), and so finds the first visit when no baseline visit is named. May be null.
participant_id_colnullThe participant's id in the participant table, when it is not named as id_col is.
baseline_visitsnullThe baseline visit, or a list of them. Null means the first visit in visit order.
baseline_stat'mean'How several baseline visits are brought to one value: one of BASELINE_STATS.
requirednullThe names of the variables a participant must have to be in the frame. Null means all of them.

BASELINE_STATS

How several baseline visits are brought to one baseline value: mean, min, max or first. first is the first baseline visit, in the order they are named in settings, at which the participant has a result. With one baseline visit they all give that visit's result.

DROPPED

The reasons a participant is not in the frame. Each entry of the frame's dropped list is { reason, variable, n }: one of these sentences, the name of the variable that could not be worked out, and how many participants. Only reasons that dropped someone are listed, in the order of the variables and then of this table.

KeyreasonWhen
NOT_IN_PARTICIPANT_TABLENot in the participant tableThe participant has results and is not in the participant table. variable is null.
NO_RESULTNo result at the visitNo row for the biomarker at the visit.
MISSING_RESULTResult at the visit is missing or not a numberThere are rows, and none has a usable result.
NO_BASELINENo baseline resultNo row for the biomarker at any baseline visit.
MISSING_BASELINEBaseline result is missing or not a numberThere are baseline rows, and none has a usable result.
ZERO_BASELINEBaseline is zeroA fold or percent change was asked for.
NEGATIVE_BASELINEBaseline is negativeA fold or percent change was asked for.
EMPTY_COLUMNColumn is emptyThe column has nothing for the participant.
VARYING_COLUMNColumn has more than one value for the participantA column read from the results rows, where the participant's rows disagree.
NOT_A_NUMBERColumn value is not a numberA column read as a number.

For a change, a fold change or a percent change, the visit is looked at before the baseline: a participant missing both is counted under the visit.

UNUSED

The reasons a row was read and not used. Each entry of the frame's unused list is { reason, table, n }: one of these sentences, results or participants, and how many rows. These count rows, not participants: a participant whose first result was used is in the frame however many later ones were not.

KeyreasonWhen
NO_IDRow has no participant idEither table.
DUPLICATE_PARTICIPANTLater row for a participant already in the participant tableThe participant table. The first row is the one used.
DUPLICATE_RESULTLater result for the same participant, biomarker and visitA usable result after the first, at a biomarker and visit a variable reads.
MISSING_RESULTResult is missing or not a numberA row with no usable result, at a biomarker and visit a variable reads.

Only the biomarkers and visits the variables read are counted: a duplicate at a visit no variable asked for is not looked at.

label(spec)

A variable in words, for an axis title or a legend. spec is a variable as written or as variable returned it.

VariableLabel
{ col: 'ARM' }ARM
{ measure: 'IL-6', visit: 'Week 4' }IL-6 at Week 4
{ measure: 'IL-6', value: 'baseline' }IL-6 at baseline
{ measure: 'IL-6', visit: 'Week 4', value: 'change' }IL-6 at Week 4, change from baseline

fold_change and percent_change read the same way: IL-6 at Week 4, fold change from baseline.

Handing the frame to R

data is what a chart draws, and it is what goes to R as the table of a statistics call: plain records of text and finite numbers, which the connection to R turns into a data frame with one column per field. The fields are named as the variables were, and gsm.bio's statistics functions take a table and the names of its columns, so the names pass straight through:

const { data } = BioViz.core.frame({ results, participants }, { y, x }, settings);
const result = await connection.run('Analyze_GroupDifference', {
  data,
  args: { strValueCol: 'y', strGroupCol: 'x', strMethod: 'wilcoxon' }
});

A field's name is the name the caller gave its variable, so it is stable whatever the biomarker is called, and two variables on the same biomarker are told apart by their names: { week4: { measure: 'IL-6', visit: 'Week 4' }, change: { measure: 'IL-6', visit: 'Week 4', value: 'change' } }. R takes any name as a column name here; a short plain one (y, x, v1) is easiest to read in R's messages.

portfolio

The chart list: an object naming every chart the library offers, in safety.viz's portfolio manifest format, version 2, so safety.viz's demo app can list bio.viz's charts and draw them beside its own on the files a study already has. It is src/data/portfolio.json, and the site publishes the same list at portfolio.json, with the format beside it at schema/portfolio.json.

One chart is not in it: the stratified survival chart reads an outcomes table, which no standard domain of the format holds, so the app could not hand it one (#63).

BioViz.portfolio.version; // 2
Object.keys(BioViz.portfolio.modules);
// ['group-comparison', 'association-scatter', 'correlation-matrix', 'biomarker-screen', 'cross-tab']
FieldWhat it says
version2, the format's version.
groupsOne group, biomarkers, labelled Biomarkers: the app lists the charts under it, on a tab of their own.
modulesOne entry per chart, keyed by its name: export, the factory's name on BioViz; title; library, bio.viz; and group, biomarkers.
modules.*.tablesThe tables init takes: results, from the labs and vitals domain (bds), required; and participants, from the subject-level domain (subject), optional. They are also the entry's domains and optionalDomains.
modules.*.settingsEach column setting of the chart, keyed as in its settings: the domain it reads, the standard column it defaults to (null where the chart has no default) and whether the chart cannot be made without it. participant_id_col reads the subject-level domain's USUBJID.
modules.*.unmappedSettingsomit: a setting with no column mapped is left out, so the chart keeps its default, because a chart refuses null for the columns it needs.

The participant is named once in each file. id_col reads it from the results and participant_id_col from the participant table, so when the two files call it differently the app maps each file's own name and the charts join the two on them. Left out, participant_id_col is id_col's name, as in every chart.

The list adds nothing to the format, and the format is safety.viz's: it is copied to src/data/schema/portfolio.json by node tools/vendor-portfolio-schema.mjs, with a record of the commit, and a test validates the list against the copy and holds each entry to its chart's own settings.

What is not here

No statistics. Deriving a change from baseline is arithmetic on one participant's own results, and a cut point is a description of the values, as a median is. Nothing in the core tests, estimates or compares: a median and the quartiles of a box belong to the chart that draws them, and every test to R.

What else is exported

ExportWhat it is
rThe connection to R and the rule for printing a p-value. Its reference is the connection to R.
coreThe variable and the frame, described above.

This page is the file docs/core.md, rendered. The site build fails when the library exports something the file does not document, or when the file documents a function the library does not export.