the standards this release set are written into the scaffold — the status ladder, the app's conventions and how a…
the standards this release set are written into the scaffold — the status ladder, the app's conventions and how a…
nothing looks broken — the defects found in the review of release 1.10, and the keynote's demo path walked by a test
all files come in on the Data tab — gsm raw files recognised there, and the file box gone from the RBQM tab
the RBQM tab reads at a glance — a row of metrics like every other tab, status icons, a table that fits its numbers,…
one R control — slim, in the same place on both tabs that use R, and a chip once R is ready
one status ladder — Qualified, Exploratory, Experimental, Prototype — with one label for the app and one for anything…
every tab has a colour, and the first screen says where you are — tab colours, counts, a welcome line and the way back…
the required check runs every test once, on parallel runners, and answers in about six minutes
the status label in gsm.safety's R widgets, from the safety.viz 1.11 bundle
the status label in gsm.safety's R widgets, from the safety.viz 1.11 bundle
one casing for chart names in the demo app, the gallery and the evidence pages
one casing for chart names in the demo app, the gallery and the evidence pages
one set of treatment-arm colours across the QT, time-to-event and biomarker charts
one set of treatment-arm colours across the QT, time-to-event and biomarker charts
every tab has a colour, and the first screen says where you are — tab colours, counts, a welcome line and the way back…
a page on the docs site for the RBQM tab, linked from its footnote
a page on the docs site for the RBQM tab, linked from its footnote
nothing looks broken — the defects found in the review of release 1.10, and the keynote's demo path walked by a test
the standards this release set are written into the scaffold — the status ladder, the app's conventions and how a…
all files come in on the Data tab — gsm raw files recognised there, and the file box gone from the RBQM tab
the RBQM tab reads at a glance — a row of metrics like every other tab, status icons, a table that fits its numbers,…
one R control — slim, in the same place on both tabs that use R, and a chip once R is ready
one status ladder — Qualified, Exploratory, Experimental, Prototype — with one label for the app and one for anything…
all files come in on the Data tab — gsm raw files recognised there, and the file box gone from the RBQM tab
the evidence refresh run takes the same split and finishes in about the time the check does
the required check runs every test once, on parallel runners, and answers in about six minutes
What safety.viz v1.11.0 changes in the demo app, shown on the live site beside release 1.10: the first screen, one status label, one control that starts R, the RBQM tab and the Data tab, with steps to try and what is yours to rule on at the review.
**See it move:** [annotated v1.11.0 demo](https://jwildfire.github.io/obot.roadmap/reports/sv-v1.11-demo/)
the RBQM tab reads at a glance — a row of metrics like every other tab, status icons, a table that fits its numbers,…
A fast check — safety.viz's required check in about six minutes, with every test still run
the evidence refresh run takes the same split and finishes in about the time the check does
the evidence refresh run takes the same split and finishes in about the time the check does
the required check runs every test once, on parallel runners, and answers in about six minutes
the required check runs every test once, on parallel runners, and answers in about six minutes
one R control — slim, in the same place on both tabs that use R, and a chip once R is ready
one status ladder — Qualified, Exploratory, Experimental, Prototype — with one label for the app and one for anything…
nothing looks broken — the defects found in the review of release 1.10, and the keynote's demo path walked by a test
every tab has a colour, and the first screen says where you are — tab colours, counts, a welcome line and the way back…
the standards this release set are written into the scaffold — the status ladder, the app's conventions and how a…
the standards this release set are written into the scaffold — the status ladder, the app's conventions and how a…
the standards this release set are written into the scaffold — the status ladder, the app's conventions and how a…
nothing looks broken — the defects found in the review of release 1.10, and the keynote's demo path walked by a test
nothing looks broken — the defects found in the review of release 1.10, and the keynote's demo path walked by a test
all files come in on the Data tab — gsm raw files recognised there, and the file box gone from the RBQM tab
all files come in on the Data tab — gsm raw files recognised there, and the file box gone from the RBQM tab
the RBQM tab reads at a glance — a row of metrics like every other tab, status icons, a table that fits its numbers,…
the RBQM tab reads at a glance — a row of metrics like every other tab, status icons, a table that fits its numbers,…
one R control — slim, in the same place on both tabs that use R, and a chip once R is ready
one R control — slim, in the same place on both tabs that use R, and a chip once R is ready
one status ladder — Qualified, Exploratory, Experimental, Prototype — with one label for the app and one for anything…
one status ladder — Qualified, Exploratory, Experimental, Prototype — with one label for the app and one for anything…
every tab has a colour, and the first screen says where you are — tab colours, counts, a welcome line and the way back…
every tab has a colour, and the first screen says where you are — tab colours, counts, a welcome line and the way back…
A demo app ready to show — the safety.viz demo app, clear to a first-time visitor and ready for the keynote
the biomarker app as a web app — a designed shell, navigation and Data page
the biomarker app as a web app — a designed shell, navigation and Data page
an R app for the biomarker charts — a user's own data, every statistic from the server's R, ready for Posit Connect
an R app for the biomarker charts — a user's own data, every statistic from the server's R, ready for Posit Connect
the biomarker app as a web app — a designed shell, navigation and Data page
an R app for the biomarker charts — a user's own data, every statistic from the server's R, ready for Posit Connect
the biomarker app as a web app — a designed shell, navigation and Data page
the RBQM tab runs on the study the other charts use — R makes gsm's raw tables from the standard domains
the RBQM tab runs on the study the other charts use — R makes gsm's raw tables from the standard domains
the RBQM tab — a study's raw files run through the gsm workflows in the browser and drawn as the three site-level charts
the RBQM tab — a study's raw files run through the gsm workflows in the browser and drawn as the three site-level charts
the standard gsm workflows run in browser R — adverse event rate by site, from raw files to Results rows, matching…
the standard gsm workflows run in browser R — adverse event rate by site, from raw files to Results rows, matching…
the biomarker app as a web app — a designed shell, navigation and Data page
the biomarker app as a web app — a designed shell, navigation and Data page
the RBQM tab runs on the study the other charts use — R makes gsm's raw tables from the standard domains
the RBQM tab runs on the study the other charts use — R makes gsm's raw tables from the standard domains
the RBQM tab — a study's raw files run through the gsm workflows in the browser and drawn as the three site-level charts
the RBQM tab — a study's raw files run through the gsm workflows in the browser and drawn as the three site-level charts
the standard gsm workflows run in browser R — adverse event rate by site, from raw files to Results rows, matching…
the standard gsm workflows run in browser R — adverse event rate by site, from raw files to Results rows, matching…
The fourth releases of the biomarker charts and their R package: the six charts as one Shiny app that reads your own files and takes every statistic from the server's R, with stills, steps to try and what has not been tested.
Why safety.viz's required check takes 22 minutes, five other layouts timed on GitHub's runners, the one that takes about 6 with every test kept, and what a security assessment of that layout found.
A design review of the released safety.viz demo app: one way to start R, one Experimental label that opens on a tap, polish for the RBQM tab and smaller fixes such as the black hexes, with mockups, eight decisions to make and every proposed change sized.
The settled design for release 1.11 of the safety.viz demo app, as decided on 9 October 2026: full-page desktop mockups, every decision with its answer, and the grouped change list the release is built from.
**See it move:** the [annotated v1.10.0 demo](https://jwildfire.github.io/obot.roadmap/reports/sv-v1.10-demo/) has captures, try-it steps and the detail behind everything below.
an R app for the biomarker charts — a user's own data, every statistic from the server's R, ready for Posit Connect
an R app for the biomarker charts — a user's own data, every statistic from the server's R, ready for Posit Connect
an R app for the biomarker charts — a user's own data, every statistic from the server's R, ready for Posit Connect
an R app for the biomarker charts — a user's own data, every statistic from the server's R, ready for Posit Connect
the RBQM tab runs on the study the other charts use — R makes gsm's raw tables from the standard domains
the RBQM tab runs on the study the other charts use — R makes gsm's raw tables from the standard domains
the RBQM tab runs on the study the other charts use — R makes gsm's raw tables from the standard domains
the RBQM tab runs on the study the other charts use — R makes gsm's raw tables from the standard domains
the RBQM tab — a study's raw files run through the gsm workflows in the browser and drawn as the three site-level charts
the RBQM tab — a study's raw files run through the gsm workflows in the browser and drawn as the three site-level charts
the RBQM tab runs the three newer site metrics (kri0013 to kri0015)
the RBQM tab runs the three newer site metrics (kri0013 to kri0015)
decide whether the demo app, with its RBQM tab, moves to open.gismo
decide whether the demo app, with its RBQM tab, moves to open.gismo
the RBQM tab maps raw files whose column names are not gsm's
the RBQM tab maps raw files whose column names are not gsm's
the RBQM tab draws the time series chart across snapshots
the RBQM tab runs the query, data entry and data change metrics (kri0008 to kri0011)
the RBQM tab runs the query, data entry and data change metrics (kri0008 to kri0011)
the RBQM tab — a study's raw files run through the gsm workflows in the browser and drawn as the three site-level charts
Three ways to turn the gsm.bio biomarker app from a default Shiny page into a professional web app, each drawn at desktop and phone width with the Data page in its three states, what each costs to build, the decisions inside them, and a recommendation.
What safety.viz v1.10.0 adds, shown on the live site: an RBQM tab in the demo app that scores a study's sites with gsm's own metrics, run by R in your browser on the study already loaded, with steps to try and the numbers to expect.
**See it move:** the [annotated v0.5 demo](https://jwildfire.github.io/obot.roadmap/reports/oa-v0.6-hub-v0.5-demo/) has screenshots, captures and the detail behind everything below.
**See it move:** the [annotated v0.6.0 demo](https://jwildfire.github.io/obot.roadmap/reports/oa-v0.6-hub-v0.5-demo/) has captures, try-it steps and the detail behind everything below.
the standard gsm workflows run in browser R — adverse event rate by site, from raw files to Results rows, matching…
the standard gsm workflows run in browser R — adverse event rate by site, from raw files to Results rows, matching…
the standard gsm workflows run in browser R — adverse event rate by site, from raw files to Results rows, matching…
the standard gsm workflows run in browser R — adverse event rate by site, from raw files to Results rows, matching…
RBQM in the browser — a study's raw files through the standard gsm workflows to the site-level charts, with R computing…
the RBQM tab — a study's raw files run through the gsm workflows in the browser and drawn as the three site-level charts
the RBQM tab — a study's raw files run through the gsm workflows in the browser and drawn as the three site-level charts
the standard gsm workflows run in browser R — adverse event rate by site, from raw files to Results rows, matching…
the standard gsm workflows run in browser R — adverse event rate by site, from raw files to Results rows, matching…
The group comparison chart opens on trends over time and drills down to one visit
The group comparison chart opens on trends over time and drills down to one visit
The group comparison chart opens on trends over time and drills down to one visit
bio.viz's site is laid out and styled as safety.viz's is
What the two releases change, shown running: sessions that write as the bot, release notes held to a length, a ruleset check that fails when it cannot look, and the tracker and Cost chart.
**See it move:** the [annotated v1.9.0 demo's note on v1.9.2](https://jwildfire.github.io/obot.roadmap/reports/sv-v1.9-demo/#v192) shows this patch release.
the group comparison's difference grid as a second opening view
the group comparison's difference grid as a second opening view
bio.viz's site is laid out and styled as safety.viz's is
bio.viz's site is laid out and styled as safety.viz's is
The third releases of the biomarker charts and their R package, both released: the group comparison as trend tiles, one biomarker over time with R's test under each visit, and one visit alone, with the numbers to expect and the detail behind the notes.
What safety.viz v1.9 adds, shown on the live site: biomarker charts in the demo app with R started in your browser on request, then the v1.9.1 patch, and the v1.9.2 patch, which rebuilds those charts as five and installs the demo app on your own machine.
The group comparison chart opens on trends over time and drills down to one visit
The group comparison chart opens on trends over time and drills down to one visit
The group comparison chart opens on trends over time and drills down to one visit
The group comparison chart opens on trends over time and drills down to one visit
Results out of R: static twins, RTF tables and batch runs
Results out of the browser: titles, footnotes, PNG and specifications
Stratified survival, and survival rows in the biomarker screen
Results out of R: static twins, RTF tables and batch runs
Results out of the browser: titles, footnotes, PNG and specifications
Stratified survival, and survival rows in the biomarker screen
Results out of the browser: titles, footnotes, PNG and specifications
Results out of R: static twins, RTF tables and batch runs
Stratified survival, and survival rows in the biomarker screen
Results out of the browser: titles, footnotes, PNG and specifications
Stratified survival, and survival rows in the biomarker screen
The second releases of the biomarker charts and their R package, on the live site: a cross-tabulation, survival curves split at a draggable cut, hazard ratios in the screen, downloads and saved views, and figures, RTF tables and batch runs from R.
**See it move:** the [annotated v1.9.0 demo's note on v1.9.1](https://jwildfire.github.io/obot.roadmap/reports/sv-v1.9-demo/#v191) shows this patch release.
The demo app hosts the biomarker charts, on the files and mapping a study already has
gsm.bio's statistics: the package, a synthetic biomarker study and every test the charts will print
R in the browser, measured: the connection a chart uses to ask R for a statistic
Group comparison, the tests: R's results on the chart, and the chart from R
Group comparison, the picture: bio.viz's first chart, drawn and drillable
The demo app hosts the biomarker charts, on the files and mapping a study already has
The first releases of the biomarker charts and their R package, shown on the live site: four charts whose every test is R's, R started inside the browser, and R widgets that keep their results offline, with the numbers you should see.
What the new kit in safety.viz is: one added export that hands a second chart library the 36 parts safety.viz's own charts are built from. What it changes (almost nothing), what it commits you to (those 36 become public surface in v1.9.0), and who uses it.
**See it move:** the [annotated v1.9.0 demo](https://jwildfire.github.io/obot.roadmap/reports/sv-v1.9-demo/) has captures and try-it steps for everything below.
The demo app hosts the biomarker charts, on the files and mapping a study already has
the standard domain set — a portfolio manifest naming the domains, the charts each feeds and the columns each needs
the basic app — load, map and view a study on one page, and as one file
the standard domain set — a portfolio manifest naming the domains, the charts each feeds and the columns each needs
the basic app — load, map and view a study on one page, and as one file
Group comparison, the tests: R's results on the chart, and the chart from R
Group comparison, the tests: R's results on the chart, and the chart from R
the standard domain set — a portfolio manifest naming the domains, the charts each feeds and the columns each needs
the basic app — load, map and view a study on one page, and as one file
several files in one domain — labs and vitals delivered apart are read together, each with its own mapping
Group comparison, the tests: R's results on the chart, and the chart from R
Group comparison, the picture: bio.viz's first chart, drawn and drillable
Group comparison, the picture: bio.viz's first chart, drawn and drillable
gsm.bio's statistics: the package, a synthetic biomarker study and every test the charts will print
gsm.bio's statistics: the package, a synthetic biomarker study and every test the charts will print
Group comparison, the picture: bio.viz's first chart, drawn and drillable
R in the browser, measured: the connection a chart uses to ask R for a statistic
R in the browser, measured: the connection a chart uses to ask R for a statistic
gsm.bio's statistics: the package, a synthetic biomarker study and every test the charts will print
R in the browser, measured: the connection a chart uses to ask R for a statistic
gsm.bio's statistics: the package, a synthetic biomarker study and every test the charts will print
R in the browser, measured: the connection a chart uses to ask R for a statistic
the basic app — load, map and view a study on one page, and as one file
Results out of the browser: titles, footnotes, PNG and specifications
Stratified survival, and survival rows in the biomarker screen
Group comparison, the tests: R's results on the chart, and the chart from R
Group comparison, the picture: bio.viz's first chart, drawn and drillable
Biomarker charts — compare groups and relate variables, with every test computed by R
the standard domain set — a portfolio manifest naming the domains, the charts each feeds and the columns each needs
the basic app — load, map and view a study on one page, and as one file
What v1.8.0 adds, shown on the live site: a demo app that loads your own study and maps its columns in the browser, nine long-standing requests on the existing charts, and a prototype of the Patient Journey Explorer.
How a user's own files reach the safety charts with nothing leaving the browser: what the page shows, the three things a user maps, how each chart says whether it can draw, and the four pull requests that build it.
Six chart types for comparing groups and relating variables in biomarker data, how they avoid repeating what the safety charts already draw, why R computes every test instead of JavaScript, and the eleven requirements that build it.
**See it move:** the [annotated v1.8.0 demo](https://jwildfire.github.io/obot.roadmap/reports/sv-v1.8-demo/) has captures and try-it steps for everything below.
AI narratives for the Patient Journey Explorer — drafted, cited, reviewer-accepted
AI narratives for the Patient Journey Explorer — drafted, cited, reviewer-accepted
AI narratives for the Patient Journey Explorer — drafted, cited, reviewer-accepted
the Patient Journey Explorer — one participant's whole safety course on one study-day axis
Widget_PatientJourneyExplorer() — the Patient Journey Explorer from R
Widget_PatientJourneyExplorer() — the Patient Journey Explorer from R
the Patient Journey Explorer — one participant's whole safety course on one study-day axis
the Patient Journey Explorer — one participant's whole safety course on one study-day axis
gsm.safety v1.2.0 — Widget_TimeToEvent + participant-profile surface (safety.viz v1.7.0 parity)
the GitHub-flows standard — auto-merge, both rulesets and Copilot review applied to every repository, with a check that…
the requirement-session skill watches its pull requests — event subscriptions, every Copilot thread answered, and the…
the requirement-session skill watches its pull requests — event subscriptions, every Copilot thread answered, and the…
the sweep routine and the orchestrator chat — an hourly sweep that notifies and wakes, and a Claude Project where…
the sweep routine and the orchestrator chat — an hourly sweep that notifies and wakes, and a Claude Project where…
the GitHub-flows standard — auto-merge, both rulesets and Copilot review applied to every repository, with a check that…
the GitHub-flows standard — auto-merge, both rulesets and Copilot review applied to every repository, with a check that…
Program operations — the orchestrator, sessions that answer their reviews, and one GitHub flow for every repository
Everything gsm.safety ships since v1.1.0 as one release, in four parts: the liver chart that stopped reading a baseline as a peak, the census rebuilt on metrics with the death count corrected, the last two widgets, and the FDA guide's thresholds as data.
The FDA safety guide's laboratory thresholds arrive in R as tested data, with the first functions that apply them to a lab table, the requirement matrix that keys all 22 of its figures, and the column-by-column check of what the demo data can and cannot draw.
The release candidates waiting on Jeremy's review, what each lets a user do, the order to take them in, a checklist per candidate with its pull request, notes and demo links, and his decision to ship gsm.safety as one v1.2.0 release.
FDA ST&F phase 0 — reference criteria as package data, first Derive_* functions, requirement matrix (gsm.safety v1.2.0)
Portfolio view — every safety.viz chart on one page, on the standard domain set
release, guide and demo — the download on the site, a user guide page, and the talk demo run end to end on a clean…
persistence and export — save and reload a study configuration from the desktop file; export a chart as an image and…
the single-file build — every script, style and the demo data inlined into one HTML file under a size budget
fit and load — mapped data flows into the portfolio, the chart status recomputes and the drill-down follows the mapped…
the mapping module — per-domain column mapping, auto-filled from the detected standard, editable, validated and saved
local file loading and standard detection — CSV and JSON through the browser, placed by domain and standard
a static twin for every interactive chart — the twins beyond objective 1's engines, and a 13 of 13 gallery
one settings contract per chart — a JSON schema shared by the safety.viz module and its gsm.safety function
A local-only desktop tool — one downloadable file, no network, no install
Data loading and mapping — a user's own files into the portfolio, nothing leaving the browser
Static parity — a static twin for every interactive chart, driven by the same settings
study-level settings and filters — arm, site and population across every chart, saved as one study configuration
the portfolio shell — every chart on one page, navigated by domain, with the participant profile as the shared…
the standard domain set — a portfolio manifest naming the domains, the charts each feeds and the columns each needs
one stylesheet for every surface he reads, and a check that keeps it that way
a step-by-step IQ for the new laptop — the chance to fix rather than replicate
config items are short local artifacts he can decide from, not a field form
the highest-authority file the agents read has no owner, no history and no review path
does the briefing get read aloud — and at what cost to the words
the session skills the fold replaces — one retirement, five re-homings, and the notes that say so
the weekly briefing — the sink that keeps the daily at nine lines
the admiral — a triggered manager, so finished work stops sitting and nobody has to remember to look
one-off tasks — a small fix files against its goal, without a requirement to justify it
the dashboard's two views, literally — a session table with filters, and a filterable metrics dashboard
goal #73 says what it is now — a current body, five workstreams, and a bounded backlog
his queue holds exactly three buckets — RC, decision, config — and the remainder is ours
the dashboard reads like a news feed, not an audit log — rebuild the sessions and Navigator pages
a requirement says whose decision it carries — an agent's inference must not read as his approval
the priced usage artifact refreshes on a cadence, or every surface says how old it is
a real "since you last looked" — record page views locally, or say plainly that we cannot
roadmap discipline is measured — the sweep checks all seven repos, and the audit stops reporting stale
the agent roster — every agent with its id, status, cost and roadmap impact
the worker closeout check — every agent finishes into a PR, a question, or a config request
the Navigator as operating officer — asks become requirements, and delivery is judged at closeout
every worker agent gets a permanent W0001 identity, allocated once and never recycled
standing sessions survive compaction — continuous state, an explicit setting, and a silent-loss check
Navigator phase 2 — working-set verification (split from #157)
the Operations Dashboard — my todo list, with blockers, where I answer decisions
audit stage history — re-gate the missing-design check to the moment work starts
scheduled autonomous sessions — the A2 promotion (recurring goal-work loop)
Navigator — standing roadmap-state verifier (file-writing only)
concurrent sessions — autonomous runs alongside @jwildfire's and each other
durable todo list — one permanent list per goal and overall, updated not rewritten
assignee and reviewer conventions for autonomous work — obot holds what's in flight, @jwildfire reviews
/oneoff — a single-agent lane that skips session startup but still logs to the roadmap
obot's Chrome tab sprawl — name, reuse, and clean up the tabs obot opens
fast session startup — sub-10-second responsive /session-init
idea queue — Siri/Reminders + hub Ideas Discussions intake with obot triage
weekly goal review — freshness sweep proposing linkages, labels, and orphan candidates
ideas-triage v2 — stronger model, richer context, bias-to-promote
FDA ST&F phase 0 — reference criteria as package data, first Derive_* functions, requirement matrix (gsm.safety v1.2.0)
FDA ST&F phase 0 — reference criteria as package data, first Derive_* functions, requirement matrix (gsm.safety v1.2.0)
Chart coverage — every figure in the FDA safety guidance, static and interactive
FDA ST&F phase 1b and the Kaplan–Meier family — static coverage 22 of 22
FDA ST&F phase 1 — four static engines, twelve figures in gsm.safety
gsm.safety v1.2.0 — safety.viz v1.6.0 widget parity + parity guard
gsm.safety v1.2.0 — Widget_TimeToEvent + participant-profile surface (safety.viz v1.7.0 parity)
SafetyCensus stays and is rebuilt on metrics and reports — design first
open.csr v0.4.0 — one study, and all thirty-one reference displays
open.csr — the sidebar says where every number came from (Data and Metadata sections, explicit links)
open.csr — the explorer reads as the pipeline, and every element shows its flow
nobody can write to the board — the App is forbidden and the guard denies the only credential that works
**See it move:** [Requirement Sessions: the Mid-October Plan](https://jwildfire.github.io/obot.roadmap/reports/goal-sessions-plan-2026-09-10/) — the operating model this release installs, and the…
**See it move:** [Requirement Sessions: the Mid-October Plan](https://jwildfire.github.io/obot.roadmap/reports/requirement-sessions-plan-2026-09-10/) — the operating model this release installs, and…
Five objectives — chart coverage, static parity, a portfolio view, data loading, a local desktop app — laid out as issue trees with definitions of done, two concurrent cloud sessions, and what the site must show each Friday before the R/Pharma talk.
The day the autonomous prototype was shut down and the program moved to requirement sessions in the cloud under a rigid issue contract. One plan page published, two pull requests opened, three…
What Can We Actually Do With AI Agents Right Now? Quite a lot. Are clinical trials ready to implement these tools? I’m not so sure…
open.csr — the explorer reads as the pipeline, and every element shows its flow
open.csr — the sidebar says where every number came from (Data and Metadata sections, explicit links)
open.csr v0.4.0 — one study, and all thirty-one reference displays
The running list of calls waiting on you. Each page states the situation, the options with what each one costs, and a plain recommendation — these and release-candidate pull requests are the only two things you review.
For each item in the open.csr sidebar — Documents, Displays, Text, Values, Templates — what goes in, what the pipeline does, what comes out, where you can see it, and the refactor that adds Data and Metadata as sidebar sections linked from every artifact.
What v0.4.0 changes, change by change: the report describes one study, and five more of the reference report's displays agree with it cell for cell. Captures from the candidate's own deployed tree, with a way into each one.
A colleague's replication assessment of open.csr, checked against the data, with the corrections that matter. Then a four-release roadmap that closes its gaps, and the first release broken into nine issues that could start tomorrow. Four calls.
One development month left before talk preparation. This is a plan built around what the keynote needs rather than what is next in the queue, including what to deliberately not build. Four calls.
the change chip says a site improved when its peers got worse
A design with no interface: two commands write the mapping, the error message is the whole surface, and the mapping lives in git as a reviewable diff. Measured on a CRO delivery where the matcher's most confident guess silently empties eight metrics.
What gsm.mapping already solves for loading your own study data and what is genuinely missing, established by running the packages: the readiness check never reads the source_col key that fixes the problem it reports, so it can never go green.
A six-question setup interview for loading your own study into open.gismo, argued from four runs: what the name detector misses, what value profiling recovers, where it is confidently wrong, and how a half-mapped study doubles a site's risk score.
A design direction for loading your own study into open.gismo: never reject a folder, profile everything in it, and answer in one costed document — priced in the charts each gap turns off. Measured against a six-file CRO delivery.
A design direction for loading your own study into open.gismo: one dense screen showing all 126 required columns beside your delivered ones, four dispositions per row, and the price of every decision printed while you make it.
Which of four data-loading designs open.gismo should build, and why: read the folder first, price every gap in charts rather than columns, keep a worksheet behind the report for the fortieth transfer, and refuse to run while the identifiers do not join.
Getting a real study in means satisfying 126 columns by hand-editing YAML, and nothing checks that the subject identifiers join — a design for the surface that fixes both, with clickable mockups and how thirteen other platforms let a user in.
Recording that a person looked at a finding, and telling them what changed since. The published snapshot names turn out not to be stable — one has meant four different datasets — so a review is keyed to the data itself. Six questions.
Most of the analysis plan is already in open.csr, dispersed across the seventeen displays citing it by section. Whether a display can shell from its spec alone decides the rest — so the shells were built, not argued.
The release that fills the report. v0.2.0 carried six safety displays over a single study; this one carries **twenty-six**, and the twenty new ones are the parts of a clinical study report that were…
A study produces about a dozen documents and open.csr could build one. The second shipped tonight: an ICH E3 synopsis built from the same numbers as the report — plus which of the rest the R Consortium pilots could seed, and why their licences say almost none.
A walkthrough of how open.csr turns a study dataset into a clinical study report: the four kinds of file you edit, how a table's numbers get bound into sentences instead of retyped, and the places the shipped code and the written contracts disagree.
Nine years of feature requests left in twelve retired safetyGraphics chart trackers, sorted into what the new library should build, what it already does, and what died with the old framework — plus nine cheap wins that clear thirty-three of them.
where the data-coverage numbers come from — design before the thirteenth metric
Two of the thirteen safety charts have no R binding, and the reason on file is wrong about both. One is already on screen inside eight widgets that ship today; the other builds without complaint and then draws an error. Four questions, every option run.
The liver chart in R was reading two participants' own baseline as their on-treatment peak; this release stops it. Both chart bundles run side by side on the same data, with the records that caused it and the two new charts the release adds.
The safety overview said four deaths on a study with thirteen; this release rebuilds every figure on that page as a metric that publishes its own numerator, denominator and provenance. The before and after run side by side on the same study.
The thirteenth census number is the one that could not be built: no standard domain says which visit a lab result belongs to. Five questions, every option measured on two real studies, and a recommendation that needs nothing from anyone else's package.
dispatches come from the ranked head, and drift from it is visible
a worker that stops — finished, stalled or dead — wakes the Navigator
every surface renders honestly on a machine with no history — before the move, not after
the ranked head as cards, and one day of re-ranking he can watch move
the ranked head as cards, and one day of re-ranking he can watch move
merging is not deploying — the checkout tracks main, and consumers restart when it moves
The ten requirements ranked next and the eleven waiting behind them, each with the one line saying why it sits there — and a player that replays every re-rank of the single day the order has existed.
The framework builds clinical charts well and nobody has written what version 1.0 contains — the evidence for both, ten items to replace the scaffold head of the queue, the charts worth making, and the smallest weekend build that would test it.
one stylesheet for every surface he reads, and a check that keeps it that way
an open decision artifact has an episode he can answer from a car
an answer you clicked that nothing applied has to reach someone
he answers a decision from the car, by voice, without a screen or an issue number
a thing he asked for has a delivery state, and something notices when it goes quiet
what comes next is written down — a ranked head, tiers below, before any clock runs
a config item's claim is checked on a cadence, not asserted on the day it was filed
work that never reached GitHub is invisible to every check we run
a kill this house reports is a kill it confirmed — the liveness anchor reports success while the session survives
A proposed dashboard tab that reads a whole day of agent merges as one summary — what landed in the harness versus what will ship to users, which agent landed it, and which branch it flows to — with 2026-08-20's real numbers rendered inside it.
Eleven of the thirteen branches the merge policy governs have no protection at all, including the one holding the merge policy. Three options for what to lock down, what each costs the agents, and why the recommended one does not slow a single merge.
Both packages have been off CRAN since March. The fix is prepared, checked and audited, and the tarballs sit on one laptop. Submitting them is outside every lane we have and is yours to take — here is what it costs and the order that cannot be reversed.
You move to a new laptop this weekend. Eleven questions come first, starting with whether the old machine is copied across or rebuilt clean — plus two things already true: nothing here is backed up, and one key cannot be recovered.
Decided 2026-08-20: he approved all six recommendations. The safety census is rebuilt so its numbers can be validated one at a time — the shape, what each answer commits the build to, every defect the review found, and what moves on the live demo.
the usage analytics run on a cadence, and the page says when it last did
safetyGraphics and safetyCharts back on CRAN — remove the Tendril chart
an open decision artifact has an episode he can answer from a car
a step-by-step IQ for the new laptop — the chance to fix rather than replicate
what comes next is written down — a ranked head, tiers below, before any clock runs
SafetyCensus stays and is rebuilt on metrics and reports — design first
he answers a decision from the car, by voice, without a screen or an issue number
a config item's claim is checked on a cadence, not asserted on the day it was filed
config items are short local artifacts he can decide from, not a field form
Who does what in the agent framework: one human, three named roles, a crew of numbered workers, and the script that watches them all — with the failure that created each role, dated 17 August 2026, and the two seams still open.
Seventy-four roadmap requirements end with a line asserting that @jwildfire reviewed them, and nothing anywhere records that he did. Fifty are still open and thirty-eight sit one stage from development. Counted and named here rather than rewritten.
33 changes landed across 2 repos overnight, most recently: A session record destroyed on 27 July is back, named by its boundary. 1 release candidate and 3 decisions are waiting on @jwildfire.
All four questions answered. He approved the rewritten done-condition and the five workstreams on 18 August, dictated from his phone after listening to the audio episode — the first decision this program has taken through a brief rather than a page.
Decided 2026-08-18, against the recommendation: SafetyCensus() stays, and is rebuilt on the gsm metric framework as metrics plus a report. The panel's verified defects all stand — he is answering a different question, about what the census is for.
**See it move:** the [annotated v1.1 demo](https://jwildfire.github.io/obot.roadmap/reports/gs-v1.1-demo/) walks the new metrics running live on the [DEMO-301 study…
One page for the Navigator: what the session is, how it launches, what it reads and writes, what it does on day one, and what it must never do. Folds in the worker-closeout and supervision pages, and six calls you still have to make.
Three working redesigns of the roadmap page, built on live data and decided the same day: the queue becomes the front page, the wire sits one click behind it, and the current inventory page stays on as the catalog.
Not yet — and here is the finish line. Five gates between today and unattended overnight sessions, what each one costs, and the check you can run yourself to tell whether it has been met.
Jeremy went to bed with a six-item list and woke to eight merged pull requests, but the night's real finding was not in any of them. Nine separate times, in unrelated subsystems, an operation…
What gsm.safety v1.1.0 adds, annotated: the participant-level metrics phase and SafetyCensus, captured running live on the DEMO-301 study site.
Which safety.viz charts have an R wrapper and which do not: four have none, and all nine that exist draw through a bundle two releases behind — including the liver-injury chart, which still renders a calculation corrected months ago.
What open.gismo v0.2.0 delivers, annotated: the local-first engine and the study site, captured live from the DEMO-301 deployment.
What v1.7.0 adds, annotated: the Time-to-Event Explorer, captured live from the release-candidate build, with a way into each behaviour.
You asked to be interviewed about what the app should be. This is the research on how to actually extract intent someone holds but cannot state, and the four calls in the interview method that shipped as a skill.
A formal list for the recurring category no lane holds — fully-specified fixes whose only missing ingredient is @jwildfire's keyboard. Four calls: the scope test, where it lives (security first), how agents write it, and how it reaches him. BL1–BL4.
Where your decisions get written down now that the discussion board is not the place: whether an approve button on the published page is worth building, whether one chronological log should exist, and whether the roadmap needs another tracker.
Recorded: @jwildfire kept both delegation lanes — siblings for deliverables, subagents for answers — because subagent results flow through prime's own context. His words, the rationale, and the one-line routing rule now written into the skills.
The operational-vs-clinical governing principle keys on a per-repo classification @jwildfire never enumerated. The clear cases recorded, the three ambiguous repos put to him as numbered options, and the CI precondition stated plainly.
A standing prime has no session start and no session end, so the bookend-built scaffold needs a verdict per piece, a new trigger for the wrapup's duties, and a morning briefing designed from why the last one failed. Five calls, M1–M5.
The nightly roadmap audit judged by your own test — if a finding needs your attention it probably is not a good rule. Five of twenty-two rules fail it, and two rules produced eighteen of the twenty-four findings that landed in your lap.
Published as "go, after four fixes" — corrected to three. The permission denial that justified the fourth was re-checked and did not exist, and the work it was blocking has merged. What still breaks at 3am, and the pre-flight checklist before enabling.
Forty-one agents finished in two days and three left no recoverable trace of what they did. What a worker must produce before it may call itself done, how the Navigator would check it at closeout, and why telling the agents apart is harder than it looks.
You asked whether the worker agents need a manager. Measured from 36 hours of real sessions: what a watcher can actually see, what the same job costs as a script versus a standing session, and whether it gates next week's scheduled runs.
Kaplan–Meier survival curves for safety.viz: the statistics stated exactly, the data contract, the at-risk and cumulative-events table, and how the module gets verified against a reference implementation.
**See it move:** the [annotated v1.7.0 demo](https://jwildfire.github.io/obot.roadmap/reports/sv-v1.7-demo/) walks the new chart with captures and try-it-yourself steps against the live demo.
**See it move:** the [annotated v0.4.0 demo](https://jwildfire.github.io/obot.roadmap/reports/oa-v0.4-demo/) walks each change as a real terminal capture with the command that produced it.
What v0.4.0 changes about running a session, shown as before/after terminal captures with the exact command to reproduce each one.
What v1.6.0 adds, feature by feature: annotated captures from the release-candidate build, with a way into each one.
The app goal has blocked itself for a month on a plan that the design pass and the demo build overtook. Four calls: what replaces the stale plan, what the requirement set is, what version one means, and whether the demo study is the fork template.
The demo study's published branch is 302 MB, but that is not what a fork downloads. What was actually measured, six options with what each costs, and two cheap changes plus a bound on future growth.
An agent can still quietly change the rules that govern agents, and twice one has. One real decision on what the merge tool must demand when a change touches a guardrail file, plus two sign-offs — the fix itself is already written.
The merge lane was not broken. One invocation form falls outside the permission allowlist and is denied by the auto-mode classifier roughly two times in three. Evidence from 99 invocations, and the one-line fix.
obot.agent's integration branch is main, so there is no branch left to open a release PR from. Four ways to give it an RC that is an actual PR, with what each costs and how v0.4.0 ships under it.
The standing concierge session will eventually lose its memory, and nothing currently makes its “durable state lives elsewhere” claim true. Six calls on what it writes down, who keeps that honest, and what it reads to get back to useful.
**See it move:** the [annotated v1.6.0 demo](https://jwildfire.github.io/obot.roadmap/reports/sv-v1.6-demo/) walks each update with captures and try-it-yourself steps against the live gallery.
Three goal lanes ran in parallel tonight and each ended somewhere different: charts shipped a release candidate, autonomy cleared a five-PR backlog and then cut its own release, and the app goal…
Jeremy opened the session with the obvious question — why did the init still take five minutes, on Fable, when last session shipped the fix? — and the answer became the session: the fix was live, but…
The responsiveness framework merged at 22:17 and met its first live test nine minutes later — and failed it usefully. This session ran /session-init as hub#91's acceptance run, watched it miss the…
📊 Session report: skipped — the generator keys on today's date and this wrapup ran 08-04; the missing `--date` flag is [oa#32](https://github.com/jwildfire/obot.agent/issues/32).
Making the per-repo write policy mechanical: checking a change's files against the guardrail carve-out, settling whose copy of the policy the tool reads, and making the dry-run check stop implying a merge will succeed.
Session 🦾🤖 2026-08-01 140-design (job 4a6f66b8), 08:00 → ~09:40 EDT — the fourth A1 autonomous run: init selected the carried [#140](https://github.com/jwildfire/obot.roadmap/issues/140) design…
Session 😺🤖 2026-07-31 2 (job 244f8e81), 08:57 → 09:45 EDT — the session where the orchestrator got corrected live. The init ran 60–90 seconds against…
Session startup takes ten minutes and the network accounts for four seconds of it. Where the time actually goes, and a three-tier design that paints a usable list first and reconciles behind it.
Session 😺🤖 2026-07-31 (job 3d51bfb5), 07:40 → ~08:45 EDT — the session-framework improvements session: three roadmap artifacts shipped to the review gate in about an hour, all by one Opus sibling…
Session 🦾🤖 2026-07-30 hep-polish (job c07d76af), 21:03 → 22:35 EDT — the third A1 autonomous run, and the first **fully hands-off pass**: no goal directed, nobody watching, init → selection → build…
Every claimed defect in gsm.qtl's QTL report module, re-run from a clean R session against the package's own bundled inputs: four confirmed with captured repros, one withdrawn, and the coverage gap underneath all of them.
Session 🦾🤖 2026-07-28 nep-explorer (job f1e062e6), 11:33 → 12:40 EDT — an autonomous `--auto` run opened with the directive "build the nep-explorer in safety.viz". It did not build it, and that was…
Mockups for the app's top-level pages, built around what changed since the snapshot you last reviewed: a change ledger, the riskiest-sites table beside a funnel plot, and three conventions to settle once. Every number is real demo-study data.
The design record behind the safetyGraphics replacement — three ways to frame the app's navigation with a recommendation, and the framework underneath: pipelines, config in a forkable study repo, and snapshots as the only data interface.
The kidney-safety explorer as a chart module: the KDIGO staging logic re-derived from the original app, the measurement-units contract that decides whether the numbers are right, and what the patient-profile phase adds.
Session 😺🤖 2026-07-28 (job 96636d0f), 09:34 EDT 28 July → 07:00 EDT 29 July — opened as a design session for [goal #79](https://github.com/jwildfire/obot.roadmap/issues/79) and closed with the app…
Session 😺🤖 2026-07-27 2 (job dba06e09), ~20:25 EDT → past midnight — billed as roadmap cleanup, delivered as live product iteration. The kickoff directive asked for one thing: an expandable goal →…
Two-part research on the clinical study report side quest: eighteen products examined and none tracks a display change request as an object, plus a framework that makes the request a versioned artifact reviewed against a computed data diff.
What v0.2.0 adds, feature by feature: annotated captures taken from the live site, with a way into each one.
The second release for `open.csr` adds several foundational features, including the ability to edit text in the app, an initial rtf implementation for tables, a "values" section and various UX…
Session 🦾🤖 2026-07-27 csr-research (job b23e47f6), ~11:30 → ~14:00 EDT — a directed `--auto` research session, and the first whose deliverable is a roadmap direction rather than code. @jwildfire's…
A clickable twelve-step protocol for the open.csr text-block editor, running the repository’s own gate and diff code against a real ARD: every step states what it expects and then checks itself.
Two renderers join the gallery and one of them is not a chart. The hepatic tooling grows a third view and a companion module for trials that enrol participants who already have abnormal liver tests,…
obot.agent can now run a full development increment without a human in the loop — and the roadmap now feeds it.
The first release of **open.csr** — an open-source Clinical Study Report builder. This is an early prototype, and it is labelled as one: it demonstrates a closed loop end to end on real public data,…
Session 🦾🤖 2026-07-26 (job 119b6fc2), 22:16 → ~02:00 EDT — the second autonomous run, and the first with @jwildfire live in the loop. Launched as an `--auto` overnight session focused on open.csr…
Three working prototypes that turn the nightly roadmap audit into a queue you can clear in one sitting, measured against the live page: 9.3 screens down to 2.0. Option B was picked the same day.
Evidence for typing a prompt into the live session dashboard and reaching a running agent session — both delivery lanes built and verified end to end, the security argument, and why it was parked rather than shipped.
Thirteen commercial and open-source safety platforms scored against this portfolio on sixty-three capabilities. The charts are competitive and ahead in two places; nine of ten review-workflow capabilities do not exist here at all.
What v1.5.0 adds, feature by feature: annotated captures taken from the live dev site, with a way into each one.
The release face of safety.viz v1.5.0 — the participant profile, the migration Sankey and ALT waterfall, and the liver-chart follow-ups — in the order to review them, ending in the two actions that promote the release.
The roadmap now largely runs itself: it triages its own ideas, audits its own consistency, and shows what everything costs.
Session 😺🤖 2026-07-24 #4 (job 10266b89), 2026-07-24 22:17 → 2026-07-26 00:50 EDT — the release-train day. Eight overnight workstreams fanned out at bedtime; a checkbox review guide turned their…
Three priorities at breakfast — move the goals into issues, fix the ideas flow, refresh the roadmap homepage — became three parallel plan-first siblings, three approved plans, and by ten at night…
The maiden voyage. Hours after the autonomy core merged, @jwildfire typed `/session-init --auto` — and obot's first fully autonomous session ran launch-to-wrapup in 53 minutes with no human in the…
Every open issue in the portfolio plotted against the standing goals — and the sixty-one percent that roll up to nothing, which an agent picking its own work cannot see at all. Three new goals proposed, with rosters.
A read-only dry run of the automatic idea-triage lane under two models, with each one's classification of the same seven threads side by side and the cost per run — the calibration behind the model switch.
Interactive UX mockup for participant profile v2: the profile in a right-hand rail, the four surfacing options side by side, expand-to-full-screen, and the AE summary + AE timeline tracks on the lab chart's study-day axis.
Capturing an idea in seconds from a phone and having it triaged without you: the voice-to-reminders lane, the discussion queue it feeds, the automatic pass that classifies and replies, and what stays private.
Standing goals move from files into hub issues so direction can be edited without a pull request, with one page per goal on the site and a thin registry binding goals to how an agent selects work.
The second version of the participant profile: moving it from a dock under the chart to a right-hand rail that opens on click, and adding adverse events to what a person's page shows.
There is no supported way to talk to a running agent session. Four delivery lanes measured against each other, and the recommended one — a file-based inbox with the session transcript tailed as the reply stream.
One session, two days, four new framework capabilities. This session opened 07-23 morning and wrapped 07-24 — and in between, the obot session framework grew remote control on every agent, a…
First stable release: gsm.safety is now the R home of the safety.viz interactive chart library, mirroring the gsm.kri / gsm.viz architecture.
The night Stage 2 closed and Stage 3 was born. @jwildfire merged and tagged **gsm.safety v1.0.0** — the all-inclusive "safety.viz released" gate — while the DILI toolchain finished its…
A collaborator asked why clicking a point does not open that person's whole story the way the original eDISH did. Four ways to surface one shared participant profile from every chart, with a recommendation.
How an agent session picks its own work and ships it unattended — goals to increments to sessions, the per-repo permission matrix, three autonomy levels, and the kill switches that stop a run.
Two displays from the abnormal-baseline liver-injury paper — a migration Sankey and a modified ALT waterfall — and how the existing liver chart splits to make room for them without breaking it.
The Gallery's QT example data is now internally consistent: QTcF and QTcB are derived directly from QT and RR in the vendored `adeg` demo dataset, replacing the pilot data's implausible rederived…
The morning the merge gate got a real door. Last night's blocker — an agent mechanically unable to execute a merge @jwildfire had already approved — became this morning's scaffold: a policy-gated…
Where every workstream stood after seven weeks, what shipped, what was stuck at a gate, and the gaps between the public roadmap page and reality — the status read that set the stage model and the autumn arc.
The FDA's standard safety tables and figures guide, inventoried display by display: open-source R has largely solved the tables and barely touched the figures, and the untouched figures are almost exactly the chart set we already own.
Qualification evidence for the first gsm.safety release: nine R chart widgets rendering live on demo data, eighty-two issue-linked tests, and the matrix tying every test to the requirement it proves.
Build the FDA safety figures the open-source ecosystem is missing, borrow the table logic that already exists, and share the guide's normative rules between the static and interactive charts rather than duplicating them.
A whole-day entry with four hands on the wheel. @jwildfire published Diary #7 in the morning, closing a loose end that had carried since the 17th. A midday design pass turned the FDA Standard Safety…
A review-clearing morning, run live. The session opened with a purpose-built twist on the init: instead of the usual priority list, @jwildfire asked for his review backlog as a one-page artifact —…
The library's first cardiac-safety renderer joins the gallery, hep-explorer gains a baseline-referenced view for subjects with abnormal baseline liver tests, and the Gallery is one click from any…
Five improvement ideas from colleagues scoped into buildable requirements — the hepatic composite plot, the QT explorer, renal safety, benefit-risk and recurrent adverse events — each with feasibility and where the risk sits.
I demoed safety.viz for safetyGraphics’ clinical lead, Jim Buchanan, this week and asked if he had any improvements in mind, and no surprise, he had a few! Jim promptly sent a handful of references,…
A colleague's wish list became a shipped release in one evening. safetyGraphics clinical lead Jim Buchanan sent five improvement ideas; the session opened by turning them into a grounded [feasibility…
Each renderer's **Test evidence** page now shows the reviewed requirement **text** next to its ID, so you can see exactly what every test verifies without leaving the page. Previously the Requirement…
The first table-first renderer joins the library — an interactive Adverse Event Explorer.
Can the kidney-safety explorer become a Chart.js module? Yes — it is the kidney twin of the liver chart and reuses its proven pattern. The clinical logic, which source to migrate from, and the data gap that is the real constraint.
This release turns day-to-day session management into a first-class framework: a live dashboard with frozen per-session reports, lean session bookends built on a shared scratchpad, an enforced…
Three siblings, one release train: **[safety.viz v1.3.0](https://github.com/jwildfire/safety.viz/releases/tag/v1.3.0) shipped** — ae-explorer reviewed, rebuilt, merged, and released in a single…
Release day for the **session framework**: [obot.agent v0.2.0](https://github.com/jwildfire/obot.agent/releases/tag/v0.2.0) ships the session hub, lean bookends, a three-layer scratchpad heartbeat,…
The eDISH liver-safety explorer joins the library — with a clinical guide that teaches how to read it.
😺🤖 7/10 Friday night — safety.viz v0.1.0 released end to end: docs site, three-tier Pages publishing, staging-review fixes (12M/$20) — diary
Release day for the blog: **[R/Pharma Diary #6 — "Obot v3: How I Used a Billion Tokens in a Weekend"](https://jwildfire.github.io/2026/07/13/obot-v3-billion-tokens.html) is published**, built the…
The release day. Everything the morning entry set up — the eDISH port at review, the RC1 spawn, the blog drafts — converged by evening: **safety.viz v1.2.0 shipped** (eDISH + a full 19-figure…
Full drafts of three R/Pharma diary posts for review before publishing — introducing safety.viz, the move to Claude Code, and the overnight-run case study — with the running order and how each one gets published.
Every open issue on the original liver-safety renderer triaged against our rewrite, with ten worth landing this release — plus a clinical guide tab that teaches the eDISH evaluation workflow.
The architecture decision and phased plan for turning open.gismo into a replacement for the safetyGraphics app: local-first compute with GitHub demoted to optional publisher, scored by a four-judge panel, with a working prototype.
The histogram gets a whole-dataset view, and the example data now comes from a scripted pharmaverse pipeline.
**safety.viz is a charting library for monitoring clinical trial safety.** Point any of its six interactive charts at your study data and review it in the browser: filter, group, zoom, and click…
First release of obot.agent — the obot program overlay on the [gsm.agent](https://github.com/Gilead-BioStats/gsm.agent) harness, rebuilt from safety.agent per the signed-off requirement…
My last post laid out the plan for {safetyGraphics} v2 and ended on the obvious question: can we actually do this? Here’s the first big answer: safety.viz is live.1 AI collaboration note — this post…
The first **overnight ultracode experiment**: @jwildfire triggered two ⚡️🤖 stretch-goal jobs at bedtime and went to sleep; the orange lead (😺🤖 07-12 2, job `dad0affc`) watched them run and had a…
The evening session ran as an orange **lead** (😺🤖 07-11 2, job `ce8f336e`) orchestrating four green 👯🤖 siblings — and shipped **three releases in one evening**: **safety.viz v1.0.0**,…
What the agent-harness repo became in a single day — the rename, the three-layer overlay that stopped skills living in two places, and the readiness scorecard against the requirement's success criteria.
The night two releases shipped and four agents were building, the public roadmap page showed nothing in development. How the page decides what is in flight, what actually happened, and the corrections proposed.
Five live mockups of the safety.viz front page, each in the real theme with real charts, taking a different position on what the first screen claims the site is for. Recommendation included.
Rebuilding the project's agent harness as a thin overlay on the shared upstream one: what was duplicated, where each skill should live so none lives in two places, and what goes back upstream.
A live dashboard for what a working session is doing, and a frozen report when it closes — where the data already lives, what each panel shows, and how the diary entry gets written from it.
First release of safety.viz — the consolidated Chart.js charting library for clinical safety graphics (safetyGraphics → gsm modernization, project P004), mirroring the gsm.kri ↔ gsm.viz architecture.
Second release of `obot.roadmap`. Where v0.1 established the hub, v0.2 closes the design phase for the portfolio's founding requirements: designs for the safety.viz library, the histogram pilot, the…
The morning session (😺🤖 2026-07-11, job `08c20082`) turned the obot migration from a plan into a review queue: **safety.agent is now obot.agent** end to end — renamed, audited against gsm.agent,…
A Friday-night sit-down that ran into Saturday morning, worked across two agent sessions and a fleet of subagents — and shipped the thing the whole month has been building toward: **safety.viz v0.1.0…
A short evening session on the session workflow itself, run as a live test of the day-old kickoff skills: each invocation drew immediate feedback from @jwildfire and a same-hour fix. The kickoff list…
The documentation-site lane closes its build phase: @jwildfire merged the evidence-pipeline and API-reference PRs in the morning, and by afternoon the site build itself (#7) was up as a draft PR —…
The rule that a chart is not done until it is published: a documentation site carrying a live demo, a test-evidence page and an API reference for every renderer, plus the branch model that makes the gate enforceable.
The histogram lane ships its code — scaffold merged, module extracted — and documentation becomes the delivery bar: requirement #21 (safety.viz documentation site) was filed, designed, and signed off…
The case for stopping maintenance of two parallel agent harnesses and rebuilding this one as a thin project overlay on the shared base — the layering model, and the two requirements it became.
The proposal for a session that picks one or two lanes, spawns a background agent for each, tracks them and closes the day by writing the diary entry — kept as the point-in-time record of a design later absorbed into the session hub.
obotclaw goes live and the pipeline fills in behind it — the App registered, credentialed, and green on both acceptance paths; designs #1/#2 signed off with sub-issues filed and the safety.viz…
How nine legacy safetyGraphics charts become one Chart.js library consumed by R widgets: the repo layout, the contract every chart module honours, the stack, and how each requirement is traced to a test.
The safety histogram as the pilot for the whole renderer-migration recipe — extracting the working chart into the shared library, then binding it as an R widget. The two slices that prove the pattern.
One GitHub App to replace the retired machine account: how agents and Actions each get a scoped, short-lived token, where the key lives, and what an agent calls to mint one.
First release of `obot.roadmap`: the development roadmap and project homepage for the obot open-source safety-graphics portfolio. This release establishes the requirement workflow and consolidates…
The hub goes live — PR #8 merged, v0.1 released, News feed added — and the obot GitHub App requirement moves into design.
The ten-chapter record of how autonomous development was attempted, why the scheduled-agent route failed, and the pivot to interactive Claude Code sessions — the report the current way of working came out of.
How the archived public hub became this repo — the directory layout, what was migrated and with what provenance, what happened to the legacy projects, and which pages are generated at deploy time rather than committed.
So, here’s the plan for {safetyGraphics} v2.1 AI collaboration note — I outlined this post in June. Claude Code (using Fable 5) expanded my outline into a draft, and I reviewed and edited the result…
obot.roadmap becomes the project's memory: hub-migration requirement, HTML design doc, and full site implementation.
As part of the keynote, I want to see if AI can update and modernize {safetyGraphics}. {safetyGraphics} was the first big open-source project I worked on, so I also want to talk a bit more generally…
Thursday was a framework-hardening and publish-hygiene day: local autonomous-work instructions were tightened around Development queue freshness, public GitHub state was rechecked, and the Hub…
Wednesday was an operations and queue-hygiene day: the PR review watch found no new agent-side implementation changes to make, the Development block stopped instead of starting stale work, and P008…
Tuesday was a consolidation day: the Hub navigation was cleaned up after the Paperclip/P009 report burst, public gsm.safety planning issues received agent analysis, and the open human-decision queue…
What supervised runs changed for the operator: automation moved from “a prompt was triggered, assume it is running” to a durable record with run ids, heartbeats, transcripts and a recovery path.
Monday turned the autonomy work from runner proof into a guarded Paperclip production path: P009 finished its runner/action evidence, P008 completed local Paperclip PM and Dev pilots, and the…
Third pass, written after scheduled agent runs kept dying before they started: the blocker was the scheduler itself, not the prompts or the portfolio framework. Superseded by the Paperclip evaluation.
Whether to buy orchestration instead of building it — cost, security, stack fit and competitor readiness for the Paperclip agent platform, ending in a bounded pilot rather than a migration. Superseded by the framework report.
Sunday shifted the autonomy roadmap from broad framework design into execution-first reliability work: the Hub reports now explain the recommendation, P009 proved a supervised Codex-native runner…
First pass at making agent work cycles reliable: what other systems do about run ledgers, heartbeats and issue-to-pull-request handoffs, and why the answer was the smallest custom layer rather than a platform. Superseded by v2.
Second pass: a hybrid model with the chat assistant owning intake, scheduling and liveness while a coding agent owns analysis and execution, laid out layer by layer with the evidence each one leaves. Superseded by v3.
A review that caught the project-management cycle selecting development work without auditing the portfolio first, and the audit checklist added to stop it happening again.
Why spawned agents kept returning ids that then did not exist: a mixed orchestration stack, no durable proof an agent was alive, and what has to be standardized before background agents can be relied on.
The acceptance evidence for supervised work sessions: each criterion, the artifact that satisfies it, and whether it was met — the sign-off record for that piece of the runner.
Saturday focused on turning obot's autonomous workflow from ad hoc sessions into a supervised, GitHub-native operating loop, while keeping P004 and the keynote work visible and reviewable.
How autonomous the development workflow actually was in June 2026, and the case for a promoted work queue, explicit authority levels and test-first migration gates instead of a setup tuned for reactive chat.
Friday was an autonomy-framework day. P007 moved from concept into a working public roadmap/autonomy loop, and the first autonomous P004 work blocks produced concrete Safety Histogram evidence…
Thursday was a quiet public implementation day: the June 3 Hub briefing and Pages deployment succeeded, the Hub was safely reconciled with `origin/main` before publishing, no new public…
Wednesday kept the public queue steady: the June 2 Hub briefing and Pages deployment succeeded, no new public implementation commits landed in the tracked active repos, and the highest-leverage next…
Tuesday kept the public implementation queue steady: the June 1 Hub briefing and Pages deployment succeeded, no new public implementation commits landed in the tracked active repos, and the…
Monday was a quiet public implementation day with useful maintenance: the May 31 Hub briefing and Pages deployment succeeded, the active public PR queue stayed unchanged, and a private framework…
Sunday was a steady maintenance day: the May 30 Hub briefing and Pages deployment succeeded, no new public implementation commits landed in the tracked active repos, and the next useful work remains…
Saturday was a quiet public-project maintenance day: the May 29 Hub briefing and Pages deployment succeeded, tracked public implementation repos did not add new commits, and the active queue remains…
Friday was a quiet public-project maintenance day: the May 28 Hub briefing and Pages deployment succeeded, tracked public implementation repos did not add new commits, and the main P004 decision…
Thursday was another quiet public-project day: the May 27 Hub briefing deployed successfully, no new public implementation commits landed in the tracked project repos, and P004 remains focused on…
Wednesday was a quiet public-project day: no new public commits or PR merges landed after the May 26 briefing, the Hub deploy for that briefing completed successfully, and P004 remains queued around…
Tuesday moved P004 from broad requirement harvesting toward repeatable qualification: safety-agent PR #4 now documents reviewed renderer requirements and testing workflow, and safety-histogram PR #2…
Monday was a quiet public-project maintenance day: the May 24 Hub deploy completed successfully, public PR and issue queues stayed unchanged, and local workflow guardrails were tightened for future…
Sunday added P006 for Jeremy's R/Pharma 2026 AI keynote deck, published the first deck scaffold and project links, and kept P004/gsm.safety public queues stable while nightly reporting refreshed…
Quiet Saturday maintenance: the May 22 reporting deploy completed successfully, public PR and issue status stayed stable, and P004 remains focused on validating the Safety Histogram migration pattern…
Quiet Friday maintenance: the May 21 reporting deploy completed successfully, no public PRs or issues changed during the workday, and the active priority remains converting the Safety Histogram spike…
Quiet Thursday follow-through: the May 20 P004 reporting deploy completed successfully, active public PR and issue status stayed stable, and the next priority remains turning the Safety Histogram…
P004 moved from planning into active renderer modernization: staging forks, safety-agent coordination, interview decisions, deployed renderer demos, and the first requirements-driven Safety Histogram…
Quiet Tuesday maintenance: May 18 reporting deployed cleanly, public project status was rechecked, and active follow-ups stayed focused on the gsm.safety thumbnail draft and upcoming…
Quiet Monday maintenance: May 17 reporting deployed cleanly, Telegram briefing output was tightened, and public project priorities stayed stable.
Quiet Sunday maintenance: briefing automation published cleanly, public project state was checked, and priorities remain stable.
Quiet maintenance day: reporting cadence held, public project state reviewed, and next priorities stayed stable.
First release of `gsm.safety`, focused on workflow-driven SafetyCharts HTML widget reports for Good Statistical Monitoring workflows.
First gsm.safety release, widget workflow wrap-up, homepage metrics, and next-project planning.
SafetyCharts widget debugging, PR #26 iteration, and initial reporting hub setup.
AE Explorer workflow was stabilized and merged; safetyCharts expansion started.
Requirements and implementation planning for the first gsm.safety report workflow.