The autonomous prototype is shut down. From here, work runs as requirement sessions: five objectives broken into requirements and tasks, each with a definition of done, one requirement per cloud session, running for as long as it takes, steered by Jeremy through the issues on this hub and from his phone. This page lays out the five objectives as issue trees, sequences their requirements into two concurrent lanes, and says what the deployed site must show each Friday until the R/Pharma talk in mid-October.
Decisions
Jeremy · 2026-09-10 · handover document and in-session answers
The fully autonomous multi-agent prototype is shut down. Work moves to requirement sessions: semi-autonomous, long-running Claude Code sessions that each pursue one clearly defined objective and take as long as they need. Not a rigid weekly sprint.
The scaffolding is retired in one pull request, tagged as the next obot.agent release; the previous tag and git history are the archive.
The desktop app is safety.viz-native: a self-contained local HTML and JavaScript app. The open.gismo app arc parks for the talk.
Sessions move to Claude Code cloud sessions, verbatim: "I also want to move everything to cloud sessions instead of running from remote control here."
Sequencing goes to this page before any session starts.
On tracking, verbatim: "I want to move to much more rigid issue tracking in roadmap moving forward. Each objective issue should have well defined requirement issues with linked task issues before the session starts. Each of these issues should have clear definitions of done. Objective is complete when all requirements and their tasks are completed per the issue definitions. Each agent/subagent should be linked to an issue or issues and associated PRs. Agents should actively comment on their assigned objectives, and standup should give a summary of which objectives are complete/in progress/blocked. All questions in standup should be tied to blocked issues."
Deadline: everything ready by mid-October, in time for the talk, and ready to go live.
Later still, verbatim: "I think that our 5 goals are too big for single sessions using /goal. I think we should probably have one session per requirement instead. Requirements should be scoped in a way that it can be clearly judged. Let's call our goals 'objectives' instead to avoid confusion. Objectives → Requirements → Tasks. Requirements and Tasks should have definitions of done that can be clearly evaluated, used by the /goal command."
Earlier the same day, on the four gating questions: the merge lane moves to GitHub rulesets with auto-merge on integration branches and a required review from Jeremy on release branches, retiring the policy script; the connected GitHub account is the actor for cloud sessions, with the drafted-by line as the record of authorship; the existing chart and app objective issues are reused and three objectives are filed new; and parity means a static twin for every interactive chart — no code export from the app.
The situation in three sentences
The last iteration built a robust agent structure and then spent it on itself: orchestration health improved while no agent-produced release ever reached Jeremy for review. Meanwhile the product is real and mid-flight: thirteen interactive safety charts are released, nine have R widgets, two more widget releases sit in review, and the FDA static-chart work has an approved design and zero implementation. Five weeks remain to bring SafetyViz to parity with the old safetyGraphics app — its charts, its data loader and mapping, and a tool that runs on a desktop — plus the modern charts it never had.
How work runs now
The issue tree is the contract
An objective is a hub issue that can complete. It lists its requirement issues and states what done means for the whole objective. It is complete when every requirement is closed.
A requirement is a hub issue in the existing five-section shape, plus a definition-of-done block: the observable end state, how it is proven, and the release it ships in. It is complete when every task is closed and the proof is posted.
A task is an issue in exactly one implementation repo, linked as a sub-issue of its requirement, with its own definition of done and a link back to the parent. A pull request closes it with the closing keyword and the proof in the PR body.
Nothing starts without the tree. The session-start check is: every node of the objective's tree exists, has a definition of done, and carries the talk milestone. If a node is missing, the session's first job is to file it and stop for Jeremy's sign-off before building.
The session
One requirement per session; an objective is the steering unit and is never a session. The session runs Claude Code's built-in /goal in auto mode, with the condition built from the requirement's definition of done: every task under it closed with its evidence, or blocked with a comment naming what is needed from Jeremy, and the requirement's own end state proven in a closing comment. The condition names how the session proves it each turn, and a turn or time bound.
The evaluator judges only what the session surfaces, so every turn ends by listing the tasks' state. Background work defers evaluation and triggers check-ins at 30 minutes, then hourly.
Fan-out uses standard tooling: subagents for bounded work, the Workflow tool (ultracode) for multi-agent stages, Claude Design (ultradesign) for visual design. Every subagent brief names its task issue; every PR cites its task. No bespoke dispatch, no worker ledgers.
Merges use the existing lane and policy file unchanged. Release-candidate pull requests remain Jeremy's review, and they are the weekly deliverables.
Where sessions run
Every requirement session is a cloud session on claude.ai/code, bound to its objective's repository — Lane A in safety.viz, Lane B in gsm.safety — in a per-repository cloud environment whose setup script installs the toolchain once and caches it. Nothing runs on the Mac, nothing is bridged through Remote Control, and the machine can sleep.
Steering is native: Jeremy opens the session on the web or the phone and answers or redirects there. A session waiting on him idles until its VM is reclaimed, which is why blocked questions also live on the issue — the answer survives the session.
The nightly standup is a scheduled cloud routine on the hub repository: it reads GitHub, renders the objectives as complete / in progress / blocked with one question per blocked issue, and publishes to the voice-readable branch. Routines run autonomously, hourly at the finest, and count against the account's daily run allowance.
Cross-repository writes — a task closed in gsm.safety, a comment on the hub objective — go through the GitHub access the cloud session already holds. The obotclaw[bot] identity does not travel: its private key sits in this Mac's Keychain, so in the cloud the actor is the connected GitHub account and the attribution line in the body says who drafted it.
Week 0 verifies two things before anything depends on them: that the built-in goal command runs inside a cloud session, and that R plus the pharmaverse packages install in the gsm.safety environment's setup script within its cache.
Comments and the standup
At session start: one comment on the requirement with the order of its tasks and which agent holds which; the objective's sign-off comment is quoted.
At each task close: the definition-of-done evidence on the task, and the closing PR.
Nightly: one progress comment on the requirement in three fixed headings — complete, in progress, blocked; when a requirement closes, one line on the objective. A blocked item is an issue carrying the blocked label whose latest comment asks Jeremy one specific question.
The standup is rendered from GitHub, never from local state: every objective with its complete / in progress / blocked counts, then a questions section with one entry per blocked issue and nothing else. It publishes to the same voice-readable branch as today. Jeremy steers by answering on the issue, or by messaging the running session.
What goes and what stays
Retired
The navigator, admiral and prime sessions and their launchers; the autonomous dispatcher and objective registry; the session bookend skills and their scratchpad, journals and dashboards; spend tooling; the stop hooks that interrupt turns; sibling-spawn briefings; the Remote Control and launchd lanes; the Keychain-bound bot push helpers.
Kept
The chart-building skills and test framework; the requirement lifecycle skills; the release-notes skill; the standup, rebuilt as a cloud routine that reads GitHub; the news feed, which shows issue transitions and artifacts in place of the diary. The merge lane becomes GitHub rulesets: checks and auto-merge on integration branches, Jeremy's review on release branches.
The five objectives as issue trees
Each card is the proposed tree. Each requirement is one session — scoped so a single /goal run can close its tasks and prove its definition of done; where a requirement below is too big for that, it splits at filing. Requirements marked as existing fold in issues already on the hub rather than duplicating them. Weeks refer to the calendar below; the repo is where the tasks live.
Objective 1 · Chart coverage: every figure in the FDA safety guidance
hub objective: reuse "Keep adding charts" · lane B · repos: gsm.safety, safety.viz, obot.agent
Done: all 22 figures of the FDA Standard Safety Tables and Figures guide (v2.0, April 2025) render as static charts in gsm.safety with visual-regression evidence and a gallery entry; all 13 safety.viz modules have an R widget; both packages are released.
Reference data and derivationsexisting: "static safety charts from FDA guidance", phase 0 · week 1
Derivation functions for treatment-emergent flags, abnormality grading, extreme-value exclusion, ULN multiples, DILI quadrants — tested against the guide's published criteria.
Requirement-matrix rows for every figure in scope; ADaM-to-mapped-domain alignment written down.
DILI quadrant scatter (F7, F8) first, as the vertical slice that proves the shared derivation layer.
Shift scatter (F15, F22); box plot over time (F10, F16–F20); dot plus risk-difference forest (F2, F3).
Each figure beside its safety.viz twin in a static gallery on the gsm.safety site; package check clean.
Static phase 1b and the Kaplan–Meier familynew · week 4
Retention and rate figures (F5, F21, F12) and the mean-change-over-time pair (F6, F9) as simple engines.
Kaplan–Meier figures (F1, F4, F11, F13, F14) by wrapping ggsurvfit, not by building a survival engine.
Done when the coverage table reads 22 of 22.
Widget parityexisting: two release candidates in review · week 1
The v1.2.0 and v1.3.0 release candidates already carry the missing widgets; Jeremy's review is the remaining step, then merge forward the same day so the stacked branches do not diverge.
The parity guard in CI stays green for every later safety.viz release.
Objective 2 · Interactive and static parity with reproducible code
hub objective: new · lane B, after objective 1's phase 1 · repos: safety.viz, gsm.safety
Done: every interactive chart has a static twin in gsm.safety driven by the same derived data and the same settings names, so a chart configured interactively renders identically as a static figure with no translation step.
One settings contract per chartnew · week 3
Each chart's settings are one JSON schema, used by the JavaScript module and accepted by the matching R function under the same names.
A round-trip test: settings exported from the interactive chart drive the static function without translation.
A static twin for every interactive chartnew · weeks 3–4
Twelve twins arrive with objective 1's four engines; the remainder (histogram, AE timelines, QT central tendency, kidney explorer, hepatic waterfall, participant profile as a listing) close the set of thirteen.
Proof: for each chart, the settings JSON exported from the interactive demo drives the static function and the two renderings agree on the numbers plotted. The last twins (waterfall, kidney, profile listing) are the cut line if the calendar slips.
Objective 3 · Portfolio view
hub objective: retitle "Build the app" · lane A, first · repo: safety.viz
Done: one page on the safety.viz site shows all thirteen charts organised by data domain on the canonical demo data, with chart navigation, shared filters, a shared participant drill-down, and a chart-status panel saying which charts the loaded data supports; released.
The standard domain setnew · week 1
A portfolio manifest: the domains (subject-level, adverse events, labs and vitals as one long-format measurement domain, ECG), the charts each domain feeds, and the columns each chart needs, validated against the existing per-chart data schemas.
Demo data is the vendored pharmaverse ADaM extract; the scripted pipeline requirement already on the hub becomes its reproducibility task.
The portfolio shellnew · weeks 1–2
One page hosting every chart through the shared chrome, chart navigation by domain, the participant profile as the shared drill-down, and a chart-status panel that mirrors the old app's chart-availability check.
Evidence: a Playwright run over every chart on the page, published with the release.
Study-level settings and filtersnew · week 2
Arm, site and population filters applied across charts; per-chart settings persisted in one study-configuration JSON that saves and loads round-trip.
Objective 4 · Data loading and mapping
hub objective: new · lane A, second · repo: safety.viz
Done: a user drops their own files into the portfolio page, the tool recognises the domains and standard, lets them correct the column mapping, and loads the consolidated portfolio with the chart-status panel reflecting what their data supports — with nothing leaving the browser.
Local file loading and standard detectionnew · week 3
CSV and JSON through the browser's file API; no upload, no server. Domain and standard detection from column names, in the shape of the old app's standard detection, with plain-language errors when a file cannot be placed.
The mapping modulenew · week 3
Per-domain column mapping, auto-filled from the detected standard, editable by hand, validated field by field, saved into the study configuration.
Proof: a deliberately non-standard dataset is mapped and every chart renders.
Fit and loadnew · week 3
Mapped data flows into the portfolio; chart status recomputes from the mapping; the participant drill-down follows the mapped identifiers.
Objective 5 · A local-only desktop tool
hub objective: new · lane A, third · repo: safety.viz
Done: a single downloadable file combines the loader and the portfolio, opens from the desktop with no network and no install, saves and reloads a study configuration, exports charts, and is released with a user guide and a rehearsed talk demo.
The single-file buildnew · week 4
A build target that inlines every script, style and optional demo dataset into one HTML file that opens from a file URL in current Chrome, Safari and Edge, under a stated size budget, with a fresh-machine test.
Persistence and exportnew · week 4
Study configuration save and load; chart export as an image, and the settings JSON that drives the static twin.
Release, guide and demonew · week 5
A release with the download on the site, a user guide page, and the talk's demo script run end to end on a clean machine.
The calendar
Two concurrent sessions, so there are two threads to steer and never more. Each lane runs one requirement session at a time. Lane A carries the app spine — objectives 3, 4 and 5 in sequence, since they compose into one product. Lane B carries the charts — objective 1, then objective 2. The last column is what must be visible on the deployed site by that Friday; it is what Jeremy reviews, and it is the honest measure of the week.
Lane A · the app
Portfolio view → data loading and mapping → local desktop tool. One requirement session at a time, each started when the previous requirement closes.
Lane B · the charts
FDA figure coverage → interactive/static parity. Starts on the reference data while the widget release candidates are reviewed.
Week
Lane A · the app
Lane B · the charts
Must show by Friday
0 · Sep 10–13
Issue trees for objectives 3–5 filed and signed
Issue trees for objectives 1–2 filed and signed; widget RCs reviewed
The retire PR merged; cloud environments for safety.viz, gsm.safety and the hub verified; every objective's tree on the hub with definitions of done; the first two requirement sessions running in the cloud
1 · Sep 14–20
Domain manifest; portfolio shell renders all 13 charts on demo data
Reference data and derivation functions merged; widget releases published
Portfolio page on the dev site; gsm.safety at widget parity; derivations in the package
2 · Sep 21–27
Shared filters and study configuration; objective 3 closes and released; objective 4's first requirement starts
DILI quadrant and shift engines: four static figures with evidence
Portfolio released on the public site; first four static figures in the gsm.safety gallery beside their twins
3 · Sep 28–Oct 4
File loading, standard detection, mapping module, fit and load; objective 4 closes
Box-over-time and dot-forest engines: twelve figures; settings contract
Drop your own CSVs into the dev site and see the portfolio; static coverage 12 of 22
4 · Oct 5–11
Single-file build; persistence and export; objective 5's first requirement starts
Phase 1b and Kaplan–Meier wrap: 22 of 22; remaining static twins
A downloadable app that opens offline; coverage table reads 22 of 22
5 · Oct 12–16
Release, user guide, demo rehearsal
Releases; slack for slips
Everything on the public site; the demo script run on a clean machine
Risks and the cut line
Lane B is the heavier lane: twenty-two static figures and thirteen twins in five weeks. The order protects the talk — the differentiated twelve first, Kaplan–Meier by wrapping, and the last three twins as the declared cut line.
Lane B cannot start gsm.safety work until the two stacked widget release candidates are reviewed; they must merge forward the same day or the branches diverge silently.
The shared derivation layer moves calculations out of shipped JavaScript modules. It is gated behind the DILI vertical slice and the existing Playwright evidence so a working chart cannot regress unnoticed.
Two objectives in the tree already carry the Experimental marking on the time-to-event chart pending external clinical review; the plan does not depend on lifting it.
The goal evaluator reads only the transcript. A session that stops reporting the tree's state will be judged on silence; the session skill makes the end-of-turn state listing mandatory.
Lane B needs R in the cloud sandbox. If the setup script cannot install R and the pharmaverse packages within the environment's limits, Lane B falls back to a routine that runs the R checks in GitHub Actions while the session edits, or to a single local session for that lane only.
The goal command is documented for headless, desktop and Remote Control use and referenced for cloud sessions; if it does not arm on the web, the fallback is the headless form inside the cloud session, which runs the loop to completion in one invocation.
Decisions needed
Decided: objective issues — reuse "Keep adding charts" as objective 1 and retitle "Build the app" as objective 3, each rewritten with a definition of done; file objectives 2, 4 and 5 new; close the autonomy objective as retired; leave the CSR and keynote objectives paused.
Decided: parity means a static twin for every interactive chart; no code export from the app.
Concurrency. Recommendation: two sessions, as laid out. A third lane is possible but doubles the steering load on the busiest weeks.
The freeze date. This page assumes ready by Friday 16 October; the exact talk date pins the last week.
Version plan for the talk. Recommendation: safety.viz v2.0.0 carrying the app, gsm.safety v2.0.0 carrying the static charts — one headline release each rather than a run of minors.
The three open release candidates (gsm.safety v1.2.0 and v1.3.0, open.csr v0.4.0). The first two are Lane B's prerequisite this week.
Decided: the connected GitHub account is the actor in the cloud; the drafted-by line records authorship; the Keychain-bound token and push helpers retire.
Decided: GitHub rulesets replace the policy script — required checks and auto-merge on integration branches, a required review from Jeremy on release branches — so a cloud session merges with a plain pull request.
Cloud environments. Recommendation: one per repository the lanes touch (safety.viz, gsm.safety, obot.roadmap), trusted network plus the package registries the R toolchain needs, and any API key as a proxy-held credential rather than an environment variable.
This week
Stop the surviving navigator session; close the two obot.agent pull requests that only serve the retired machinery; push the eleven worktrees holding unpushed work and remove all sixty-eight.
Connect GitHub for cloud sessions and create the three cloud environments; run a first throwaway cloud session in safety.viz to confirm the goal command arms and the toolchain installs.
Open the retire pull request on obot.agent with the rewritten overlay contract, and a small requirement-session skill: the tree check, the condition template, the comment cadence, and the standup routine's prompt.
File the issue trees for objectives 1–5 with definitions of done and the talk milestone, for Jeremy's sign-off.
Start Lane A on objective 3's first requirement and Lane B on objective 1's, in the cloud.