Eight per goal, drawn from the migration's unfinished edges, the FDA static-display plan, the open.gismo app-first arc, the autonomy learnings, and the September talk. Each was checked against every existing issue title in all five repos — none of these duplicates something already filed.
Two hard facts shaped the charts list. First, of the nine renderer requirement matrices harvested into obot.agent/docs/requirements/, two have neither a shipped module nor a requirement issue: paneled-outlier-explorer (114 harvested rows) and web-codebook (223 rows — the largest harvest of the nine). Everything else in that set has shipped. Second, the FDA ST&F plan identified fifteen figures with no open-source R implementation, and the Kaplan-Meier family is the biggest of those holes. Neither fact has an issue behind it today.
A correction worth carrying: the workspace notes still say "the other 8 RhoInc renderer forks stay remote until their migration starts." That was true in June. Today seven of the nine harvested renderers have shipped as safety.viz modules, and the remaining backlog is two renderers plus nepExplorer (already covered by hub#35).
Already filed here: nepExplorer #35, QT Phase 2 #37, FDA static charts #9, forest plot #38, value tree #39, recurrent AE #40, hepExplorer follow-ups #88, participant profile v2 #75. The eight below fill the gaps those leave.
Port the last mainstream safetyGraphics renderer — small-multiples outlier detection across every measure at once — into safety.viz with its full done-gate.
Why now: it is the only remaining renderer with a reviewed 114-row requirement matrix sitting unused, and gs#20 already built the R-side report workflow against the legacy version. Finishing it is what lets hub#1's success statement — all the renderers, one library — be said out loud in September.
Decide whether the 223-row web-codebook harvest becomes a safety.viz module, becomes part of the app's data-review surface, or is formally retired — and record the decision where the migration story can cite it.
Why now: it is the largest unresolved item in the migration and the only one where "not done" and "not doing it" are indistinguishable from outside. A codebook is arguably an app feature, not a chart, so this decision also shapes the open.gismo scope. Cheap to decide, expensive to keep ambiguous through a keynote.
A KM curve module with risk tables, censoring marks, and confidence bands, plus its static twin in gsm.safety.
Why now: the FDA ST&F assessment named the KM family as the largest figure group with no credible open-source implementation, and time-to-first-event is the display safety reviewers ask for immediately after incidence. It is also the clearest "we built what the guidance asks for and nobody else has" slide in the deck.
An EAIR display — events per participant-year with confidence intervals, by group and by preferred term — as a first-class renderer rather than a column in a table.
Why now: every existing AE display in the portfolio counts participants; none adjusts for time at risk, which is the first question a reviewer asks about an imbalance in a trial with unequal follow-up. It reuses the ae-explorer data contract, so the marginal cost is low.
A categorical baseline-to-worst shift view — grade in, grade out, counts in the cells — complementing the continuous scatter that shift-plot already provides.
Why now: the continuous shift plot answers "how much did values move"; regulators ask "how many participants crossed a grade boundary", and that is a different display. It is the highest-frequency table in the ST&F lab set and it has no interactive twin.
Add a CM domain track to ae-timelines and the participant profile, so treatment of an event sits on the same study-day axis as the event.
Why now: participant profile v2 (hub#75) is bringing AE domains into the profile rail already; adding CM while that axis work is open is far cheaper than revisiting it later. It is also the single most common question asked of a hepatic signal: what were they taking.
One documented path from any safety.viz module to a publication-quality static image, and from any gsm.safety widget to the same output in an R report.
Why now: the static ST&F work (hub#9) is about to create a second rendering path; deciding now whether static output is a separate engine or an export of the interactive one prevents two portfolios that drift. Submission-ready output is also the answer to "can I actually use this at work", which is the question the talk will get.
Define and test the size a renderer must survive — participants, records, visits — with decimation or aggregation strategies where it does not.
Why now: every demo dataset in the portfolio is small by construction, so nobody knows where the modules break. The first person to point a renderer at a 10,000-participant study will find out in public, and there is a live demo scheduled in September.
Already filed here: the D1/D2/D6 design discussion #34, and nothing else. Two of the eight below do not depend on that decision at all and could be filed tonight.
Complete the GitHub App installation on jwildfire/open.gismo so agent commits, pushes and PRs there carry the bot identity like every other repo.
Why now: hub#79's own boundaries name this as a hard prerequisite — no merge increment may target the repo until it lands — and it is the only blocker on the list that is a five-minute action rather than a decision. It is not filed anywhere.
Enumerate what safetyGraphics does today — charts, mapping, filtering, settings, export, the Shiny surface — into a checkable matrix, and mark what open.gismo v1.0 must match to earn the word "replacement".
Why now: the goal's title promises a replacement and no artifact defines what that means, so v1.0 has no acceptance test. It is independent of D1, it is the natural companion to the renderer requirement matrices that already exist, and it turns a claim into a gate.
The step where a user points their columns at the standard domains — the most-used screen in safetyGraphics — as a first-class part of the app, backed by gsm.mapping.
Why now: without it the app only works on data that is already in the right shape, which is nobody's data. The open.gismo plan lists mapping ambition as an open decision (D3); this makes it a requirement rather than a footnote.
Wire the Widget_* bindings into the app so the safety.viz portfolio appears inside open.gismo rather than only in R sessions and the gallery.
Why now: this is the join that makes the two arcs one product, and it is the demo that carries the talk: an audience idea becomes a chart, and the chart appears in the app. Neither repo has an issue describing the seam.
A public template repo containing a synthetic study, its config, and a working pipeline, so anyone can fork it and have a running instance in minutes.
Why now: decision K3 chose this direction on 2026-07-19 and it was never filed. It is also the most convincing possible answer to the keynote's inevitable question — "how do I try it?" — and it needs to exist before the talk, not after.
Export a self-contained HTML report — charts, flags, run metadata — from any completed study run in the app.
Why now: the report is what leaves the tool and gets shared, which makes it the app's real output. The gsm.kri reporting precedent means the shape is known, and it gives Phase 3 (publish and share) something concrete to aim at.
Turn decision D6 into a requirement: named, restorable snapshots of a study's config plus results, so a run can be reproduced or compared later.
Why now: D6 was decided on 2026-07-19 (snapshots are in v1) and has no issue. Reproducibility is the claim the regulated audience will test hardest, and retrofitting snapshots after the storage layer settles is markedly more expensive.
Take the plan's Phase 1 — error handling, validation messages, run reliability on the local backend — out of the report and into the roadmap as a requirement with sub-issues.
Why now: the app goal currently has nothing selectable in it. Even under D1's unresolved architecture question, local-core hardening is work that survives every possible answer, so filing it costs nothing in optionality and gives the goal a pipeline.
Already filed here: --auto v1 #18, ideas-triage v2 #58, PR-ready marking #70, weekly goal review #87, plus two shipped. Two of the goal's own named futures have no issue; they lead the list.
Point a session at a goal and let it choose the increment: read the goal's sub-issues, apply the boundary prose, propose a pick, and start — the thing --auto does, available on demand and attended.
Why now: hub#73 names it as a future and nothing tracks it. It is also the cheapest way to test whether goal boundaries actually produce good picks, with @jwildfire watching, before the same judgment runs unattended at 2am.
The next tier of standing grants, informed by the maiden run: which repos, which branches, which actions move from ask-first to proceed, with the carve-outs restated.
Why now: the goal lists it as the trust-accrual step and one A-run has now completed end-to-end. Grants are a policy-carve-out change, so this needs @jwildfire's hand on it regardless — filing it makes the conversation scheduled instead of incidental.
A second agent, with no memory of the build, checks the unattended PR against its requirement and the evidence gate, and reports disagreements rather than fixing them.
Why now: the maiden run flagged its own deviation, which is encouraging and also exactly the case where self-report is weakest. Independent verification is the mechanism that lets grants widen without review time growing — the load-bearing piece of every later autonomy step.
Name the ways a run dies — stall, expired login, hook denial, merge conflict, empty selection — and define for each whether the run retries, degrades, or wakes someone.
Why now: the overnight runbook already carries this knowledge as folklore (watch for the 40-minute stall; a login-expired stall needs a manual attach). Encoding it is what turns a supervised overnight lane into an unsupervised one.
A lightweight claim registry so concurrent sessions announce the repos, branches and files they are working in, and a session that would collide picks something else.
Why now: four sessions ran in parallel tonight and stayed out of each other's way because a human briefing said so. That does not scale to a lane that starts itself, and worktrees prevent file collisions but not two agents deciding to fix the same issue.
Refuse to start building on a requirement whose Design or Data sections are empty or whose sub-issues are unfiled — advance the pipeline instead, and say so in the digest.
Why now: the charts goal already encodes this rule in prose ("an anchor requirement still in its lifecycle yields a pipeline-advancement increment"). Making it a check rather than a paragraph is small, and it stops the failure mode where an unattended run builds the wrong thing confidently.
Every unattended run reports what it spent — tokens, wall clock, model mix — beside what it produced, and the numbers accumulate somewhere reviewable.
Why now: model allocation is already a deliberate per-task decision, but there is no feedback loop telling anyone whether the expensive choices paid. It is also the number the keynote audience will most want and the one hardest to reconstruct after the fact.
Each A-run publishes a permanent page: what it selected and why, what it changed, what it verified, what it declined to do, and where the deviations were.
Why now: the renderer work already has a done-gate that says evidence must be on the site; autonomous work has a diary entry, which is narrative rather than evidence. In a regulated audience, "the agent shipped it" is only credible if the agent's own trail is inspectable.
Already filed here: the live audience demo #74. The deck itself (#10) and the diary series (#22) exist but are not linked to the goal — see Part 2. The eight below are the pieces neither of those covers.
The talk's spine as a document: the argument, the beats, and a slide-by-slide inventory marking which slides need an asset, a demo, or a number.
Why now: the goal names it as a candidate child and it is the artifact every other keynote item depends on — you cannot build the metrics pipeline or the demo fallbacks without knowing which slides need them. September is close enough that the arc should be stable while the material is still accruing.
For every live element in the talk: a pinned build, an offline path, a rehearsed fallback, and a rule for when to abandon it and show the recording.
Why now: the talk's centerpiece is agents building software live in front of an audience over conference wifi, against GitHub, with tokens that expire in about an hour. Each of those is a known failure mode with a known workaround; the plan is what makes them non-events.
One real session rendered end to end — the prompt, the decisions, the tool calls, the review gates, the diff, the cost — as a page the audience can open afterwards.
Why now: it is already promised publicly as the Part-7 teaser on the keynote page, the session framework now records everything it needs, and it converts the talk's most hand-wavy claim ("the agent ran the session") into something a skeptic can read line by line.
Generate the talk's counts — issues shipped, PRs merged, releases, renderers, tokens, sessions — from the hub and the repos, refreshable on the morning of the talk.
Why now: those numbers will be quoted, screenshotted, and checked. Hand-counting them once means they are stale by the talk and wrong in the recording; generating them means the slide is true and the method is itself a small demo.
The short link and QR code on the slide, pointed at the Ideas board, with moderation and a rate cap decided in advance.
Why now: hub#74 lists moderation and rate limiting as open design questions on a public repo whose link will be broadcast to a room. The intake lane needs to be tested with more than one person before it is tested with three hundred.
A full dry run against the clock, including the live segments, with the cut list decided while there is still time to cut.
Why now: the goal names it as a candidate child. A talk with unattended agent work inside it has a variance problem no amount of slide polish fixes, and the rehearsal is what converts that variance into a decision about what to drop.
A pass over everything the talk points at — repos, issues, diary, dashboards, demo data — for anything not intended for a public audience, with a documented checklist.
Why now: the talk's whole premise is showing the real working repos rather than a sanitized demo, which is the strongest thing about it and the reason a review is needed. Doing it once, early, is far cheaper than discovering something on stage.
One page published the day of the talk: the deck, the repos, the diary series, the forkable demo study, and a plain "start here" path.
Why now: the interest curve peaks in the hour after the talk and decays fast. The page is mostly assembly of things that will already exist, so its only real cost is deciding in advance that it exists — which is the kind of thing that does not happen at 11pm the night before.
The three new goals from Part 2 arrive with rosters, so they do not need requirement candidates to be viable. These are seeds, not a full spike.
Every candidate above was matched against the titles of all 182 issues in obot.roadmap, obot.agent, safety.viz, gsm.safety and open.gismo, open and closed, on the concepts they name — paneled, codebook, Kaplan-Meier, exposure-adjusted, grade shift, concomitant, static export, performance, mapping, parity, demo study, snapshot, verification, grants, cost, retry, conflict, rehearsal, QR, metrics, fallback. Three near-matches exist and are cited rather than duplicated: gs#20 (the legacy paneled-outlier report workflow, closed), sv#49 (exposure track inside hepExplorer, a different thing from an EAIR display), and hub#9 (FDA static charts, which C3 and C7 extend rather than repeat).