🍊😺 obot
Launched
Commit8d9ba9b

changelog v4.8.1 is current with this build

What changed Β· full log

Reports

AI-generated reports for the obot portfolio, following the gsm.roadmap artifacts pattern: one folder per report, containing a self-contained index.html (plus any assets) and a README.md recording provenance, sources, and assumptions. The site deploy workflow publishes this folder as-is.

The reports below were migrated from the archived obot-claw hub (July 2026). New reports land here under the same contract; see the design doc for requirement #7.

Every artifact says what it is β€” in one line, on the page

The news feed's description line is what decides whether a row is worth opening. It comes from the artifact's own page head, written when the artifact is written:

<title>What the page is called</title>
<meta name="description" content="What this page contains, and why you would open it.">

Put it directly after <title>. One line, 40–260 characters, and it must clear the plain-English bar (@jwildfire, 2026-08-15):

node scripts/check_artifact_descriptions.mjs verifies every artifact, and the deploy runs it before publishing: a page with no description fails the build. If one ever reaches the site it renders as a loud ⚠ NO DESCRIPTION strip rather than a plausible sentence β€” a fallback that reads as intentional is how the old one survived six weeks.

Emphasis is structural, or it is a callout

@jwildfire, 2026-08-16: "I really don't like the randomly bolded sentences in the middle of paragraphs. Call things out in modals if they're super important, but no more random inline bold." He reads these pages fast, often on a phone. Bold applied to a quarter of a page emphasises nothing, and the skim it was meant to help is the thing it destroys.

The full rule, with the sweep that established it, is in decisions/README.md and requirement #198.

Index

Report Date Status
bio.viz and gsm.bio v0.2.0 β€” annotated demo 2026-10-04 Current β€” the review surface for the second releases of the biomarker chart library and its R package, captured with real webR against the live dev sites before release prep: the cross-tabulation with the shared cut rule (Response by CRP at its median, p = 0.305), stratified survival with a draggable cut (CRP at 2.783: hazard ratio 3.523; dragged to 4.34: 2.153), hazard-ratio rows in the biomarker screen, every chart's titles, own footnote, PNG and CSV downloads and saved specification, the two new widgets offline, the static figures gallery, an RTF table, and a batch run of two specifications saved from the browser across all 12 biomarkers. Every number read from the page or file and recomputed in desktop R with base R and survival alone, recorded in the README (hub #359–#362)
safety.viz v1.9.0 β€” annotated demo 2026-10-03 Current β€” the review surface for the safety.viz v1.9.0 release candidate, captured with Playwright and real webR against the live dev site at 3acaf62: the demo app's Biomarkers tab with bio.viz's four charts (17 of 17 charts ready on the pilot study), R started on request with one press (13.26 MB, once) and R's tests after it, re-checked in desktop R from the app's own data files, the single file offline saying why it has no statistics, the chart list's format version 2 on the Domains page, and the kit of 36 parts that becomes public surface with this release (hub #366, #354)
bio.viz and gsm.bio v0.1.0 β€” annotated demo 2026-10-03 Current β€” the review surface for the first releases of the biomarker chart library and its R package: the four charts on their live demos with the numbers to expect (IL-6 at Week 4, 1.235 between the arms; TNF-alpha with IL-10, r 0.6384), the R check page and what starting R in the browser costs, gsm.bio's seven statistics functions, and its four widgets opening offline with R's results stored. Every number read from the live page or from desktop R at capture time, recorded in the README
The safety.viz kit β€” what changed, and what did not 2026-10-03 Current β€” a one-page explainer for the biomarker charts objective: safety.viz gained one export, SafetyViz.kit, a frozen list of 36 parts its own thirteen charts are built from, so bio.viz can build on the same sidebar, filters, listing, profile rail and Chart.js instead of copying them. No chart, schema, default or screenshot changed; the one commitment is that the 36 become public surface when v1.9.0 ships. The 36 parts grouped by purpose, how it was proven, two rough edges, and who uses it next
safety.viz v1.8.0 β€” annotated demo 2026-10-02 Current β€” the review surface for the safety.viz v1.8.0 release candidate (sv#158), captured with Playwright against the live dev site at 001ce91: the demo app that puts every chart on one study and says which ones your data can feed (hub #352, on the standard domain set of hub #325), loading your own files and correcting their column mapping in the browser, including as one HTML file opened from disk, nine requests from the original renderers' trackers, and the Patient Journey Explorer as a prototype whose requirement stays open (hub #349). Its AI narrative layer was taken out of the release after review and is kept on a branch
The read, and the bench behind it β€” the recommended path 2026-08-27 Current β€” the synthesis of the four data-loading directions designed the same night, judged from three seats and reduced to one path: the product reads the folder exactly as the CRO sent it, answers with one document that prices every gap in the charts it turns off rather than the columns it is missing, and refuses to run while an identifier does not join. The two front-runners turn out not to be rivals β€” the intake report is a document, og doctor is a diagnostic, and each fixes the other's stated weakness β€” so ten grafts are carried across with their sources named, one addition (answers replay on the next transfer) is flagged as the graft with the least evidence behind it, and seven things are rejected by name including the wizard's shape, confidence scores, an app database and a second mapping layer. What gsm.mapping already solves is stated so nobody reads this as one β€” renaming, typing, domain assembly, and everything source_col cannot say arriving as RunQuery steps β€” and re-measured here, ApplySpec()'s own Columns not found in source data error is unreachable code: purrr::keep() deletes the offending entries before the guard tests for them, so a three-entry spec over a two-column file returns two columns with no message, warning or condition of any kind. Row D1 of the platform gap analysis is re-read platform by platform β€” the count of five survives and three of the five names do not: tidyCDISC documents no mapping surface at all, Spotfire publishes nothing at this grain, JReview's is administrator-configured, while Medidata, Veeva and Empirica have one the row never named. Five things the survey could not establish are stated as such, including what elluminate and Medidata do when a mapping does not fit β€” the failure behaviour this whole design is about. Two clickable mockups: eight steps from a six-file CRO folder to a running study, and the bench worksheet behind the report. Six decisions R1–R6 close the page, each carrying the answer it argues for. Nothing built, nothing filed, no package source modified
The Mapping Bench β€” every column, every consequence, on one screen 2026-08-27 Current β€” one of four competing directions for the open.gismo data-loading session, committed to a mapping surface for a clinical data manager who knows their data and wants control: all 126 required columns beside the delivered ones, four dispositions per row (Bound / Derived / Constant / Declined), and the price of every decision printed while it is made. The hard case was executed, not reasoned. A liver panel pooled from three laboratories, run against gsm.mapping 1.1.3 and gsm.core 1.2.0: 14,300 rows and 765 participants delivered, 10,076 rows and 548 participants survive the participant inner_join, and 417 reach the eDISH plot β€” 55% of the study, with no warning anywhere. The two defects mask each other exactly, so what arrives at the chart has no NAs, one unit per analyte and 16 Hy's-Law candidates over a denominator nobody was told about. A first attempt aborts with 29 rlang frames and Could not convert string '' to DOUBLE. The file proposer is the argument for asking rather than taking: built and run, it gets seven of fourteen domains right, correctly abstains on one, and is confidently wrong on six β€” it puts the liver-lab file on Raw_DATACHG at 75% name coverage, because nine of fourteen domains want a studyid, a subject id and a date. Declining the seven undelivered domains still leaves 26 of 41 data-driven workflows running, computed from the project's own YAML, and the page names the five that name no raw column at all (srs0001 re-normalises rather than failing). Central engineering claim: source_col plus generated RunQuery steps express everything, so nothing is needed from Gilead-BioStats. Clickable bench, seven decisions D-B1–7, and six weaknesses stated in the order the author would attack them β€” beginning with the fact that this is the first thing in open.gismo that genuinely needs an app. Nothing filed, no package source modified
The intake report β€” read their data and report back 2026-08-27 Current β€” one of four competing directions for the open.gismo data-loading session, committed to a single angle: the product reads whatever folder it is given and answers with one costed document, priced in the charts each gap turns off. The argument in one screen: a six-file CRO delivery generated from demo-301's own data β€” 60,938 rows, a SAS transport, one liver panel split across three lab vendors β€” was copied into input/ as the README instructs, and og_validate() replied with twelve lines of file not found naming none of it. The finding that decides the direction: the flagship hepatic explorer passes every name-based check on that delivery and renders 0 of 4 series, because the chart asks for four literal strings ("Alanine Aminotransferase") and the vendor writes ALT β€” and gsm.safety::Input_HysLaw carries the same defaults, so the Hy's Law metric fails identically and invisibly. A surface that only reconciles column names ships both failures through. Four detectors written and run: 6 of 6 files read including the transport; 15 candidate keys yielding 12 relationships with 4 broken-by-prefix and no false positives; 5 derivable columns unlocking 10 displays, two of them marked as assumptions rather than facts. A resolver chains all 39 metric and module workflows back to the 56 raw columns they consume, so every finding is quoted in displays: 3 render complete, 6 render wrongly, 10 are one answer away, 20 have no data at all β€” and answering everything reaches 19 of 39. Clickable mock where the ledger and the YAML diff move together. Five named weaknesses, two of them serious: the first report is a wall a wizard beats on time-to-first-chart, and the answers have nowhere to go because the gsm.mapping spec accepts only type and source_col. Decisions C1–C5; nothing filed, no package source modified
The mapping surface β€” what gsm.mapping already solves, and what it does not 2026-08-27 Current β€” the step before everything open.gismo does well: how a person with a real data extract gets their own study into a project folder, extending row D1 of the platform gap analysis. The correction is the finding: og_validate() never reads source_col, the one spec key whose purpose is to say my column is called something else β€” 94 correct source_col lines written into a real project produced a byte-identical error report, while gsm.mapping::Ingest() mapped the same files perfectly (14 of 14 columns, 1,000 rows). The readiness check and the engine disagree about the same project, so that screen can never reach green for a delivery not already named with gsm's internal names. The job is a quarter the size it looks: only 23 of the 90 column declarations are named by any metric β€” Raw_SUBJ$subjid carries 23, $invid 12, $country 11, $timeonstudy 9 β€” and the Site Risk Score has no spec at all, so a partial mapping changes its denominator rather than failing. A clickable four-step mockup whose suggestions are computed at page load, run against a CRO delivery hand-authored before the alias table it is matched against (scores Partial ADaM, 19 of 42). Six decisions D-MAP1–6; nothing filed, no package source modified. Written concurrently with, and not merged into, the companion design session
CSR template objects β€” what the report framework still has to be taught 2026-08-26 Current β€” a study produces about a dozen documents and open.csr could build one of them; this records the second, and reads the R Consortium submission pilots for what else could be borrowed. Shipped: the ICH E3 Annex I synopsis as a real template object (open.csr PR #29), twenty-five sections assembled against the same test study from the same analysis results datasets β€” 254 randomised and 217 patients with a treatment-emergent adverse event are the same figures in both documents because both bind the same named value, and the six displays renumber from Table 14.1.1 in the report to Table 13.1 in the synopsis from one unchanged specification. The licence answer is the finding: all twenty-six RConsortium submission repositories were checked against the API, and the ones holding the useful artifacts β€” the ADRG, the define, the eCTD skeletons, pilot 6's ARD-based table program β€” carry no licence file at all, which is all rights reserved; the ones that do carry a licence are GPL-3.0, one-way incompatible with this Apache-2.0 project. Nothing was adapted. The twelve documents a CSR framework plausibly needs, each with what stands between open.csr and building it, and the two that need neither new prose nor a decision β€” the post-text display package and the abbreviated report. A gap analysis against the reference report for the test study puts open.csr at six displays of thirty-two, and names the six as the safety spine. Three things stated as unestablished rather than guessed, including the PHUSE ADRG template's terms
open.csr β€” the data design framework 2026-08-26 Current β€” how open.csr turns study data into a Clinical Study Report, written for a colleague who has never opened the repository: the four source directories a contributor edits, and one real display followed end to end. analysis.yaml says what to compute and display.yaml says how to show it; the pipeline writes an ARD of 236 rows for the demographics table, plus HTML and submission RTF from the same rendered cells. Every statistic has an address, so prose cites t-demographics:sex:p;variable_level=F;group=Total and the build substitutes 56.3% β€” 181 bindings resolve across the demo report, and CI fails any digit in the finished prose that did not come from one. Named values give a reused number a single name with a closed four-operator arithmetic, deliberately closed so the R builder and the JavaScript gate can each evaluate it and agree. The ICH E3 template is two files β€” 119 sections of what a CSR is, and what this report puts in 18 of them β€” with table numbers assigned at build time rather than typed. One diagram of the four parts and their connections. Five findings: not one committed artifact names its commit (all 13 ARDs record a null commit, the ledger an empty string) so the README's "reproducible from its commit" is true as design and unevidenced in the files; named values reach the synopsis but not the report (11 of 15 cited there, including both derived ones, while the report's only citations sit in an unapproved draft); the README and site described three components where the framework has four; four checkable gaps in the interface contracts; and a named-value citation leaving no entry in a block's bindings array, so anything counting citations from the assembled JSON undercounts. Corrected the same evening when a second template object landed 45 seconds before the companion PR merged β€” the superseded finding is kept on the page with the date it changed. Companion in-repo document and site page shipped with it
Legacy renderer trackers β€” what survives the move to safety.viz 2026-08-24 Current β€” every open issue in the twelve retired safetyGraphics / RhoInc renderer trackers, read and sorted for #33: 282 issues, 144 worth migrating, 84 already covered, 54 obsolete, 0 unassessed. Coverage is checked against safety.viz source, its reviewed requirement matrices and its per-module coverage docs rather than module names, and every "already covered" prints the file line, requirement row or test that covers it β€” 232 of the 282 carry such a check, the other 50 are marked judged. Nine quick wins clear 33 of the migration candidates between them, each citing the thing in the new architecture that makes it quick: two modules already normalise a filter start, five modules build their measure list from one identical line, hep-waterfall resets in four lines, hep-explorer already ships the standing caution qt-explorer wants. Two claims were executed rather than reasoned and came out opposite ways β€” columnPlan(1, …) confirms ae-explorer draws no count column at all for a single-arm study (a defect the port inherited from aeexplorer#148), while listVisits() confirms shift-plot already drops null visits. The brief was wrong about one thing and the page says so: the preserved sweep held bodies but not comment text, so 135 comments across 90 issues were re-fetched β€” 21 issues have an empty body and their whole request in the comments. Nothing filed anywhere; no write of any kind to RhoInc or SafetyGraphics
The ranked head β€” the next ten and the bench 2026-08-21 Current β€” the two tiers of the requirement queue as cards, generated rather than written: the ten carrying top10 with their rank and the one line saying why each sits there, and the eleven on the on-deck bench, deliberately unranked and drawn that way. Title, state, milestone and sub-issue progress are derived from GitHub at build time off the same join the five-minute sweep uses, so the page and the sweep cannot disagree about what the ten are. A player replays every re-rank there has been β€” six commits against rank/top10.json, reconstructed from the store's bytes at each one, with cards moving up, entering and leaving, and the commit's own argument beside each frame. The honesty is the design: ranked order goes back to one commit on 2026-08-20 at 12:39 and the masthead says so, the rails are the only thing drawn to scale and they stop where the record stops, the scrubber steps by commit rather than by time, and reversals are shown as reversals β€” #260 was benched at 19:41 and restored at 21:41 the same evening when three agent commits landed under his name. Membership from the label events reaches back to 2026-08-18 21:32 and is drawn as a separate record. Requirement #297
gsm.qtl report_qtl.yaml β€” upstream audit 2026-07-29 Current β€” an upstream audit of gsm.qtl's QTL report module, triggered by running it in demo-301 and by @jwildfire's pushback on a first pass that got the diagnosis wrong ("Not seeing these in production use"). Every claim re-executed from a clean R session against gsm.qtl's own bundled inputs rather than the demo study's: four defects confirmed, one withdrawn. Ranked β€” the compreas fill tests is.na() only while every dataset in the ecosystem uses "", so the QTL0002 discontinuation listing shows 78 rows where 8 is correct while the same page's headline metric says 14 (silent wrong result); an unqualified pull halts the workflow at step 9 under both shipped drivers; the SUBJ/STUDCOMP join drops invid and the render dies; the report is silently written to tempdir(). Underneath all four: the module is run by no test, example or vignette, yet ships byte-identical to studies via gsm.library. The withdrawn claim is kept on the page with the reason it failed. Written as draft issues with copy-pasteable repros, verbatim console output and rendered before/after screenshots β€” nothing filed against Gilead-BioStats; drafts in obot.agent PR #62
Overview & domain dashboards β€” design 2026-07-28 Current β€” the second design pass of the day on goal #79, directed after #134 shipped: what the top-level pages show. Three mockups β€” Study Overview built around what changed since the snapshot you last reviewed (a review-period control scoping the page, a new/persisting/resolved change ledger, flag tiles and the real enrolment curve), RBQM with the riskiest-sites table @jwildfire asked for beside a funnel plot because 110 of 143 sites score exactly zero and the top five includes sites with two and three participants, and Safety as three options with a recommendation. Three components to settle once (change chip, suppression rule, review state) and eight questions, all answered the same evening. Two findings surfaced rather than buried: the site risk score silently divides by 114 instead of 178 because two 32-weight KRIs are all-NA, and snapshot-over-snapshot trends are blocked by og_run stamping every snapshot with today's date. Grounded in ~20 products surveyed and the primary regulatory texts read in full; every number is real demo-301 data, every screen is a mockup
The app β€” design record 2026-07-28 Current β€” the design pass behind goal #79 (the safetyGraphics replacement, open.gismo arc), in two halves: Part 1 puts three shell/navigation directions on the table β€” A Monitor, B Gallery, C Study site β€” recommends C's frame with B's gallery and A's tiles, and works the two surfaces @jwildfire picked (the chart-viewing workflow, and snapshots / the demo study) plus the provenance chip that traces a number back to its snapshot, data version, package pins and pipeline run; Part 2 states the data & config framework under it β€” workr pipelines, config-as-code in a forkable study repo, Project Snapshots as the only data interface, a Domain as a config entry rather than code β€” with a real / aspirational / inconsistent inventory read from source. Sixteen open decisions (D-APP0–7, D-FW1–8). Every screen is a static mockup and the DEMO-301 study is fabricated. Direction approved in session and built as #134 (demo-301 v0); GitHub-role context on #34
open.csr β€” the change request nobody tracks 2026-07-27 Current β€” roadmap research for goal #112, two parts from a four-lane research fan-out: Part 1 deep-dives 18 CSR-automation and review products on the TLF change-request seam (none tracks a request as an object; none re-flags prose when a table's numbers change; new competitor Clymbr Hub is assembling the TLF side; cards 0.8.0 shipped the ARD diff primitive; ARS verified to have no lifecycle model) and Part 2 proposes the display change-request framework β€” CR-as-versioned-artifact, review tier proven by the computed ARD diff (the SAP/shell asymmetry made mechanical), version-vs-fork rules, drafts watermarked in the PR preview, dual sign-off (biostatistician signs the ARD diff, writer signs the binding impact report) β€” with five shippable increments and five open questions (D-CR1–5) at @jwildfire's review gate
open.csr text-block editor β€” live protocol 2026-07-26 Current β€” the editing surface shipped by open.csr#9 (part B of #113, reader β†’ editor), demonstrated by running it: the repository's own gate module (text-core.js), diff writer (editor-core.js), a real 84-row CDISCPILOT01 ARD and two prose blocks are inlined into the page, so both editors are live and a twelve-step protocol drives them β€” type a number by hand and the numeric-fidelity gate fails it in the sentence; bind the same number and it passes. Each step declares what it expects and then checks itself (12/12 against bb58906). Covers the four ways a binding goes wrong, scale/digits qualifiers, the patch (hunks offset past the frontmatter, so no browser edit can touch approval state) and the multi-block patch bar
safety.viz v1.5.0 β€” annotated demo 2026-07-25 Current β€” visual companion to the release plan for #114 item R1: what v1.5.0 adds, feature by feature, as four annotated walkthroughs β€” participant profile and its v2 rail plus adverse-event domain (sv#105, sv#112), migration Sankey + ALT waterfall (sv#97), the eDISH follow-ups (sv#110) and pre-filled axis limits (sv#108). Four short screen captures (draggable Hy's-Law cut-lines, profile click-through, Sankey β†’ composite hand-off, waterfall hover) plus annotated stills, all taken with Playwright against the live dev site β€” the build this release promotes β€” with numbered steps into each demo. The profile captures were re-taken late on 2026-07-25 against the rail, after sv#112 merged and removed the dock
Audit view redesign β€” three ways to clear the queue 2026-07-25 Current β€” design prototypes for #109, driven by the real nightly ledger (33 findings, 22 rules): three working compact views β€” A Ledger (one table, rule bands, detail in place), B Rail (master–detail), C Sweep (rule-first batch). The queue drops from 9.3 screens to 2.0 and from 125 px to 31 px per finding; βœ“/βœ— per row and per rule, collapsible search/sort/filter sidebar, staged decisions that print the repository_dispatch body instead of sending it. Decided the same day: Option B ships, dispatch stays per click, D3–D7 to the recommendation (rule-band reject confirms above 3, activity log as a fold under the table, rule reference kept, run status as a row pill plus one panel); PR #110 unblocked
Dashboard chat β€” working prototype 2026-07-25 Parked 2026-07-26 (backlog) β€” evidence for #77: a file-based per-session inbox delivered by a Stop hook (working sessions) or a persistent Monitor (idle sessions), with the transcript JSONL tailed as the reply stream and a loopback-only local server hosting the live page; both lanes verified end to end (0.9 s idle, next turn boundary while working); decisions D1–D6 awaiting @jwildfire; companion to design #77 and obot.agent PR #50
Platform gap analysis β€” what the other safety platforms ship that we don't 2026-07-25 Current β€” external-landscape survey: 13 safety-monitoring / clinical-review platforms plus 2 reference catalogues, 63 capabilities scored against the portfolio (17 have, 7 filed, 37 missing or partial). Headline: the chart migration is nearly complete and ahead of the field on abnormal-baseline DILI, but 0 of 10 review-workflow capabilities exist here β€” review state, change-since-last-review, annotation, issue tracking, alerting. 12 ranked requirement proposals, deduplicated against the goal atlas (defers to its C3–C8, A3, A6, A7); nothing filed
The Goal Atlas β€” first sweep of the roadmap against the goal layer 2026-07-24 Current β€” first full issue sweep after the goal layer shipped (#53/#71): 182 issues, only 29 of 74 open ones reachable from a goal, 45 unclaimed including 20 requirements; four rollup visualizations, three proposed new goals (workbench / one-product / evidence) with rosters, nine proposed relinks, and 32 candidate requirements across the four goals; all proposals awaiting @jwildfire β€” nothing filed or linked
Participant profile v2 β€” where the profile lives 2026-07-24 Current β€” interactive UX mockup for #75: the profile in a right-hand rail, the four surfacing options switchable live against the real chart + profile modules, expand-to-full-screen, and the AE summary / AE timeline tracks on the lab chart's study-day axis; decisions D1–D9 awaiting @jwildfire; companion to design #75
Participant profile β€” one drill-down module for every safety.viz chart 2026-07-22 Current β€” triggered by collaborator feedback on the composite view's missing eDISH-style click drill-down; four surfacing options (A dock / B drawer / C view / D host-composed) with a recommended hybrid + cohort stepper; decisions D1–D4 awaiting @jwildfire; relates to sv#87/#88/#53/#91 and hub #43
obot portfolio β€” executive overview 2026-07-21 Current β€” accomplishments to date, status across 7 workstreams, next steps (this week + Aug–Sept), gaps G1–G7, and the roadmap.html diagnosis behind generator v1.8.0; feeds hub #31. Direction update same night (banner in report): stage model set β€” Stage 2 = safety.viz fully shipped, Stage 3 = autonomy goals (charts + safetyGraphics-replacement app), Phase 4 = Sept talk prep; decisions on #10/#18/#34
FDA ST&F β€” static display strategy 2026-07-21 Current β€” full inventory of the 60 tables / 22 figures in guide v2.0; 15 figures unserved in open-source R, 12 with a shipped safety.viz twin; architecture + phasing recommended, decisions D1–D4 awaiting @jwildfire; input to hub #9 (P005)
Safety graphics β€” improvement requirements & feasibility 2026-07-17 Current β€” 5 colleague ideas scoped; all 4 sources reviewed; Initiative 01 (hep composite, sv#67) + Initiative 03 (QT Phase 1, hub #36 / sv#68) building; renal ties #35
nepExplorer β†’ safety.viz β€” migration assessment 2026-07-15 Current β€” GO (phased); decisions D1–D3 awaiting @jwildfire; input to hub #29 / #33
Hep-explorer β€” upstream backlog & a clinical-guide section 2026-07-12 Current β€” proposal awaiting @jwildfire (A/B/C decisions); companion to safety.viz PR #44
open.gismo v1.0 β€” design & roadmap 2026-07-12 Current β€” decisions D1–D6 awaiting @jwildfire; companion to open.gismo PR #1
Roadmap usage audit β€” the public story lags reality 2026-07-11 Current β€” tier-1 corrections awaiting @jwildfire
safety.viz homepage β€” five layout directions 2026-07-11 Current β€” awaiting @jwildfire's pick (safety.viz#29)
safety.agent harness proposal (#17/#18) 2026-07-04 Current
Autonomous PM/Development framework report (10 chapters) 2026-06-06 β†’ 06-11 Current β€” flagship; Chapter 10 covers the Claude Code migration
Autonomy audit and refactor development framework 2026-06-05 Current
PM agent and portfolio framework review 2026-06-06 Current
Subagent failure deep dive 2026-06-06 Current
Work-session supervision acceptance evidence 2026-06-06 Current
P009 supervised runner user summary 2026-06-08 Current
Framework options v4 β€” Paperclip evaluation 2026-06-07 Superseded by the framework report
Framework options v3 2026-06-07 Superseded by v4
Framework options v2 2026-06-06 Superseded by v3
Framework options v1 2026-06-06 Superseded by v2

Superseded versions are retained deliberately (design decision D3) β€” the memory philosophy favors preserving the decision trail.

Decision artifacts

decisions/ holds the pages an autonomous session writes when it hits a call it cannot make β€” situation, options with what each costs and forecloses, a plain recommendation, and what unblocks on each choice. One folder per decision, dated. These and release-candidate PRs are the only two things @jwildfire reviews, per the release-candidate framework; the contract is documented in decisions/README.md.

Session reports

sessions/ holds the frozen per-session operational records produced at wrapup by the session hub (requirement #24). It follows its own flat contract β€” one self-contained HTML file per working session, named by the diary slug β€” documented in sessions/README.md.

Adding a report

  1. Create reports/<kebab-name-with-date>/ with a self-contained index.html.
  2. Add a README.md: how it was generated, sources, assumptions, LLM disclaimer.
  3. Add a row to the index above (newest current work at the top of its group).