roadmap audit · decision artifact · 2026-08-15

The roadmap audit audits itself

A full audit of jwildfire/obot.roadmap ran tonight under the new standing grant to auto-accept high-confidence findings. 32 changes were applied without asking. But the sharper instruction was the aside: “generally speaking, if it needs my attention it probably isn’t a good audit rule.” Applied as a test to all 22 rules, it fails five of them — and 18 of the 24 findings that landed in your lap came from just two rules. One broken rule was fixed and shipped tonight; the rest are the decisions below.

45
live findings, 22 rules
32
changes auto-applied
24
findings that needed you
2
rules producing 18 of them

Decisions

@jwildfire · 2026-08-15 · in chat

“I reviewed the roadmap artifact. It’s good overall. R4 is the only place we disagree. I do want requirements tied to a single release. If a requirement is too big for that, split it into multiple requirements. It’s fine to defer sub tasks, just make a note on the original requirement, file a new requirement and then transfer the deferred tasks. Other than that note, your recommendations are approved.”

Six of the seven calls are adopted as recommended — the four verified-shipped closes plus the standing grant to close future ones the same way, re-homing the three orphaned pieces of work, retiring the stale checkbox list, parking long-untouched work automatically instead of asking why it stalled, catching a missing design when work enters development rather than after it ships, and taking goals off the delivery board.

The seventh is rejected and replaced. The proposal was to teach the audit that some requirements stay open on purpose because a later phase is coming. He does not want requirements that stay open on purpose. A requirement is tied to exactly one release; if the scope is bigger than one release, it is more than one requirement. Deferring is still allowed — it just has a procedure, in this order: note the deferral on the original (what is deferred and why), file a new requirement for the deferred scope with its own milestone, transfer the deferred sub-tasks to it, and the original then closes with its release. “Phase two is coming” is never a reason to hold a requirement open, because phase two is a different requirement.

Carried out the same morning. The four approved closes are closed. Both known violations of the new rule were split by the procedure above: the kidney-explorer requirement closed with the release that carried its population screen, and its unbuilt patient-profile half moved to a requirement of its own; the liver-explorer follow-up closed with the six enhancements it delivered, and its three unshipped items were transferred to a new requirement rather than re-filed, so their scoping and history survived the move. A sweep then found five more requirements in the same shape — they are in §5. The rule itself was written into the requirement-authoring path, where requirements are created, rather than into the audit that catches them afterwards.

The short version. The roadmap was in better shape than the finding count suggests. Most of the 45 findings were mechanical and are already fixed. The real defect is in the rule set: three rules ask questions instead of proposing changes, and a question can never resolve itself — it re-fires every night until you answer it. Fix those three and the nightly needs-you rate drops from 53% to about 13%.

The seven calls — all now answered

Three were one-off tidy-ups of the roadmap itself; four changed how the nightly audit behaves so it stops handing you the same work every night. Six were adopted as recommended; the last was replaced by your own rule. Kept here as the original proposals, with each outcome marked.

  1. Close four issues that shipped weeks ago and were never closed, and let the audit close future ones on its own when it can prove the same three things. — Recommended: yes to both.C1
  2. Re-home three live pieces of work whose parent requirements were closed out from under them — a keynote stylesheet, a lab-chart rollout across six renderers, and one item that is probably now moot. — Recommended: re-home two, drop the moot one.O1
  3. Retire the three-week-old checkbox list of items still waiting on you and re-file only the four that are genuinely still live, as ordinary issues. — Recommended: retire and re-file.H1
  4. Stop the audit asking “why has this stalled?” — have it quietly move long-untouched in-flight work back to Backlog instead, which you can undo by dragging it forward. — Recommended: adopt the automatic park.R1
  5. Stop the audit nagging for design write-ups on work that is already built — catch the missing design at the moment something enters development, when it can still change the outcome. — Recommended: move the check to that moment.R2
  6. Settle whether goals belong on the delivery board, so the audit stops asking you the same question every night. — Recommended: goals are not board items; take the two that are on it off.R3
  7. Teach the audit that some issues stay open on purpose because a second phase is planned, so it stops proposing to close them. — Rejected. Replaced by your rule: a requirement covers exactly one release, and a planned second phase is a second requirement.R4

Answered in chat on 2026-08-15; the decision as given is recorded in full at the top of this page. The Q&A thread remains as the working discussion, but this page is the record. The short codes are only there so a reply can name one call without quoting it.

0. What was done about all this, the same morning

Everything below was carried out on 2026-08-15 by an unattended session, under the decision above. Nothing was deleted.

CallDone
Close four shipped issues C1All four closed on the evidence already in their threads: the Safety Histogram pilot, the obot GitHub App, the safety.viz docs site, and the gsm.safety branch model. The standing grant to close future ones the same way is now written into the audit’s rule set.
Re-home three stranded pieces of work O1The shared keynote stylesheet moved under the keynote goal. The lab-family dock rollout got a requirement of its own under the charts goal, with its issue transferred into it. The upstream-contribution item was closed as no longer applicable — it would have to land in a repository outside your org, where the no-writes rule means there is nowhere for it to go.
Retire the checkbox list H1Closed, with every one of its eleven unticked boxes accounted for in the closing comment: seven were dead letters, decided by events. The four still live are now ordinary milestoned issues — whether the public roadmap keeps publishing what this program costs, the Goal Atlas accept-or-reject, which of the platform gap analysis’s twelve proposals to file, and confirming the ideas-triage requirement can close.
Park stalled work instead of asking why R1Shipped. The rule that produced thirteen of your twenty-four findings now moves long-untouched in-flight work back to Backlog on its own. Dragging the card forward is how you say it is live again.
Catch a missing design at the start R2Not shippable the same morning, and filed rather than faked: it needs the audit to remember stage transitions, which it cannot do from one night’s snapshot. Now a requirement of its own (#178)a new requirement rather than reopening the closed one that owns these rules, which is your rule applied to the audit that proposed changing it.
Goals off the delivery board R3Shipped, and the twenty-night question is closed. The two goals that were on the board are off it; the convention is written into the hub README beside the goal label; the rule now enforces it per goal instead of debating it. Removing a board item deletes nothing — both goal issues and all their links are untouched.
One requirement, one release R4Your rule, in place of the recommendation. Written into the requirement-authoring path, the task-decomposition guidance, the lifecycle documentation and the audit’s own policy; applied retroactively to all seven requirements already in violation (§5); and the released-and-open rule re-specified around it, with a test pinning the absence of any escape hatch.

0.5  A second pass, later the same day

A follow-up run once all of the above had landed found three things the first pass could not have seen, because they only exist after the changes were made.

Two rules started fighting over the same two cards

Making the stalled-work rule park things automatically is what took it from thirteen findings that all needed you to none. But it also made it executable — and it now collided with the rule that promotes a requirement once every task under it is closed. On the abnormal-baseline liver tools and the open.csr submission outputs, both rules fired the same night and proposed opposite moves: park it back to Backlog, versus promote it to Review or close it. Two automatic rules disagreeing about one card means whichever runs last decides where the card lands — which is the unpredictability the audit exists to remove, and nothing caught it because until this morning the two were never both automatic at once.

Fixed: a requirement whose every task is closed is finished, not stalled, so the promote rule owns it and the park rule stands aside. That is the same “one rule owns one situation” pattern the released-and-open rule already uses. Verified against live state afterwards: no card anywhere carries two automatic proposals. (4bddb57, tests written first.)

An eighth requirement in violation of your rule

The morning sweep found seven. The open.csr submission outputs requirement was an eighth and was missed, for an understandable reason: its delivered scope reached the app’s development tier on the same day the release went out rather than being described in release terms, so it did not read as spanning two releases. It did. RTF tables, the values store and the sidebar all shipped in the editing release on 27 July; the part that decides how editing reaches the reading view shipped in nothing, and nineteen days later had still not moved — for exactly the reason your rule names, that the requirement carrying it already read as done.

Split by the procedure: a deferral note on the original, and the deferred part filed as its own requirement under the open.csr goal with a milestone, carrying its four options and their recommendation across verbatim so the decision loses nothing in the move. No tasks to transfer — the original’s only task covered the shipped parts and is closed. The limit worth remembering is in the sweep, not in the rule: “spans two releases” is easy to see when a requirement says so and easy to miss when it does not, so the audit catching it on the next pass is the safety net.

The audit was flagging its own question queue

The four decisions rescued from the retired checkbox list were filed as ordinary issues — and were promptly reported as defects for being tracked by nothing, which was true and was also the entire point of them. Rather than teach the audit to ignore them, they were put on the board, along with the archive watch item. Hiding them would have recreated precisely the failure that let the checkbox list rot for three weeks: a decision queue that no surface tracks is a decision queue nobody reads.

Where it ended. Findings went from 45 to 5 across the day. Four of the five are automatic proposals to close work that is finished — the liver tools, the open.csr outputs, and the time-to-event chart family — and closing is the one thing agents cannot do here, so they are staged for you. The fifth is a missing design section on the demo study, which the stage-history requirement is the fix for. Nothing is left that needs a judgement from you and has nowhere to go.

1. What was applied without asking

Every change below is reversible and none of it deleted anything. Issue closes were deliberately not auto-applied — they are a state change with consequences, and agent closes are hook-blocked in this workspace by design. They are batched into the first decision below: closing four issues that already shipped C1.

ChangeWhereWhy
milestone backlogFive parked items: reviewer notes in the PR template (#69), the safety-histogram fork close-out (#81), legend and hover on the hep-waterfall panels (#83), the participant-profile copy gap (#119), and the Lex Jansen archive watch (#145)Your 2026-08-14 rule: no work starts until a milestone is assigned. These sat at Backlog with none.
milestone 2026q3Six live items: canonical demo data for the renderers (#25), the open.csr report builder (#111) and its demo surface (#113), the 2026-07-25 review checklist (#114), the four-documents-disagree draft-PR contradiction (#152), and the demo-301 pipeline that has been failing silently (#153)Same rule; these are in Development or actively live, so they get the current delivery slot rather than backlog.
created milestone backlogobot.agent, safety.vizNeither repo had one, so your milestone-before-work rule was unsatisfiable outside the hub. Created rather than reported, because the rule is the point.
milestone backlogThe three orphans: upstream contributions to the shared harness (obot.agent#14), the shared keynote stylesheet (obot.agent#15), and the lab-family dock rollout (safety.viz#99)Real open scope stranded under hub requirements that were closed above them. A milestone makes them pickup-ready; deciding where they should now live is the second decision below O1.
assignee @jwildfireThe Navigator requirement — a standing roadmap-state verifier (#157)Open requirement with no assignee.
board → BacklogThe same Navigator requirement (#157)On the board with no Status at all.
board → ReviewThe abnormal-baseline DILI tools — migration Sankey and ALT waterfall (#43)All three of its sub-issues are closed while the board still said Development.
board → DevelopmentRoadmap transparency — requirements tied to goals, releases, and live status (#31)Sat at Requirement Gathering with its only sub-issue closed and the page long since live. It now carries the open em-dash bug (#94).
board Released → Developmentopen.csr submission outputs — RTF tables and listings, a values store, in-Reader editing (#129)The board was wrong, not the issue. Its own last comment says Parts A and B shipped and Part C still awaits your option pick. Released was a false claim.
gave it a goal: keep adding chartsParticipant-level metrics and flags in gsm.safety (#138) and the safety-histogram fork close-out (#81), both now under the charts goal (#78)The metrics requirement was the one genuinely goalless requirement in the hub; the fork close-out was in no structure at all.
gave it a goal: more autonomyThe four-documents-disagree draft-PR contradiction (#152), now under the autonomy goal (#73)A contradiction in how the harness is supposed to open its own pull requests is an autonomy defect.
gave it a parentLegend and hover on the hep-waterfall panels (#83), now under the hepExplorer follow-up requirement (#88)Interaction work on those panels is exactly what that follow-up requirement covers.
gave it a parentReleases rendering as an em-dash on the roadmap page (#94), now under the roadmap-transparency requirement (#31)It is a straightforward defect in the page that requirement built.
gave it a parentThe demo-301 pipeline failing silently (#153), now under the demo-301 v0 requirement (#134)A broken pipeline belongs under the demo it breaks — and that demo already sits under the build-the-app goal, so the fix inherits a goal for free.
wrote down the shipping evidenceSafety Histogram pilot (#2), obot GitHub App (#3), safety.viz docs site (#21), gsm.safety branch model (gsm.safety#44)Proof that each one shipped is now recorded in its own thread, so the decision to close them is a one-click call rather than a re-investigation.
code fix, merged to mainThe goalless-requirement check now walks the whole ancestor chain (17881b3, closing #133)It used to look only at an issue’s direct parent, so properly-nested requirements were reported as having no goal. See §3: it was producing false alarms, and a false alarm is the purest example of what you described.
board → ReleasedThat same ancestor-chain fix (#133)Closed by the code fix above; the board had not caught up.
board → BacklogThe four issues that were just given parents: hep-waterfall legend/hover (#83), the em-dash release bug (#94), the draft-PR contradiction (#152), and the demo-301 pipeline failure (#153)A side-effect of the links above — see §3.5. Linking something as a sub-issue silently adds it to the board with no Status, so each one had to be staged straight afterwards.

Live findings after these changes: 45 → 28. Of the 28 remaining, most are the decisions below; the rest are covered by the rule changes proposed in §3.

2. The decisions — seven, all in one place

These are the findings that survived auto-application. Each is here because doing it silently would have cost something irreversible, not because the rule was clever.

Close four issues that are provably shipped C1

Four issues are open, sit at Released on the board, have zero open sub-issues, and carry in-thread evidence of shipping written by the sessions that shipped them. This exact set was already written out for you, as an unchecked box, on the review checklist handed to you on 2026-07-25 — the one that was never worked to the bottom (#114). So this has been sitting in your queue for three weeks.

C1-a — recommended

Close all four as completed. Then extend the auto-accept grant to closes that meet a hard bar: board says Released, zero open sub-issues, and shipping evidence already written in-thread. That bar is what made these four safe; it is checkable by the rule rather than by judgement.

Costs
Nothing reversible is lost — reopening is one click.
Forecloses
Nothing. The bar excludes every ambiguous case, including the nepExplorer migration (#35), which is open on purpose for a second phase — see the last decision on this page.
C1-b

Close all four, but keep closes permanently outside the grant. Every future one comes back to you.

Costs
This is exactly the rule shape you objected to: the released-but-still-open check keeps generating queue items for you forever.
C1-c

Leave them open — you are using “open at Released” deliberately as a watch list.

Costs
Then the board stage, not the issue, is the thing that is wrong, and the rule should be re-pointed at that instead.

Re-home three pieces of work stranded under closed parents O1

Each of these is an open piece of work whose parent requirement was closed as completed above it — so the work survives, but it hangs off something finished and is invisible from any goal page. All three now have a backlog milestone so they are pickup-ready. Moving them means unlinking them from a closed parent, which is a structural change I did not make unasked.

The stranded workThe parent that closed above itReading
Stage upstream contributions back to the shared gsm.agent harness (obot.agent#14)Rebuilding the agent as a thin overlay on that harness (#17)Probably moot. gsm.agent is a Gilead-BioStats repo, and the standing rule is no writes outside your own org, so there is nowhere for the contribution to land. Untouched since 2026-07-11.
Land the shared keynote stylesheet (obot.agent#15)The same thin-overlay rebuild (#17)Still wanted. It belongs under the R/Pharma keynote goal (#72). Nothing blocks it except that an issue can only have one parent.
Roll the lab-family dock out across the six remaining long-lab renderers (safety.viz#99)The participant-profile drill-down module (#45)Real, unscheduled work. The participant-profile module itself shipped; rolling its dock across the other six renderers never did.
O1-a — recommended

Move the keynote stylesheet under the keynote goal; close the upstream-contribution item as no-longer-applicable; and file a small follow-up requirement for the six-renderer dock rollout under the keep-adding-charts goal, with the existing issue re-parented to it.

Costs
One new hub requirement.
Unblocks
The closed-parent check goes quiet, and the dock rollout becomes work an autonomous session can pick up on its own.
O1-b

Move all three to their nearest live goal and decide nothing about scope now.

Costs
The upstream-contribution item sits in the backlog forever as something nobody can ever act on.
O1-c

Leave them where they are; a closed parent is acceptable bookkeeping.

Costs
The rule keeps firing every night, and nothing hanging off a closed parent is reachable from a goal page.

Retire the three-week-old checkbox list of items waiting on you H1

On 2026-07-25 a session handed you a single long issue — a morning review-and-release guide — in which every line was a checkbox, and ticking one counted as your approval for the agent to act (#114). Nine boxes were ticked; eleven never were. It is quietly the single largest source of “waiting on @jwildfire” anywhere in the hub — and most of what is left has been overtaken by events. Here is every unticked box, checked against live state tonight.

Seven are dead letters. Each asks you to decide something that has since decided itself:

Four are genuinely still live. These are the ones the recommendation below re-files as ordinary issues, so read them as what you would be agreeing to keep:

H1-a — recommended

Close the checklist and re-file only those four live items as their own issues: the public cost data, the Goal Atlas decisions, the platform gap proposals, and the ideas-triage close-out. A checkbox guide is a session instrument — it goes stale the moment that session ends, and because it has no milestone, no goal and no board stage, nothing else in the system can see the work trapped inside it.

Costs
Four small issues to triage instead of one long checklist.
Unblocks
The Goal Atlas is the biggest single win sitting on the table: six pre-computed re-links you can apply in minutes, plus thirty-two requirement candidates already analysed and written up.
H1-b

Work the checklist to the bottom in one sitting, then close it.

Costs
Roughly half an hour on decisions that are three weeks cold — and the Goal Atlas and the gap analysis both need real thought, which is not what a checkbox is for.
H1-c

Leave it open as a rolling list.

Costs
It has produced no movement in twenty days. It is invisible to the roadmap page, to the board, and to every audit rule but one.

3. The rule set, judged by your own test

You wrote: “generally speaking, if it needs my attention it probably isn’t a good audit rule.” That is a sharper claim than it looks. It says the measure of an audit rule is not whether it detects something true — all 22 of these detect true things — but whether the thing it detects can be resolved by the system that detected it. A rule that ends in a question is a rule that has outsourced its own job.

Here is every rule that fired tonight, scored that way.

RuleFiredAuto‑appliedNeeded youFalse pos.Verdict
STALLED-IN-FLIGHT130130fails — replace
DESIGN-MISSING5050fails — re-gate
GOAL-BOARD-INCONSISTENT1010fails — needs a default
GOALLESS-REQUIREMENT3102fixed tonight
OPEN-IN-RELEASED5131marginal
CLOSED-PARENT-OPEN-SUBS2020marginal — but rare
UNTRACKED-TASK7700passes
MILESTONE-MISSING5500best rule here
SUBS-DONE-PARENT-OPEN2200passes
ASSIGNEE-MISSING1100passes
UNSTAGED-BOARD-ITEM1100passes
Plus 10 rules that fired nothing (CLOSED-NOT-RELEASED, OFF-BOARD-REQUIREMENT, BOARD-DUPLICATE, GOAL-NO-MEMBERS, MERGED-PR-OPEN-TARGET, OPEN-PR-CLOSED-TARGET, PR-NO-REQUIREMENT, REQUIREMENT-LABEL-MISSING, AUTO-DRAFT-CONFLICT, PROMOTED-IDEA-OPEN) and one muted (HARD-WRAPPED-BODY). A quiet rule is a passing rule.

The pattern is not subtle. Every rule that passed proposes a specific change — set this milestone, move this stage, assign this person. Every rule that failed proposes a question — “establish what is blocking it”, “write the Design section”, “settle the convention”. The dividing line is not confidence, difficulty or domain. It is grammar: imperative rules resolve themselves, interrogative rules queue for you.

Stop asking “why has this stalled?” — park it instead R1

Thirteen findings — 29% of the entire audit, and every one of them yours. Its detection is fine (board stage is Development or Review, no activity for 14 days, no open PR references it). Its proposal is “Establish what is actually blocking it and either move the stage back with a reason, or record the blocker.” That is a request for an explanation, and no amount of GitHub data can supply one. So it re-fires every single night, for the same 13 issues, forever.

Two deeper problems. First, stalled is not a defect — you run a wide program and park things deliberately; the rule treats normal behaviour as a fault. Second, it double-counts: the DILI tools requirement (#43) fired here and under the all-sub-issues-closed check, where the very same situation produced a stage move the audit could apply by itself. When two rules see one fact and only one of them proposes a change, the other is noise.

R1-a — recommended: turn the question into a reversible default

Replace it with STAGE-DECAY: in-flight, no activity for 30 days, no open PR → automatically move the item back to Backlog and comment saying so. Mechanical, auto-applied, no question asked.

The insight is that “why is this stalled?” and “should the board still claim this is in flight?” are different questions, and only the second one matters to the roadmap. Parking is reversible and costs you nothing; if the rule is wrong, you drag it forward, and that act is the activity that clears the finding. The roadmap stops claiming 13 things are in flight when nothing has moved in three weeks.

Costs
Board churn on genuinely-slow work. Mitigated by 30 days rather than 14, and by skipping anything with an open PR or an open sub-issue.
Unblocks
13 findings become 0 — some auto-parked, the rest simply not defects.
R1-b — narrow it to contradiction only

Keep a rule but fire it only where the stage contradicts observable evidence: all sub-issues closed while the stage is below Review, a release shipped that closes the issue while the stage is below Released, the last linked PR merged more than N days ago while the stage says Development. Every one of those has a concrete stage move attached. Silence on mere quiet.

Costs
Genuinely stuck work stops being surfaced anywhere. Acceptable if the weekly goal review — the freshness sweep that already proposes linkages and orphan candidates (#87) — is the place for that instead.
R1-c — demote to a counter

Keep detection, drop it from the findings list: render “13 items in flight with no movement” as a dashboard number, not 13 rows awaiting decisions.

Costs
Least change, least value — the number is only actionable if someone reads it, which is the same problem one level up.

Stop demanding design write-ups for work that is already built R2

Five findings, all landing on you, all proposing “Write the Design section from the work as it actually stands.” But every one of the five is already in Development: the remaining-renderers release (#29), the hepExplorer follow-ups (#88), the open.csr change-request loop (#130), the roadmap-hierarchy work (#132), and the demo-301 build (#134). Retro-writing a design document for work that is already built is documentation theatre: it cannot inform a decision that was made weeks ago.

This one is also mis-classified. It is not a cleanup finding at all — it is a work order, and work orders belong in the backlog where they get a milestone and a stage, not in a findings list that re-renders them nightly.

R2-a — recommended

Fire it only on the transition: a requirement moving Design → Development with an empty Design section. At that moment it is a real process violation with a real fix (write the design before building). Once the item is past Development, stop asking — and instead auto-file a small documentation task under the requirement, so the gap is tracked as work rather than as a nightly reproach.

Costs
Needs the audit to remember the previous stage; it currently sees only the present snapshot. That is a genuine addition, roughly a stage-history field on the ledger.
Unblocks
5 findings → 0 tonight, and the rule starts catching the violation when it can still be prevented.
R2-b

Keep it as-is but reclassify it from a finding to a backlog generator: on first sight, auto-file a “write the design for #N” task and never report it again.

Costs
Five new backlog items nobody asked for. Honest, but it moves the pile rather than shrinking it.
R2-c

Retire it. A design document that nobody missed while the work was being built was not load-bearing.

Costs
Loses the one gate that has actually improved requirement quality this year. I would not.

Settle whether goals belong on the delivery board R3

One finding, and it is a genuine inconsistency: the keynote goal (#72) and the autonomy goal (#73) sit on the board, while the charts goal (#78), the build-the-app goal (#79) and the open.csr goal (#112) do not. The rule’s proposal is “Settle the convention” — which is a question, not a change. It has re-fired every night since 2026-07-25, and the same question is also sitting unanswered inside the three-week-old checklist above. Twenty nights of asking you the same thing is the clearest possible illustration of your point.

R3-a — recommended

Adopt goals are not board items and make the rule enforce it: goals live forever, the board tracks delivery stages, and a permanent item has no stage to be in. Concretely, remove #72 and #73 from the board (a removal, so it needs your word) and have the rule flag any goal that appears on it.

Costs
Goals stop showing in board views. They are already surfaced by the goal pages and the hierarchy section, which is the better surface.
R3-b

Adopt the opposite — every goal on the board at a fixed Backlog status — and have the rule auto-add missing ones. Also self-resolving, and needs no deletion.

Costs
Five permanent rows in a board built for things that move.
R3-c

Retire the rule. It has produced exactly one finding in three weeks and no harm has come from the inconsistency.

Costs
None measurable. This is a defensible answer.

R4  Superseded One requirement, one release

What was proposed, and why it was wrong

The proposal was to teach OPEN-IN-RELEASED that some requirements are open on purpose: honour a phased label or a stays-open marker in the body and stay silent. The kidney-explorer requirement was the example — its population screen shipped in a release, its patient-profile half had not been built, and its last comment said in as many words that it would stay open for phase two.

@jwildfire rejected that outright, and the rejection is more interesting than the proposal. The rule was not producing a false positive. It was correctly reporting a defect that the proposal would have taught it to ignore. A requirement that stays open across releases has no release it can be said to have delivered, so nothing about it is ever finished, and every downstream surface — the board, the release notes, the roadmap page — has to carry an item that means “partly done, indefinitely”. Adding a marker would have made the audit quiet about exactly the thing worth being loud about.

R4 — decided by @jwildfire, 2026-08-15

A requirement covers exactly one release. If it is too big for one release, split it into more than one requirement. Deferring sub-tasks stays allowed, with a procedure, in this order:

  1. Note the deferral on the original requirement — what is being deferred and why, in the issue itself, so the next reader does not have to reconstruct it.
  2. File a new requirement for the deferred scope, carrying its own milestone.
  3. Transfer the deferred sub-tasks to it — move them, do not re-file them, so their scoping and comment history moves with them.
  4. The original closes with its release.
Costs
More requirements, and a real authoring step at the moment scope is deferred. That step is the point: it forces the deferred half to be named and milestoned while someone still remembers what it is, instead of surviving as a sentence in a comment on a closed-in-spirit issue.
Unblocks
Every requirement acquires a release it delivered, and therefore an end. The board stops carrying items with no reachable final state, and release notes can be assembled from requirements rather than from memory.
R4-a — original recommendation, rejected

Honour a phased label or a <!-- stays-open: reason --> marker in the body and stay silent; treat recent activity on a Released-and-open issue as a stage error rather than a close candidate.

Why rejected
The first half institutionalises the defect. The second half survives — see the verdict below.
R4-b — not adopted

Add a Partially Released board stage and let the rule read that.

Why not
Same objection, one stage worse: it gives permanent-partial an official name on the board.

Verdict on the rule: sound now, but it needs re-specifying — in the opposite direction

You asked whether OPEN-IN-RELEASED is a good rule under your own test now that the new rule outlaws its blind spot. It is — and it becomes one of the better ones — but it does not get there by being left alone. Three changes, and one of them is the reverse of what R4-a proposed.

  1. Delete the escape hatch instead of adding one. The rule's whole ambiguity was that a Released-and-open issue might be a defect or might be intentional, and it had no way to tell. Your rule removes the second category entirely: intentional no longer exists. So the rule must have no marker, label or comment that silences it — any such hatch would restore exactly the ambiguity that made it marginal. Released-and-open is now always a defect, and the rule's job is to say which of the two defects it is.
  2. Two outcomes, both applicable without asking. If the issue has open sub-issues, unresolved options in its Design section, or recent activity, the board stage is wrong — move it back, do not close. Otherwise the work is done and unclosed — close it. Neither branch needs a human. This is the surviving half of R4-a, and it is what caught the open.csr stage error correctly last night.
  3. Add the split as a third outcome, and keep this one a proposal. A Released requirement with open sub-issues is now a rule violation in its own right, not merely a stage error — it is a requirement spanning more than one release. The audit can detect it precisely and can prepare the whole split: draft the new requirement, list the sub-tasks to transfer, name the release the original delivered. It should not file it unattended, because choosing the new requirement's milestone is a scheduling judgement, and because a wrongly-split requirement is expensive to unpick. So: auto-apply the close and the stage fix, propose the split.

Worth saying plainly: this stops being an audit rule at the moment it matters most. The audit catches a requirement spanning two releases weeks after the split became necessary, when the deferred half's rationale has already faded. The rule belongs where requirements are written, not where they are inspected — which is why it went into the requirement-authoring and task-decomposition guidance rather than only into the rule set. The audit is the backstop for requirements written before the rule existed, and there are seven of those.

3.5  What happened when the fixes were applied

Applying 27 changes produced 6 brand-new findings — a useful accident, because it exposed two rule interactions that only show up when the audit is actually acted on rather than read.

The general lesson: a rule set is only testable by applying it. These three interactions had been latent since the audit shipped on 2026-07-24 because findings were being read, not executed in bulk. Worth re-running the audit immediately after any apply batch and treating a rule that reliably creates work for another rule as a defect in the first one.

4. What this does to the numbers

Tonight: 45 findings, 32 changes applied unattended (27 from the audit, 5 cleaning up after it — §3.5), 24 findings needing you — 53%. Live findings closed the night at 28.

With six recommendations adopted and the seventh replaced by your rule, the same roadmap state produces roughly 6 findings needing you, about 13%. The six that remain are genuine one-off scope judgements, like deciding where a stranded piece of work should now live, rather than the same question re-asked every night. Your replacement of the seventh does better than the forecast rather than worse: the rejected version would have silenced a class of finding, while the rule you gave instead eliminates the condition that produced it — and the sweep in §5 has already cleared every existing instance, so that rule starts tomorrow night from zero.

The one structural addition worth naming: three of the five failing rules fail for the same missing capability — the audit has no memory of stage transitions. It sees only tonight’s snapshot, so it can detect “this is in Development with no design” but not “this entered Development without one”. Adding a stage-history field to the ledger is what turns DESIGN-MISSING, a sharpened STALLED-IN-FLIGHT and part of OPEN-IN-RELEASED from interrogative into imperative. That is the single highest-leverage change to the audit, and it belongs inside the existing nightly-roadmap-audit requirement that owns these rules (#92) rather than in a new one.

5. The retroactive sweep — seven requirements, all closed

Your rule is not only forward-looking: any requirement already open across more than one release is in violation of it today. So the roadmap was swept for that shape — an open requirement whose delivered scope has shipped in a release while other scope remains — and every case found was split by the procedure and closed the same morning. Seven in total: the two known ones, and five the sweep turned up.

The pattern in all seven is the same and it is worth naming, because it is what the rule is actually for. None of these was neglected work. Each delivered its headline scope, shipped it, and then kept a small unshipped tail alive indefinitely — and the tail did not move, in one case for a month past its own release, precisely because the requirement carrying it already read as substantially done. An open requirement is the worst possible home for deferred scope: it is finished enough that nobody re-reads it, and open enough that nothing prompts anyone to.

RequirementWhat it delivered, and whereWhat was deferred, and where it went
Kidney explorer
nepExplorer migration (#35)
The KDIGO creatinine scatter and stage summary — the population screen, in full — in safety.viz v1.6.0, one day before this audit.The per-participant kidney profile, never built and never decomposed, gated on a dataset that does not exist yet. Now #168. No sub-tasks to transfer.
Liver explorer follow-ups
hepExplorer enhancements (#88)
Six of eight enhancements — draggable cut-lines, marginal box plots, dropped-record downloads in v1.5.0; study-day animation, sparklines, quadrant polish in v1.6.0. Two releases, which is the violation.Three items that made neither release: the exposure track and P_ALT estimate, the v1.2 interaction follow-ups, and the waterfall panel legends. Transferred to #169. Two false milestones corrected on the way.
Autonomous operations
(#18)
Autonomous session mode — the --auto skills, goal layer and launcher — in obot.agent v0.3.0 on 2026-07-26. Unattended sessions have run on it since 2026-07-24.Scheduled and recurring sessions. No new requirement was needed — that scope was already filed as a requirement in its own right and was merely nested here; it was re-homed to the autonomy goal.
Session hub
(#24)
The live dashboard, the wrapup report step and the scaffold channel in obot.agent v0.2.0 on 2026-07-15. Every session since has produced a report through it.Rendering those reports as diary tabs — parked at backlog for a month because the live view covered the in-session need. Now #170.
Canonical demo data
(#25)
The migration onto pharmaverseadam in safety.viz v1.1.0 on 2026-07-12 — the half that made the demos defensible to a clinical reviewer.The scripted pipeline that would build the bundle reproducibly. Still a hand-trimmed extract, and every renderer needing a cohort the real data cannot show has bolted another injection script onto it. Now #171.
safety.viz v1.3
(#29)
The renderers it was named for: the hierarchical adverse-event table in v1.3.0 on 2026-07-16, evidence pages carrying reviewed requirement text in v1.3.1.The clearest case of the seven. Four platform-hardening items, none of them renderers, none moved in the month since the release the requirement is named after — three still milestoned v1.3.0. Now #172.
Roadmap transparency
(#31)
The roadmap page itself — requirements tied to goals, releases and live status — in obot.agent v0.2.0. Live and in daily use since.Nothing, strictly. Its one open item is a defect in the delivered page, not held-back scope, so no new requirement was filed: a bug does not keep its requirement open. Re-homed to the autonomy goal as an ordinary issue.

Two of the seven needed no new requirement at all, and that distinction is worth carrying into the rule: the split is for deferred scope, not for every open thread. A defect found after release is an ordinary issue against shipped work. Scope that already has its own requirement just needs re-homing. Filing a new requirement for either would be paperwork, not tracking.

Net effect on the roadmap: eight requirements closed (the seven here plus the fourth verified-shipped close in the batch you approved), five new requirements filed, all carrying milestones, and eleven sub-tasks either transferred or re-homed with their history intact. Nothing was deleted. Every closed requirement now names the release it delivered, and there is no longer a single requirement in the hub that is open with its board stage at Released.

6. Already surfaced elsewhere — not re-litigated here