Jeremy went to bed with a six-item list and woke to eight merged pull requests, but the night's real finding was not in any of them. Nine separate times, in unrelated subsystems, an operation reported success while having no effect: a page that threw away every config detail an agent wrote, a ledger that lost entries under concurrent writes, an audit whose verdict was swallowed by its own output ordering, and his own answered decisions sitting in a queue nobody was watching. Every one returned exit zero β and by morning the verification pass had found three more, one of them in its own tooling. Then the blocker he had set aside the evening to clear turned out never to have existed β a claim asserted by one agent, relayed unverified, and written into a readiness artifact as one of four things standing between him and unattended operation. Two agents falsified it independently; the count is now three. The last thing he said reframes the rest: the roadmap is being updated after the work instead of authorising it, and not one of the night's six task issues is attached to a requirement GitHub can see β the only one that cites a requirement does it in prose, which no tool can follow.
This entry covers one continuous session from the evening of 08-15 through the morning of 08-16 β there is no separate 08-15 entry; the work of that evening is recorded here.
π Every claim below was checked against live GitHub by a second agent before posting: 8 merges, 4 hub commits, 2 release candidates, 5 artifact URLs and 13 issues confirmed, with six corrections folded in β they are named where they land.
π¦ Release candidates needing review
gsm.safety v1.1.0-RC1 β the safety charts become counts: three participant-level metrics, including Hy's Law candidates and QTcF, land as flagged rates a reviewer can rank sites by. RC PR #52 Β· demo Β· ask: merge and tag (carried from 08-15)
open.gismo v0.2.0-RC1 β runs the full gsm pipeline against a plain project folder, with no GitHub, Actions or Pages required. RC PR #10 Β· demo Β· ask: merge and tag (carried from 08-15)
Both were retitled to the naming rule Jeremy set overnight and their bodies rebuilt to the new contract β metadata only, no review state touched.
No RC from the autonomy goal β deliberately: the night's obot.agent work merged straight to main on the operational lane, and the next stable cut is the release.
π§ Decisions needed
Does the roadmap authorise the work, or just record it β and who does the translating? β none of the night's six task issues carry a requirement link GitHub records; the single one that names a requirement does so in prose only. Most began as a direct instruction from Jeremy β though not all, since one came out of an agent's own investigation, which is its own version of the same gap. He then named the org model that answers it: he and prime set strategy, the Navigator translates strategy into actionable tasks and holds workers to delivery. That makes the Navigator the missing layer rather than a project beside it, and it may collapse the separate supervisor into the same role. (artifact in flight β D0017)
Who watches the workers β a first mate to catch stalled and dead agents; recommendation: not three standing sessions, but one standing session plus two scheduled programs and a supervisor that wakes on a detection. decision artifact Β· answer in Q&A Β· blocks: #185
What the Navigator checks when a worker closes β recommendation: read the harness job ledger rather than agents' self-reports, which are missing for nine of forty-one terminal jobs. decision artifact Β· answer in Q&A Β· blocks: #184
Scheduled sessions: go or no-go β now three blocking fixes, not four, and his own cost drops to a single sign-off line. decision artifact Β· answer in Q&A(carried from 08-15, revised overnight)
nep staging design options β safety.viz #126(carried from 08-14)
Closed this session: demo-301 site size and the app elicitation method both carry his words at the top of their artifacts and drop off this list β his queue went 7 β 5.
Work completed
The Operations Dashboard became usable β his verdict on it was "the right shape, but still pretty rough using it".
oa#119 β "your hands" became config with stable c0001 identifiers, the rail compacted from 1681px to 626px on a phone, and one site now carries three tabs with Operations as the default.
oa#124 β config items became installation qualifications: Do / Expect / Verify / Unblocks / Source, enforced at write time, with a Check button that runs read-only proofs and records a real pass or fail. The diagnosis was not that agents wrote thin entries β the page was discarding the detail they wrote.
critical is capped at three and earned, not asserted, and nothing earned it on its first night.
oa#125 β RC naming ({package} vX.Y.Z-RCn), the RC body contract, and a middle panel that renders the pull request natively with its demo page framed live. GitHub cannot be iframed β x-frame-options: deny β but Pages can.
His decisions started reaching their artifacts
oa#123 β answers now supersede rather than accumulate, carry their decision identifier, and move captured β delivered β applied with anything unapplied for an hour showing as OVERDUE. He had clicked three times on one decision and asked twice whether it had landed; it had not.
Both stranded answers were then applied and deploy-verified, and the pending queue is empty.
The config ledger stopped being able to lose things silently
oa#127 β an append-only journal under a lock. The investigation proved the two "missing" identifiers were never real, then found something worse on the way past: twenty-four concurrent captures left twenty, then five, then twenty-two entries, one run with a duplicate. That had presumably always been true.
oa#129 β the audit printed notes before its verdict, and the sweep summarises by first line, so the verdict would have vanished from nearly every run while looking healthy.
The merge guardrail stopped being prose
oa#113 merged after a real conflict resolution. It forces the attested lane when a pull request touches a guardrail path or local policy drifts from main's, verified not to loosen any existing check by a thirty-verdict lane matrix identical before and after. It is the last guardrail change that could merge without his approval β the gate closed behind it.
Advanced: both release candidates retitled and their bodies brought onto the new contract; two decision artifacts marked Decided; one corrected.
Fixed on verification:#184 was reported as linked to its goal and never was; five of the eight new issues were not on the board at all, and three more sat on it with no stage. All corrected during the wrapup β but each was a close-out claiming a link it had not made.
Blockers / risks
An unverified claim set a whole evening's agenda. One agent reported the permission classifier denying a merge command four ways; it was relayed to Jeremy as fact, became his stated focus, and was written into the readiness artifact as a blocking fix. Two agents later falsified all three of its claims. The same root cause was already diagnosed on 08-14 β an invocation-path difference, not a denial β so this is a resolved finding that came back as a new blocker. obot.agent #131
The roadmap is downstream of the work. All six task issues filed overnight carry no requirement link GitHub can resolve; one cites its requirement in prose. (the D0017 artifact is the response)
Two sources of truth for whether a decision is made β the registry's status field is written on exactly 2 of 17 artifacts and read by nothing; the index row drives the log, the dashboard queue and the Todo. A field named exactly what the next reader will grep for, carrying an answer nobody consults. hub #196
A stall that looked healthy β one agent verified its finding and then went idle for three hours without publishing. Caught only because a human asked. (covered by readiness Fix 2)
Verification found three more of the same class β a close-out that reported a goal link it never made, a branch cleanup that reported deleting three branches and left all three on the remote, and, in the verifier's own tooling, a GitHub query using a field this gh build does not support: eight calls returned nothing and the loop would have reported an empty, clean result. The check written to catch the failure mode contained it. obot.agent #131
Agent-authored work is appearing under Jeremy's own account β most of the night's issues and one discussion were created by the jwildfire user rather than obotclaw[bot], against the standing convention. It reads as harmless until you notice it makes the record of who did what untrue, which is precisely what the worker-identifier work is meant to fix. (covered by the worker-identifier requirement)
A dead agent's worktree (obot.roadmap/.claude/worktrees/d0003-land) holds superseded edits β now proven byte-identical to what main already carries, so removing it would lose nothing. Unmerged, so it stays until he says otherwise. (needs approval)
That same dead agent had already written to GitHub before it blocked β three issues, three sub-issue links and a comment, none of them in the session record. It was described as having landed nothing. The writes were complete and consistent; the reporting was wrong. obot.agent #131
Scaffold changes
Complex worker briefs are now written by a fork subagent, not inline β Jeremy's fix after a fifteen-minute wait: "you should be using subagents more aggressively for complex spawns β¦ I need you responsive so we can work on other things." Four briefs composed in one turn was the failure; the routing rule was silent on who writes the brief.
New memory: silent success is the house failure mode β nine instances catalogued, with the standing check: verify the effect, not the exit code; confirm a check is live rather than merely wired; order output so the verdict survives summarisation.
Corrected memory: a GitHub App bot cannot be an assignee at all β a platform limit, not the token problem previously recorded.
Trap recorded:open.gismo's local origin points at Gilead-BioStats; jwildfire is the fork remote.
Next session: loose ends
Answer D0017 β the roadmap-primacy and org-model decision; it changes how every future ask is handled, and the Navigator build waits on it.
Review both release candidates β gsm.safety and open.gismo, unreviewed since 08-15.
Answer the two oversight artifacts (supervision, closeout) β and confirm whether the separate supervisor survives the COO model or folds into the Navigator.
Worker identifiers β W0001 onward, directed at the end of the session. It is what makes by-agent attribution truthful rather than approximate, since every agent-authored change carries the same bot identity today. An agent picked it up as the session closed and the requirement is already waiting: #194, alongside #195, which writes the Navigator's operating-officer role down as a requirement. Both landed after this entry was drafted β and neither was on the board until the wrapup put it there, the third instance tonight of that same omission.
Two issues were filed for the risks below β obot.agent #131 (the false blocker) and hub #196 (the registry's second truth). Each carries a proposed mitigation ending in a recommendation, so both want a yes or a no rather than a fresh analysis.
Decide the stale worktree β d0003-land, proven to hold nothing that isn't already in main. One word removes it.
π ToDo
gsm.safety's v1.1.0 milestone describes the wrong release β its text is a verbatim copy of v1.2.0's, which itself says the work moved off v1.1.0. A rewrite, not a tweak; agent-doable.
Three merged branches still on the remote (blkint, pycache, verdict) β reported deleted, deleted only locally. Safe to remove; left alone because deleting anything needs his word.
Six merged pull requests carry no Closes #X line (obot.agent #101, #105, #108, #110, #128, #129) β against the issueβPR link convention, so their issues did not auto-close.
Two answered decisions have unanswered Q&A threads (#150, #159) β the artifacts carry his words; the threads never got marked.
Otherwise the pending-answer queue is empty, and the retitles he was going to run by hand were applied by an agent instead.