You said the goals are mostly fine and the work is not well tracked, and that what I was describing were requirements under increased autonomy rather than a new goal. That is right, and this page acts on it. The tracking gap is real and measurable — of 129 things that shipped across the seven repositories in the last fortnight, 59 had no requirement above them — but it largely closed itself on the evening of 15 August, and the rate since then is 13%. What did not catch up is the goal. Its stated scope names five pieces of work, four of which are not what anyone has built for a month; its done-condition describes a pipeline rather than the operating model that replaced it; and its thirty-seven open requirements sit in one flat list with nothing saying what any of them belongs to. This page proposes a new body, five workstreams to group those thirty-seven, a bounded answer for the backlog of history, and — last, because it matters least — a recommendation on the rename.
“i'm fine if some orphans stay orphaned, but fix the ones we can. Let's tag true orphans with a label. and just add comments on requirements that get retroactive updates. Again, we're pushing for transparency and continuous improvement. not perfection.”
That is the opposite of what this page recommended. The recommendation was to draw a start line — call everything shipped before the evening of 15 August history, stop counting it, and never backfill — and the concierge that carried the question to him and the operating officer both said the same. He took neither that nor a full backfill. Attach what can honestly be attached, label the rest as settled rather than pending, and say so on the face of every requirement that gains a child after the fact.
He refined it once, when the chronology problem was put to him: a good many of those items were finished before the requirement they would attach to was ever written, so a forced parent would be a fiction. Hence “the ones we can” rather than all of them, and a label for the rest that reads as a settled state instead of an outstanding task.
Done the same morning, and verified against GitHub rather than against an exit code. Fifty-one items were attached — ten closed issues re-parented, forty-one merged pull requests re-linked with a dated retroactive note — and all fifty-one were re-read afterwards to confirm the structural link actually changed. A hundred and forty-six were labelled as accepted orphans across all seven repositories. Twenty-eight requirements were told, on the issue, that they had gained a child after the fact and why. Twelve further candidates were rejected on reading them, because each only mentioned a requirement in passing and a wrong parent is worse than none. The check now stops counting a labelled orphan, so the number falls to real drift: measured at 06:05 on 17 August, immediately after the work, the count across all recorded history went from a hundred and ninety-seven to one, and that one survivor was eight minutes of genuine drift caught by the check itself. It is not a standing figure and should not be read as one — a day later the check reports nine, which is new work orphaning at the ordinary rate rather than the fix coming undone. Work: #233 and obot.agent#164.
The title becomes Increase Autonomy, as recommended, and the loop framing goes into the goal's opening sentence rather than into the name.
That is this page's recommendation taken whole, including the part that matters more than the word: the public-facing name stays something an outside reader can parse, the false claim about the work living in one repository goes, and the loop gets described at the altitude where it has room to define itself. The option naming schedules is ruled out with it, on the ground that scheduled sessions have never run.
Applying it is his, not an agent's. The goal's own text reserves its direction and membership to him and autonomous sessions propose rather than edit, so the new title is posted to the goal as a comment, together with the body this page proposed, and an agent applies it on his word. Nothing downstream moves when it does: the public URL, the registry entry and the selection lane all read the goal slug, not the title. Comment: #73.
“I'm listening to the podcasts. They're pretty good, actually, so good job. Um, I'm just gonna try dictating directly to you from my phone. That seems easier than a than a Siri link. We'll see how that goes. Um, I listen to the one about the goals, and I approve. Yeah. That all sounds fine. The goal groups makes sense. I do have a question about whether everything... you mentioned there were forty something issues attached. I'm assuming those are all requirements, but I'm not really sure. I guess it doesn't really matter. It... it's approved either way. Um, so I approved both of your recommendations.”
Both recommendations, taken as recommended. The goal's done-condition becomes the operating model rather than the pipeline — a week in which he answers decisions and reviews release candidates and nothing else needs him (G1) — and the thirty-seven open requirements are grouped into the five named workstreams, carried as labels rather than as a new tier of issues (G2). With the orphaned history and the rename already settled on 17 August, all four questions on this page are answered and nothing here is his any more.
That ratifies something that was already live. The five labels were applied at the operating officer's direction, not on an answer of his, and this page said so plainly while the question was open — defensible only because a label is reversible, which was the reason it preferred labels in the first place. It is no longer a proposal wearing a label's clothes. The label descriptions still read “proposed, pending his confirmation”; correcting them is follow-through, not part of this record.
Where his words came from is part of the decision, not colour around it. He listened to the audio episode of this page and dictated his answer straight into chat from his phone — the first decision this program has taken through a brief rather than a page. It did not use the Siri lane, and he said why: “That seems easier than a than a Siri link.” Requirement #265 was written that same morning around the opposite assumption, that a hands-free answer would come through Siri and Reminders, and it should be read against this rather than the other way round. He also quoted nothing — no question code, no issue number — which is the constraint #265 states and this exchange confirms: the words that reached the record were “the one about the goals” and “both of your recommendations”.
What follows from the approval is somebody else's to do, and none of it happened here. The goal's title and body are his to apply — both are posted on #73 as a comment, because the goal reserves its direction to him — and requirement #226 stays open until they are. This record is #271.
He wondered, while approving, whether the forty-odd labelled issues are all requirements, and said it did not change his answer. He is owed the number anyway. Counted on 2026-08-18 from the goal's own sub-issue connection rather than from the figures earlier on this page: forty-five open issues carry one of the five workstream labels, every one of them a child of #73, and none carries two. Forty-three of the forty-five are requirements. Two are not.
The two are #94, recent releases rendering as an em-dash on the roadmap page, filed 25 July; and #152, the four documents that disagree about whether an increment pull request should be a GitHub draft, filed 14 August. Both are small items filed straight against the goal rather than wrapped in a requirement that would only restate them at greater length — the lane #234 proposes and has not yet settled, which these two predate. So the answer is: almost all, and the two exceptions are the shape the labels are meant to sort rather than a defect in them.
One thing the forty-five does not cover, stated because a count that flatters is worse than none. Twelve open children of the goal carry no workstream label at all on the same date. One is deliberate and says so on its own issue: the safety-census review is product work on a clinical package that landed under this goal because the review happened here, and a taxonomy that absorbs everything stops telling him anything. The other eleven were all filed on 18 August, after the labelling pass. That is the goal gaining children faster than a pass over it, which is its ordinary state rather than a hole in the grouping.
I proposed filing a new goal to cover the operating system built over the past week — the dashboard, the operating officer, worker identity, the decision pipeline. You turned it down and said why:
“You own the strategy for the project with me, not the goals. I think the goals are mostly ok — I want to keep them very high level. I do agree with your sentiment that the work isn't well tracked though. I think what you're describing are requirements under increased autonomy.”
The diagnosis was right and the remedy was wrong, and it is worth being precise about which half failed, because the failing half is not the obvious one. The work is not unlinked. Goal #73 has fifty sub-issues and every substantial thing built this week is among them. What is missing is any structure between the goal and its thirty-seven open children, and a goal body that describes them. A new goal would have added a fifth top-level thing to maintain and left both of those untouched.
Four goals, kept very high level: autonomy, charts, the app, open.csr. Nothing here proposes a fifth. Everything below is a change to goal #73's body, to how its requirements are grouped, or to its title.
The goal was written on 24 July and last touched on 14 August. Reading it today against what exists:
The scope lists, in order: hardening the unattended lane from its first run, a second version of idea triage, trigger reliability for that pipeline, a broader grant matrix, and a goal-selection command. Of those five, idea triage and trigger reliability are still open and untouched since July, the grant matrix and the selection command were never built, and only the unattended-lane hardening resembles recent work. Meanwhile the things that did get built — an operating officer that judges delivery, a local dashboard where you answer decisions, permanent worker identities, a decision-answer pipeline, discipline checks across all seven repositories — appear nowhere in the scope.
Done currently looks like this: an idea becomes a requirement, becomes a build, becomes a draft pull request, becomes a digest, unattended, with you reviewing at the gates. That was a fair description in July. It has no place in it for an officer that turns your asks into requirements before they become work and judges whether a finished worker actually moved the roadmap — which is the thing that was actually built, and which you approved as a design on 16 August. A pipeline has stages; this has a role in it. The condition also predates the scheduled-sessions assessment, worker identity, the delivery record and the dashboard.
The title reads increased autonomy in obot.agent. The work now spans seven repositories, a public site and a local dashboard, and the discipline checks were explicitly widened past the hub at your instruction. This is the one part of the title that is factually wrong rather than merely inelegant, and it is a reason to rename that survives independently of which name you prefer.
“The work isn't well tracked” is the half you care about, so it should be a number rather than an impression. Every shipped item across the seven project repositories was checked for the structural parent link GitHub records — not a reference in prose, which does not count and has twice been mistaken for one.
| Repository | Shipped, last 14 days | No requirement above it | Rate |
|---|---|---|---|
| obot.agent | 76 | 34 | 45% |
| safety.viz | 15 | 12 | 80% |
| obot.roadmap | 30 | 6 | 20% |
| gsm.safety | 3 | 2 | 67% |
| open.gismo | 2 | 2 | 100% |
| open.csr | 2 | 2 | 100% |
| safety-histogram | 1 | 1 | 100% |
| Total | 129 | 59 | 46% |
A further 101 orphaned items sit outside the fourteen-day window, giving 160 in total across recorded history as this page counted it. The delivery that followed counted 197, because it swept every repository the policy file lists rather than the seven inventoried here, and swept closed issues as well as merged pull requests. Same population, a wider net. The worst offender in absolute terms is the operating-system repository itself, which is its own kind of finding: the machinery for tracking work was built without tracking the work.
Splitting the same measurement at the evening of 15 August produces a break sharp enough to date:
| Window | Shipped | Untracked | Rate |
|---|---|---|---|
| Before 15 Aug, 20:00 (within the fortnight) | 84 | 53 | 63% |
| After 15 Aug, 20:00 | 45 | 6 | 13% |
The same break shows in the operating-system repository's own issues, and there it is almost perfectly clean. Of the twelve issues created there before that evening, none has a parent requirement. Of the seventeen created after it, sixteen do. The one exception is a note about an unverified claim rather than a piece of work.
The six items that remain untracked since the break are not features. Three are paperwork pull requests that carry no closing keyword — a bytecode cleanup, a status-line fix, the writing convention — and three are hub fixes from the spike-deletion sequence. The check counts a pull request as parented only when it closes an issue, so a small direct fix registers as an orphan by construction.
Filing discipline turned on halfway through 15 August and has held since. The problem to solve is not “start linking work” — that is happening. It is a stale goal, a flat list of thirty-seven requirements under it, and 160 items of history nobody is going to retro-file.
One caveat worth keeping. The discipline check that produces these numbers went live on 16 August, and this is the second day it has run; the break it reports coincides with the period when the people filing issues also knew a check was coming. The measurement is of filing behaviour, not of whether the requirements filed are good ones — that question is the next section.
My prior going into this was that the recent requirements were receipts — filed after the fact to describe work already done. Reading them, that is wrong and worth correcting plainly. They are well-formed: most open by quoting you directly with a date, state a business need in words, and carry a design section. Twenty-one of the twenty-four filed between 14 and 16 August are still open, which is not what a receipt looks like.
The actual defect is altitude. Each one is scoped to a single thing you said, so a day in which you make eight remarks produces eight requirements, all at the same level, all parented directly to the goal. That is how forty-eight percent of the goal's fifty children came to be filed in three days. Nothing is mis-filed; there is simply no layer between a goal you want kept very high level and thirty-seven peers.
Sorting all thirty-seven by what they are actually about produces five groups. They were proposed when this page was written, applied as labels while the question was still open, and approved as the grouping on 18 August — see the decisions section at the top.
Everything you personally look at: the local dashboard, the public hub, the roadmap page, the diary. Includes making them read like news rather than an audit log, holding your queue to three buckets, telling you what changed since you last looked, and rendering honestly on the new machine where none of the local history exists.
Who runs the work when you are not there. The operating officer itself and its second phase, the rule that every worker finishes into a pull request, a question or a config request, in-flight supervision, waking on a worker that stops, the roster of every agent with its cost and impact, and the permanent identity each worker carries.
The mechanics of a run: starting fast, closing fast, surviving a memory ceiling overnight, running alongside other sessions, the single-agent lane that skips startup, the scheduled recurring loop, and not littering the browser. This is the oldest cluster — most of it dates from July and none of it has been touched since.
The checks that keep the roadmap honest: measuring discipline across all seven repositories, making every issue reachable from a goal, gating the audit's design check at the moment work starts, the weekly freshness sweep, and the two idea-intake requirements that feed the plan in the first place.
The release lane and what a piece of work says about itself when it arrives: the release-candidate framework, resolving the four documents that disagree about draft pull requests, reviewer notes, who holds and who reviews an in-flight item, and marking whether a requirement carries your decision or an agent's inference.
The bot avatar (#76) is the weakest fit in the set and is in group five only because it concerns the identity attached to work you review. It would be equally defensible to close it.
I had suggested four groups: surfaces, the officer, provenance and identity, and roadmap discipline. Sorting the actual thirty-seven breaks it. Provenance does not hold together as a group — worker identity belongs with the officer that uses it, and the decision-attribution requirement belongs with the review lane — and the four miss two clusters entirely: the seven requirements about how a session behaves, which is the whole of the standing July work, and the six about how work reaches you. A taxonomy derived from three days does not cover a goal that is a month old.
Five names need somewhere to live. The obvious options are a new tier of grouping issues under the goal, or labels on the requirements.
Grouping issues would give each workstream a body and a done-condition, which is genuinely useful — and it would also be a fifth goal wearing a hat, five times over. You would have five more things whose membership someone maintains and whose text goes stale exactly as #73's did. Labels are cheaper and reversible: the goal stays the single parent, the dashboard and the hub can group by label, nothing new appears in your queue, and the officer can apply them as it files. The cost is that a label carries a name and nothing else, so the explanation of what each workstream is has to live in the goal body — which is where this proposal puts it anyway.
Fifty-nine inside the fortnight, 101 outside it. The check reports the count rather than hiding it, so this will keep appearing in the officer's state file every five minutes until something is decided about it.
The argument in this section is the one he overruled, and it is left standing rather than rewritten so the record shows both what was recommended and what he chose instead. Backfilling is not worth doing. Attaching a requirement to a merged pull request from three weeks ago does not make the plan more accurate — it manufactures a record of intent that never existed, and it is the kind of tidying that consumes a night and changes nothing you would ever read. The alternative is to declare a start line: everything shipped before the evening of 15 August is history, the check's window is set to begin there, and the count of what fell outside is stated once rather than re-reported forever.
The one qualification is safety.viz, at 80% untracked over the fortnight and outside the operating-system work entirely. That is the charts goal's problem rather than this one's, and it is worth knowing about, but it does not change what to do here.
You offered three names. Before weighing them: a goal whose body and structure are current will work under any of them, and one whose body is stale is not fixed by a better title. If only one thing on this page gets decided, it should not be this one.
With that said, renaming costs almost nothing, and that is worth knowing before choosing. Nothing in the tooling keys on the title. The goal's page on the public site is named from a goal-slug: autonomy comment in the issue body, and the policy file that makes the goal selectable keys on the same slug — so the title can change while the public URL, the registry entry and the selection lane all stay exactly as they are.
Fixes the one part of the current title that is untrue, keeps the word that the artifacts, the policy file, the goal slug and every prior decision already use, and asserts nothing that has not been built. Its weakness is that it names a direction rather than a thing: the last week's work was not “more autonomy” in any measurable sense, it was building an operating model, and autonomy is the outcome that model is supposed to produce.
Names two mechanisms. Goals exist and carry real weight. Schedules do not exist — the readiness assessment now in your queue answers “not yet” against five unmet gates, the first of which is a machine that stays awake, and the scheduled trigger has never fired once. Naming a standing goal after a mechanism that has not run means the title makes a claim the program cannot support, and would need changing again if that lane slips. This is the one I would rule out.
The most accurate description of the work. It names what is being built — the loop that runs itself, from your ask through requirement, work, judgement and surface, back to you — rather than a dial being turned. It is also the only one of the three that accommodates all five workstreams above without strain, which is a real test: the surfaces, the officer, the session mechanics, the discipline checks and the review lane are recognisably parts of one loop, and are not recognisably parts of “autonomy”. Its cost is legibility. The hub is public and feeds the conference diary; a reader who arrives at the goals page learns nothing from a coinage, whereas “autonomy” lands immediately.
Take “Increase Autonomy” as the title, and put the loop framing in the goal's opening sentence, where it has room to define itself. That keeps the public-facing word an outside reader can parse, drops the false claim about scope, and still gets the clarity — at the altitude where a coinage can be explained rather than one where it has to stand alone.
Offered so the change is concrete rather than described. This is a proposal for you to apply — the goal reserves direction and membership to you, and autonomous sessions do not edit it.
Build the loop that runs the roadmap without you in it: your ask becomes a requirement, the requirement becomes work, the work is judged when it finishes, and what you need to see reaches a surface you read — with your review gates exactly where the operating contract puts them. Expand how much of that loop runs unattended as trust accrues.
Five workstreams: the surfaces you read; the operating officer and the workers it supervises; how a session starts, survives and ends; whether the plan still matches what shipped; and how finished work reaches you. Boundaries unchanged — merges only through the merge tool with approval, the policy and goal carve-outs stay yours, nothing is deleted without approval, and work found mid-run is filed rather than built.
A week passes in which you answer decisions and review release candidates, and nothing else needs you: every ask you make becomes a requirement without you filing it, every worker that finishes is judged against what it was asked to do, every piece of shipped work has a requirement above it, and the surfaces tell you what changed since you last looked without your having to ask.
The current condition describes work flowing from idea to digest unattended. The proposed one describes a week in which you answer decisions and review release candidates and nothing else needs you. The difference is that the second has the officer in it and the first does not.
The alternative is leaving one flat list, or creating five grouping issues. Labels keep the goal as the single parent and add nothing to your queue; grouping issues would carry more context and also five more things to maintain.
They will otherwise be re-reported every five minutes indefinitely. Backfilling manufactures intent that never existed; bounding declares 15 August the start line and states the historical count once.
“Increase Autonomy”, “Increase Autonomy using goals and schedules”, or something around “Loop engineering”. Renaming is free — nothing keys on the title, and the public URL comes from the slug.
Everything above was measured on 2026-08-17 from live GitHub state and the machine's own records, not carried from a briefing.
obot.agent/tools/navigator/checks.mjs, ORPHAN_QUERY), run against all seven project repositories; parenthood taken from the structural field only..claude/session-hub/navigator-state.md, swept 06:00 on 2026-08-17.obot.agent/goals/registry.json, obot.agent/scripts/policy.json, and the slug parser in obot.roadmap/scripts/lib/collect/goals.mjs.Drafted by Claude Code using Opus 5 for @jwildfire.