One page for the whole role, because you asked for one. The Navigator's job is improving operations — the templates, the dashboards, the audit framework, the process mechanics everyone else works inside. Turning your asks into requirements and checking that workers delivered are two instruments of that job, not the job itself. This page covers what the session is, how it starts, what it owns, what it decides without asking, what it escalates, and the one line it must never cross. It also settles something measurable: six pieces of work shipped overnight with nothing in the roadmap recording them — and almost every one of them was operations improvement, which is to say the Navigator's own work, dispatched ad hoc because the role did not exist. It folds in the two pages from last night that each answered part of this, so there is one place to decide rather than three.
“I'm good with D0017 recommendations. Implement. Let me know when the agent is active. One thing for the queue: I want the opsdb/sessions page refactored to show a list of all agents along with thier ID, status, cost and the impact they had on the roadmap.”
All eight calls adopted as recommended. The day-one scope stands and the boundary moves with it — the Navigator authors and repairs the plan, and never touches the work (N1). Structure is the goal set and the shape of the plan itself; everything inside that shape is the Navigator's (N2). The requirement floor is consequence rather than size, with every exemption recorded in one line and the rate reported (N3). Escalation inherits the critical bar the dashboard already enforces, with one clause added: an escalation must name the specific thing only he can do (N4). Every judgment gets a line in the delivery record, calls that change the plan get permanent identifiers, and he reviews them in batch once a day (N5). A standing session he can talk to that thinks only on a trigger, with no polling loop (N6). The separate supervisor folds in, and its requirement closes into the Navigator's rather than being built beside it (N7). Delivery is judged at closeout against GitHub rather than against a job's own account of itself; last night's five linkable issues get linked, the sixth cites the decision behind it, and the rule is forward-only from there (N8).
One sequencing change came with the approval, recommended by the concierge and accepted along with the rest: the disagreement between the nightly audit and the wrapup verifier is resolved before the four new audit checks ship. Two independent checks reported different board state within hours of each other on this date — the audit named four requirements off the board, the verifier found nine more — and neither is assumed correct. Adding rules on top of a check whose reliability is unverified builds on sand, so the Navigator's first job is to find out which one was wrong and fix it.
The ask in his last sentence is not one of N1–N8. It becomes a requirement of its own, written by the Navigator as the first exercise of the translation job this page describes: refactor the Operations Dashboard sessions page to list every agent with its identifier, status, cost, and the effect it had on the roadmap.
Implemented the same morning, in the order the sequencing change required. The audit-versus-verifier disagreement was resolved first, and the answer was that the two checks never disagreed: the nightly audit had not run, and its twenty-two-hour-old total — four findings, none of them about the board — was relayed as that morning’s board state. The verifier queried GitHub live and was right. The Navigator session then went live as job b510658b, and the roadmap-discipline checks shipped across all seven project repos, gated on work done rather than on filing, under the Navigator’s own requirement #200. One correction to the page above: the four checks are three, because an answer unapplied past an hour was already live and is now covered by a test rather than rebuilt. The queued sessions-page ask became requirement #199, judged a new requirement rather than an amendment because the Operations Dashboard requirement declares itself one release and names later surfaces as separate requirements. First live run of the new checks: fifty-six pieces of shipped work with no requirement above them inside a fortnight, and a hundred and eleven older — the Navigator’s backlog to triage, not his to read.
"I can go either way on the manager, but strongly suspect we need an agent for the navigator. My main concern right now is that you and the workers you create aren't impacting the roadmap, which means none of this is sustainable. Every time I ask for something the roadmap should change in a concrete way. Every decision should result in new or updated issues. Every worker should advance a requirement through its lifecycle. That's the overall approach for this project and it currently feels like an after thought." — @jwildfire, 2026-08-16
"The overall approach is that prime and I set the strategy (think ceo) and then the navigator translates it to actionable tasks (coo) and makes sure the workers are delivering." — @jwildfire, the same morning
The complaint is measurable, and it was measured before this page was written. Six pieces of work were filed and shipped overnight in the agent's own repository, all closed by morning, all carrying a release milestone. None is attached to a requirement — the link GitHub actually records is empty for every one of them, including the one that names its requirement in a sentence of prose. And the requirements were there the whole time: four of the six were Operations Dashboard work, and that requirement gained no child issue, no comment and no edit across the entire night. A fifth was the release-candidate naming rule, which belongs to the release-scaffolding requirement — a requirement that names that exact framework inside its own scope.
The night was not roadmap-free, and the difference is the whole diagnosis. Eight requirements were filed overnight: three new ones from the design agents, and five as follow-through on decisions you had just made. Every task issue carried a milestone. So the roadmap moved, repeatedly — but only in the lane where a decision page had already forced it. In the lane where you simply ask for something, it did not move at all. The fault is direction of flow: the roadmap records work afterwards instead of authorising it beforehand.
The cause is structural, not careless. Today an ask reaches the concierge, the concierge writes instructions for a worker, the worker files an issue describing the work it is already doing, and the worker ships. That issue is a receipt. Nothing in the chain opens a requirement, so nothing updates one — and the role doing the dispatching is the one built to answer you in seconds, which is the opposite of what writing a requirement needs. The intake problem is a role problem, not only an ordering problem. Your org chart names the person who should own it.
Look at what shipped overnight, not just at what failed to be recorded. Three passes on the Operations Dashboard. Config items rewritten as installation qualifications. A release-candidate naming rule and a pull-request body contract. A pipeline so an answer you click actually reaches the page. A ledger so the config list cannot lose an entry quietly. And this morning, a scorecard of where the nightly audit is blind.
Every one of those is operations improvement — which is to say, every one of them is the Navigator's own job, dispatched ad hoc by the concierge because the role did not exist to hold it. The six untracked issues were not a discipline lapse. They were a backlog belonging to a role nobody had created yet, executed by whoever was nearest. That is a far better reason to stand it up today than any argument from principle, and it also tells you what its first week looks like — because that backlog is not finished.
All six are unattached, not four and not five. Two earlier counts were low: the remote-control change was missed entirely, and the dashboard iteration that mentions its requirement in prose was counted as linked when GitHub records no link for it either. Five of the six belong to requirements that are open right now; the sixth traces to a decision you signed that day.
And one of the six was not your instruction: the config-ledger fix was filed by an agent that noticed two allocated identifiers with nothing behind them. The other five trace to things you said.
"I suspect that prime (and I) own goals and overall roadmap structure. Coo owns everything else. Requirements, tasks, lifecycle, milestones, etc. coo can propose structural changes, but we make the call on implementing or not." — @jwildfire, 2026-08-16
"I'm more interested in COO improving operations. Updating templates, tweaking dashboards, refining audit framework. That is its job. Splitting goals is strategy that ceo owns." — @jwildfire, 2026-08-16
Read together, those two draw the line the rest of this design hangs on. An ambiguous boundary here is how the Navigator ends up either paralysed or overreaching — and the second quotation rules out one direction firmly: the Navigator does not propose strategy. Splitting a goal, retiring one, deciding what the program is for — none of that is its business to raise. What it owns instead is bigger than bookkeeping: the machinery everyone works inside.
| You and the concierge | The Navigator | |
|---|---|---|
| Owns | Strategy: which goals exist, what the program is for, and the shape of the plan — what stages a requirement passes through, what counts as a requirement, what a milestone means across the portfolio | Operations: the templates, the dashboards, the audit framework, the process mechanics — plus requirements, tasks, lifecycle movement and milestone assignment |
| Decides | Whether the shape or the goals change | Everything inside that shape, without asking — including improving the machinery itself |
| Raises | — | Nothing structural, as a matter of course. Only the narrow case where its own operational work is genuinely blocked on a strategy call — and that is an escalation, not a proposal |
Concretely on the plan: creating a requirement, scoping it, splitting it, moving it to design, attaching a milestone, filing its tasks, closing it — all the Navigator's, no permission needed. Adding a goal, retiring one, changing what the lifecycle stages are, or changing the rule that a requirement covers exactly one release — all yours.
And concretely on the machinery, which is the larger half: sharpening an issue template, adding a check to the nightly audit, retiring a rule that turned out to be noise, reshaping a dashboard section, tightening the release-candidate contract, fixing a process that keeps producing the same mistake. The Navigator does that work continuously and ships it, the same way any other worker ships — a requirement above it, a task beneath it, a release it lands in. It is a builder, not an inspector. Enforcement is one of its instruments; improving the machine so enforcement is needed less often is the actual job.
The concierge should not be filing requirements or task issues at all. Last night's six untracked issues were not a discipline failure — they were the concierge doing the Navigator's job without the Navigator's tooling, judgment or time budget. Move the job to the role that owns it and the failure has no mechanism left to occur through.
An earlier draft of this design gave the Navigator a channel for proposing structural changes. That is cut. Strategy is not something it raises, and building a pipeline for it would invite exactly the drift the ownership line exists to prevent: an officer that spends its attention lobbying about goals instead of improving the machinery.
What survives is one narrow case: when the Navigator's own operational work is genuinely blocked on a strategy call — it cannot finish a thing you asked for without knowing whether a goal covers it — it escalates, in a sentence, and carries on with everything else. It does not stall, and it does not pre-build toward the answer it would prefer. An officer that quietly makes the alternative expensive has taken the decision by taking the option away.
"I said I want a feel for how we work together. Me, you, coo. I want ceo and coo to work very closely. If it proposes a change I want your recommendation on if we should implement it and how. Does the proposed change make sense? The coo reports directly to you. You report to me. I might talk to it directly sometimes if there are specific operational things I want done, but you are still my primary point of contact." — @jwildfire, 2026-08-16
So: the Navigator reports to the concierge, and the concierge reports to you. The concierge stays your primary point of contact, and you may talk to the Navigator directly whenever you want something operational done. Three consequences have to be designed rather than assumed.
When the Navigator proposes an operational change, you do not want it forwarded — you want a recommendation attached: does this make sense, should we do it, and how. That is a review with judgment in it, and a review is only cheap if the thing being reviewed arrives in a reviewable shape. So a proposal from the Navigator must carry five things, and a proposal missing them has failed its format regardless of how good the idea is:
A proposal that forces the concierge to reconstruct the reasoning has not saved anyone anything; it has moved the work.
You should be able to give the Navigator an operational job straight, without routing it through the concierge — that is what working closely means. The risk is two records that disagree, and the concierge answering questions about a state it does not have.
The rule that prevents it is small: a direct task is handled exactly like any other — requirement first, delivery line at the end — and the Navigator tells the concierge in one line that it happened. One record, one queue, no divergence, and the concierge is never surprised by work it did not dispatch. It costs the Navigator a sentence.
| Reaches you | Stops at the concierge | |
|---|---|---|
| Critical items | Yes — immediately, under the existing capped bar | — |
| Proposals needing a decision | Yes — with the concierge's recommendation attached | — |
| The Navigator's own judgment calls | Only if you ask | Yes — reviewed daily by the concierge |
| Confirmed closeouts and shipped improvements | In the daily summary | Reviewed in detail there |
That answers a question left open earlier in this page: the primary reader of the Navigator's decision record is the concierge, not you. Which means it should be written dense and complete — everything needed to audit a day's judgment in one pass — rather than trimmed for a phone. You see the critical items and the proposals; the concierge reads the rest and is answerable for having read it.
The failure mode to watch is the obvious one: a chain of three where the middle link filters everything becomes slower than the two-link version it replaced. The protection is that the direct channel to the Navigator stays genuinely open, and that the concierge's duty is to evaluate, not to gate — nothing waits on its review to happen except the proposals that were going to need your decision anyway.
A standing Claude session named 🧭🤖 obot-navigator, launched from the workspace root the same way the concierge is, and reachable the same ways — from the terminal, from the web, or from your phone.
# the launcher to be written today, mirroring scripts/obot-prime obot.agent/scripts/obot-navigator # which is exactly: cd /Users/jwildfire/Documents/obot2 claude --bg --permission-mode auto --remote-control \ -n "🧭🤖 obot-navigator" --model opus /s-navigator
Three properties matter and each is copied deliberately from the concierge's launcher, which already works:
Its behaviour lives in a new skill, in the same three-piece shape every other role already uses. All three are required — the middle one is the piece that is easy to forget, and without it the skill simply does not load:
obot.agent/skills/navigator/SKILL.md # the contract .claude/skills/navigator -> ../../obot.agent/skills/navigator # workspace overlay symlink obot.agent/commands/s-navigator.md # the /s-navigator entry point
The Navigator must be able to start cold and know the state of the world in one read, because it will be restarted often in the first week. It reads, in this order:
.claude/session-hub/navigator-state.md — what the mechanical sweep saw in the last five minutes: open release candidates, recorded decision answers, the config-ledger verdict, recent events. This file already exists and is already current..claude/session-hub/prime-state.md — the concierge's durable state. Read-only, always. The Navigator never writes it..claude/session-notes/{today}.md — the shared scratchpad every agent logs to.The first three are already assembled into a single bundle by the concierge's rehydrate tool, which the Navigator can reuse unchanged, read-only.
The Navigator already exists as a machine, and it is running right now. A script fires every five minutes under the system scheduler, checks seven repositories, writes a state file, and appends its events to the shared scratchpad. It has run 233 times, most recently minutes before this page was written, and its last exit was clean. It caught a stale claim on 2026-08-14 and it now carries the config-ledger integrity check as well.
That machine is not being replaced — it becomes the Navigator's eyes, and the new session becomes its judgment. Nothing about the split is philosophical; it follows from what each half can actually decide.
| Question | Who answers it | Why |
|---|---|---|
| Which release candidates are open and waiting on you? | The five-minute script | A list. No judgment involved. |
| Has a worker gone quiet? | The five-minute script | Measured: the typical gap between a worker's recorded actions is about twenty seconds, and a thirty-minute-silence rule fires roughly once in a day and a half. A script does this for nothing. |
| Which task issues have no parent requirement, across all seven project repos? | The five-minute script (to add) | The exact check that would have named all six of last night's by morning. The sweep already visits all seven repositories for release candidates, so it is one more field on a walk it already makes. |
| Does this ask need a new requirement, or a change to an open one? | The session | Judgment. Requires reading the open requirements and deciding scope fit. |
| Did this worker attach to the right requirement, and did the stage really move? | The session | Judgment. A script can see that something was touched; it cannot see whether it was the right thing. |
| Should this deferral have become a requirement of its own? | The session | Judgment, and the one most often skipped today. |
The sweep is and remains the sole writer of its state file; the session writes its own, .claude/session-hub/delivery.md, append-only, one line per closeout. That is not fussiness. The last two nights produced repeated cases of one process quietly overwriting another's file, and a shared file here would put the delivery record — the thing this whole role exists to produce — squarely in that class.
The nice part is that the two files rejoin themselves without any new plumbing. The sweep reads the delivery file and renders it as a ## Delivery heading inside its own state file; and the dashboard tab that shows you that state file already renders any heading it finds — it was deliberately built that way, so a new section appears without a line of code changing. So the day-one delivery record reaches your dashboard through machinery that already exists:
node obot.agent/tools/ops-dashboard/ops-dashboard.mjs --serve --open
# the /navigator tab, already live, renders whatever headings the state file carries
There is already a nightly roadmap audit: twenty-two rules, run at half past three every morning, reading the hub's issues and board, the pull requests across the seven portfolio repositories, the ideas queue and the design documents on disk. It is good machinery and the Navigator should lean on it rather than grow a second one. So it is worth being exact about what it did with last night's six real failures: it caught one and missed five.
| The failure | Audit | Why |
|---|---|---|
| Requirements filed overnight that were unparented or off the board | Partly caught | Hub issues are what it reads — but it reported four, and a separate check later found nine more |
| Six task issues shipped with no requirement above them | Missed | Wrong scope — its untracked-work rule only examines hub issues. Now settled as a defect to fix, not a choice |
| Decisions you answered that nobody applied | Missed | It has no view of decision pages at all |
| A false claim left standing in a published page | Missed | No rule reads the prose of an artifact — and none could; this needs judgment |
| The registry and the published index disagreeing | Missed | The site reads one of the two and nothing compares them |
| Workers that ended having produced nothing | Missed | It sees GitHub, not agent runs |
An earlier draft of this page said the hub side was clean and the gap sat entirely in the spoke repositories. That is wrong and it has been removed. The requirements filed overnight were not clean either: one was never actually linked to its goal, and nine needed to be put on the board after the fact by a check running hours later. They are all correct now — verified this morning, every one parented and boarded — but they were corrected, not filed correctly.
So the discipline gap is not confined to the spoke repositories. It is wherever work outruns the record, which last night was everywhere. The intake path is still the biggest hole; it is not the only one.
The reason it is blind is worth sitting with, because it explains why a new rule alone would not have fixed this. A spoke-repo issue is visible to the audit only if a hub issue already lists it as a child — which is the very link that was missing. The check cannot see the failure, because the failure is the absence of the thing that would make it visible. Something has to look from the other end, which is precisely what widening the scope does.
@jwildfire, 2026-08-16: "The audit def needs to include all project repos not just roadmap. Coo watches tasks too and tasks live across the project." That is decided, not offered as an option below.
It is also the change that makes the headline failure detectable at all. The six untracked issues were invisible for exactly one reason — they were in a spoke repository, and the audit only looked at the hub. Without extending scope the Navigator would be blind to the precise problem that created it.
Scope becomes the seven project repositories the status page already tracks: safety.viz, gsm.safety, obot.roadmap, obot.agent, open.csr, open.gismo and demo-301.
What it costs, stated rather than waved past: the audit's snapshot has to read issues across seven repositories instead of one, so each run makes several times the API calls it makes today. That is affordable, but it is not free — if it pushes a run past a comfortable window, the honest fix is a slower cadence for the cross-repo pass rather than pretending the cost is zero. The five-minute sweep is a separate, much cheaper thing and is unaffected.
Widening the scope without widening the rule's precision produces a mess on the first morning. The blunt version — "every spoke issue needs a hub parent" — fires on 26 of the 41 open spoke issues today, and on 14 of the 16 in the agent's own repository. Nobody would trust that list; it would be muted inside a week, and the real signal with it.
Gate on work done, not on filing. An issue needs an ancestor once it has produced a merged pull request or been closed — not the moment somebody opens it. That fires on all six of last night's and on almost nothing else, which is the difference between a rule that survives and one that becomes noise.
And the backlog is the Navigator's, not yours. Whatever the wider scope turns up on its first pass — some of those 26, plus whatever else seven repositories are hiding — it works through and triages itself under the judgment grant. You see only what clears the critical bar. That is the whole point of having granted it: a scope change that would have been unusable when you were the reader is straightforwardly useful when it is not.
Two independent checks disagreed about the same state of GitHub within hours of each other. The nightly audit reported four requirements off the board; a separate verifier, run later the same morning, found nine more that the audit had not named. It may be a stale snapshot, it may be different scope, it may be a real defect — and it is not established which of the two was wrong. Neither is assumed correct here.
It matters because this design leans on the audit heavily. A check whose results are not reproducible cannot be the thing an officer relies on, and finding out which one was wrong is close to an ideal first task for the role: it is operations improvement, it is mechanical, and the answer either restores confidence in the audit or names a defect worth fixing before anything is built on top of it.
The split between the two mechanisms follows straight from this list. Mechanical, and belonging in the sweep: a spoke issue closed or a pull request merged with no requirement above it; an answer you recorded still unapplied after an hour; the registry and the index disagreeing; an agent exiting having produced nothing. Judgment, and belonging to the session: whether a stale claim in a published page is actually wrong, and whether a given piece of work should have had a requirement at all.
One more consequence of the delegation grant lands here. Several checks worth having were never adopted because their false-positive rate was too high to put in front of you. With the Navigator reading them instead, that calculation changes — the audit can afford to be more sensitive than it could when you were the only reader, because a judge now sits between the rule and your attention.
This is the one that runs when nothing else is happening, and it is the reason the role is worth a session rather than a script. The Navigator carries a live backlog of operational improvements and ships them the way any worker ships: a requirement above the work, a task beneath it, a release it lands in. Sharpening a template, adding or retiring an audit rule, reshaping a dashboard section, tightening a contract that keeps producing the same mistake.
It does not need to be told to do this, and that is the point — the other two jobs are reactive, and a role with only reactive jobs sits idle between your asks. The backlog is not hypothetical either: this morning's audit scorecard names four checks that should exist and do not, and last night produced six more improvements of exactly this kind without anyone holding the role.
You say something to the concierge that implies work. The concierge answers you immediately — that conversation is strategy and it goes through nobody — and then hands the ask, verbatim, to the Navigator with one line of context. The concierge does not write worker instructions any more.
The Navigator then, in order: reads the open requirements under the relevant goal; decides whether this is a new requirement or a named change to an existing one; writes it, with a milestone, linked to its goal; and only then writes the worker's instructions, derived from the requirement rather than from the conversation. It reports one line back to the concierge — the requirement's link and the worker's name — which the concierge relays to you.
The rule that makes this safe is that nothing in it is on your response path. If the Navigator is slow, you do not notice; only the worker's start moves, by a few minutes.
The Navigator learns a worker finished from the harness's own job records — the per-job state file and event timeline the machine already keeps for every agent, at no cost and with no calls to GitHub. Three details are load-bearing, all measured rather than assumed:
One thing the ledger cannot be trusted for: the list of what a job produced. Nearly half of the jobs measured recorded no children at all, including one that merged three pull requests and filed two issues. The Navigator therefore judges against GitHub, not against the job's own account of itself. As of this morning the ledger holds one stalled corpse that nobody has noticed — still sitting there, never marked terminal.
On each closeout it asks the three questions in the table above and appends one line to the delivery file: the worker, what it produced, the requirement it moved, and a verdict — confirmed or drift.
When the roadmap did not move, the Navigator fixes the roadmap — and never the work. That distinction is the whole safety story. It may attach the issue to the requirement it belonged to, amend the requirement, or file the missing one, because those are the plan and the plan is its job. It may not touch a line of what the worker produced, merge anything, or publish anything. Every repair it makes to the plan is logged as a call you can read and reverse, so "it tidied up" never becomes "nobody knows what happened" — the failure this role exists to surface must not be absorbed by the role itself.
One section, appended through the day, answering the question you actually asked: what did this day of agents do to the roadmap. Confirmed movements, drift, and anything that closed producing none of the three outcomes a worker is allowed to finish into.
"I am fine delegating a lot of judgement calls to COO. It can escalate critical items to me." — @jwildfire, 2026-08-16
That sentence sets the default, and the default is the opposite of how the system behaves today. The Navigator decides; it does not queue. Almost everything the sweep and the closeout check surface should end with the Navigator resolving it — attaching the issue, amending the requirement, granting or refusing an exemption, judging that a deferral deserves a requirement of its own. Only a critical item reaches you. The present failure is precisely that everything funnels into your queue, and a role that merely adds a better-organised queue has not helped.
"Critical" already has a bar — it should not get a second one. The Operations Dashboard shipped one last night: a pin capped at three items, earned only by a blocking reference confirmed still open on GitHub or by a computed condition, never by an agent's opinion. On its first night nothing earned it. That is exactly the right threshold to inherit as the escalation line, and inheriting it means there is one definition of urgent rather than two that drift apart. You have said you will be annoyed by anything marked critical that is not, and that if something needs your attention it probably is not a good rule in the first place.
It changes what is worth detecting, in your favour. A check too noisy to put in front of you can be perfectly good when the Navigator is the one reading it, because the Navigator can absorb an ambiguous signal and judge it. So the audit can afford to be more sensitive than it could when you were the only reader — rules previously rejected for false positives are worth revisiting now that a judge sits between them and you. That widens what the system can catch without widening what you have to read.
A delegated decision that leaves no record is indistinguishable from no decision. If the Navigator is making calls on your behalf and they vanish into a log nobody reads, you have lost sight of what was decided for you — which is the same complaint that started all of this, moved up one level. Trusting the agent is not a design; the record is.
Two levels, because not everything deserves the same weight:
You review them in batch, not one at a time — a single "calls made for you" section in the day's delivery record, read once, most days in under a minute. Escalation is the exception, batch review is the norm.
When you disagree after the fact, reversal is cheap by construction. The Navigator only ever changes the plan, never the work, so undoing a call means reopening or editing an issue — seconds, no code involved, and both the original call and the reversal stay on the record against the same identifier. That property is the reason the "never touches the work" line has to stay absolute even as the judgment grant widens.
The Navigator may not merge, may not publish, may not correct another agent's work, and may not edit the concierge's state. Those are absolute, and they are what makes the judgment grant above safe to give: the wider its discretion over the plan, the harder the line around the work has to be. A role that can both decide freely and reach into what was built is one that can quietly rewrite history; a role that decides freely but only ever moves issues around is one whose every mistake is visible and undoable in seconds.
The Navigator as already approved is a file-writing verifier, explicitly "never a conversational dependency", observe-and-report only. Writing requirements is authoring, not verifying, and talking to the concierge makes it conversational. So the operating-officer role does not fit inside the boundary you already signed, and pretending otherwise would be a redefinition by stealth.
The proposed replacement line is narrower than "verifier only" and keeps the original intent intact: the Navigator authors the plan and never edits the work. It may write requirements, amendments, task issues and its own delivery file. It may not touch anything a worker produced, anything published, or anything the concierge owns. The original rule existed to stop a watcher silently repairing what it watches — that rule survives untouched.
Getting this wrong just moves the bottleneck, so it is worth being exact. The concierge has been doing the translating, badly, and that is what produced the five receipts.
| The concierge | The Navigator | |
|---|---|---|
| With you | Strategy, questions, status — answers in seconds, always | Only when you want to talk about the plan itself |
| An ask that implies work | Answers you, then hands it over verbatim. Writes no brief. | Files the requirement, then writes the brief from it, then spawns the worker |
| A worker closing | Nothing | Reads the timeline, judges the roadmap, appends the verdict |
| State it owns | Its own durable state file | The delivery file. Reads the concierge's state, never writes it. |
The one rule that must not bend: the concierge never waits on the Navigator before replying to you. It replies, then hands off. If that inverts, the Navigator becomes the latency problem you already cut the concierge off for.
On every ask, one question: is this about the shape of the plan, or about something inside it? Shape — a new goal, a change to what the lifecycle is — stays with the concierge and you. Everything else goes to the Navigator. That is a decision small enough to make in a second, which is the point: it does not reintroduce slow thinking on your response path.
You wrote, at around ten in the evening: "New rule for release candidate names: {package} Vx.x.x-RCx. No other summary allowed … always link to news.md in all RC PRs."
| What happened | What would happen | |
|---|---|---|
| 1 | The concierge wrote a worker brief from your words | The concierge answers you in seconds, routes it as requirement-level, hands it to the Navigator verbatim |
| 2 | The worker filed an issue describing its own work | The Navigator reads the open requirements, finds the release-scaffolding one already scoping "a documented release-candidate framework", and judges this an amendment rather than a new requirement |
| 3 | — | It amends that requirement with the naming rule and the body contract, one line saying what changed and why, then files the task as a child of it |
| 4 | The worker shipped; the requirement was never touched | The worker's brief is written from the requirement; it ships; at closeout the Navigator confirms the requirement moved and the task closed against it |
| Result | A shipped change nothing in the plan records | An amended requirement, a linked task, a confirmed closeout — and one line for you to read |
Note what the second column does not contain: a new requirement written to describe work already done. The Navigator recognised an existing one and amended it, which is the judgment the whole role turns on and the thing no rule could have produced.
Stand the session up and have it ship an operations improvement you can look at tonight — the four audit checks that missed last night's failures — while running the two reactive jobs alongside it: your asks become requirements first, and every worker that closes gets one line of verdict you can read.
The four checks are named by this morning's scorecard and are all mechanical: an issue closed or a pull request merged with no requirement above it (gated on work done, not on filing); an answer you recorded still unapplied after an hour; the registry and the published index disagreeing; an agent exiting having produced nothing. All four run across the seven project repositories, not just the hub — that is the scope change you called for this morning, and it is what makes last night's failure visible to a machine at all. It is textbook work for this role: it improves the machinery rather than policing it, it closes the holes from last night permanently, and it lands in an afternoon.
You asked for a feel for how the two of you work together, and that means watching it do something rather than watching it report. By tonight there is a shipped change with a requirement above it, a task beneath it, and a delivery record saying so — which is the whole thesis of this page demonstrated on itself.
Waits until it has earned it: counting exemptions and reporting the rate; a delivery view in the dashboard rather than a rendered file; folding the stall alarm in; letting the sweep wake it automatically on a detection; and any of the more sensitive audit rules that were previously rejected as too noisy for you.
Four things you settled this morning are built into the design above rather than re-asked here: that the Navigator is an agent; that its job is improving operations rather than clerking the roadmap; that it owns requirements, tasks, lifecycle and milestones while you own goals and strategy; and that it may decide within its domain and escalate only critical items. What is left is where the edges sit.
Two halves of one afternoon's work. Is "ship the four missing audit checks, take every ask as a requirement first, judge every closeout" the right first day? And do you accept that the Navigator now authors — requirements, and changes to the machinery itself — which the version of it you approved a few days ago was explicitly forbidden to do, being defined as a file-writing verifier, observe-and-report only, never conversational?
That boundary has to move for the role to exist at all, and moving it quietly would be a redefinition by stealth. The proposed replacement is narrower than it sounds and keeps the original intent: the Navigator authors and repairs the plan, and never touches the work. The old rule existed to stop a watcher silently mending what it watches — that survives untouched, because a plan repair is visible and undoable in seconds while a work repair is neither.
The safer alternative is to start with reporting only: let it judge closeouts and write the record, but ship nothing and leave the concierge dispatching for another week. It halves the risk, it leaves the actual complaint unfixed — dispatching without a requirement is the step that caused it — and it gives you a role you cannot form an opinion about, because you would be watching it observe rather than watching it work.
You said you suspect the split is goals and roadmap structure for you, everything else for the Navigator — worth ratifying rather than inheriting. Goal-splitting is not asked about here; you have settled that it is strategy and yours.
The reading built into this design: structure is the goal set plus the shape of the plan itself — what stages a requirement passes through, what counts as a requirement, what a milestone means across the portfolio, and rules like "one requirement covers exactly one release". Everything inside that shape is the Navigator's: creating requirements, scoping and splitting them, moving stages, assigning milestones, filing and closing tasks, all without asking.
Two edge cases worth naming, because both arrive in the first week. Splitting one requirement into two is inside the shape and the Navigator just does it; deciding that requirements should have a new stage is the shape and comes to you. Likewise, retiring an audit rule that turned out to be noise is operations and needs no permission; changing what a milestone means across the whole portfolio is not.
A rule that charges a full requirement for every one-line fix will be abandoned inside a week, and "use judgment" is not something you can hold anyone to. So the floor has to be written down.
The proposal is that the floor is consequence, not size: a requirement is needed whenever the work changes something you or a future reader would notice — anything that ships in a release, changes a rule or a contract, or belongs in the release notes. Exempt: a typo, a dead link, a failing check, or reverting something broken within the hour. A one-line change that alters a rule still needs one; a fifty-line documentation tidy does not.
What stops the exemption swallowing the rule is that it is written, not assumed: the worker records the call in one line on its own issue, so exemptions are countable. If much more than a third of a night's work is claiming exemption, the floor is in the wrong place and the count says so before you have to notice.
You are fine delegating judgment and want only critical items. "Critical" already has a definition that shipped last night in your dashboard: a pin capped at three, earned only by a blocking reference confirmed still open or by a computed condition, never by an agent's opinion. On its first night, nothing earned it.
Inheriting that bar means one definition of urgent rather than two that drift apart, and it is already calibrated to your standing objection — that you will be annoyed by anything called critical that is not. The proposed addition is one clause: an escalation must name the specific thing only you can do. If it cannot, it is a call the Navigator should have made itself.
Structural proposals do not go through this gate at all. They reach you because they need your permission, not because they are urgent, and they wait in the batch review rather than interrupting you.
This is the risk inside the delegation grant, and it should not be left as "trust the agent". A delegated decision that leaves no record is indistinguishable from no decision — and losing sight of what was decided for you is the same complaint that started this, one level up.
Two levels are proposed. Every judgment gets one line in the day's delivery record: what it decided, on what, and why. A call that changes the plan gets a permanent identifier — filing a requirement, amending one, granting an exemption, closing something as out of scope — allocated in the same never-reused style as the config items, so anyone can point at one and reverse it.
The primary reader of that record is the concierge, not you. It reviews the Navigator's calls daily and is answerable for having done so; you see the critical items and the proposals. That settles the format question: the record should be dense and complete enough to audit a day's judgment in one pass, rather than trimmed to be readable on a phone.
Reversal is cheap by construction, and that is the property that makes the grant safe: because the Navigator only ever changes the plan, undoing a call means reopening or editing an issue. Both the call and its reversal stay on the record.
Neither half of the Navigator's job is periodic: it thinks when you ask for something, and when a worker closes. At last night's volume that is roughly a dozen to two dozen wakings, a few cents each, plus one pass at the end of the night. A session that instead polls every five minutes reaches the same answers for ten to fifty dollars a day — up to three hundred where its memory expires between wakes, which a five-minute poll makes likely rather than unlikely.
The wrinkle worth naming: "standing" and "waking" are the same session here. It stays up so you can talk to it from your phone; what should never exist is a timer that wakes it to look around when nothing has happened. The stall-watching that would justify polling is already done, for free, by the script.
You said you could go either way on the manager. Under the model you have now described it has no separate job left: "makes sure the workers are delivering" is the second half of the Navigator's own sentence, and the stall-watching half is a script either way — so a separate role would mostly own a scheduled job.
The honest counter-argument is urgency: a stalled worker needs a response in minutes, roadmap judgment can wait until closeout, and one role holding both risks the urgent thing queueing behind the slow one. That is solved by the mechanism split rather than by another person — the five-minute script raises the alarm on its own and can replace a dead worker without waiting for anything to think.
What you lose by folding: nothing measurable. What you gain: four roles instead of five, and exactly one officer accountable when work ships and the plan does not move.
The closeout contract says a worker ends in one of three things: a pull request for a planned release, a question for you, or a request for config work. Only the first advances a requirement — and that is right, not a gap. A worker that correctly stops to ask you something should not be marked down for failing to move a stage. So: a pull request moves the requirement into implementation and closes it on merge if complete; a question parks it with the blocker named on it; a config request does the same. Two of the three annotate rather than advance, and today a worker that stops for an excellent reason leaves its requirement looking untouched — indistinguishable from one that did nothing.
When work fits no requirement there are three cases and only two mean "file one": a genuine new need, file it; a piece of something already open that nobody linked, attach it; neither, and the work should not have been started — stop and ask, rather than writing a requirement backwards from finished code.
Which answers last night's six. Five of them are not missing a requirement at all — they are missing a link to requirements that are still open: four pieces of Operations Dashboard work, and the release-candidate naming rule that belongs to the release-scaffolding requirement. Attaching them and adding one line to each parent recording what shipped is about a quarter of an hour, and it produces a true record rather than a reconstructed one. The sixth traces to a decision you signed that day, so cite it there and leave it.
The alternatives cost either most of a working day writing requirements backwards from finished code — which teaches the system that shipping first is fine because the paperwork follows — or nothing at all, which leaves the most-worked-on thing in the program showing no work against it.
Deliberately not here: whether the standing sessions survive their own memory being compacted. That is the concierge's problem, it has its own requirement, and folding it in would make this page about two things.
What this page replaces. It folds in the worker-closeout decision and the worker-supervision decision of 2026-08-15, at your request for one document rather than several. Both pages remain, unchanged and cited; their questions are restated here in this design's terms, and both now point here. Nothing from either is withdrawn.
Sources. The six overnight issues, their milestones, bodies and parent links, read from GitHub on the morning of 2026-08-16; the child-issue and comment history of the two requirements they belonged to; the eight requirements filed overnight and their authors; the live sweep state file, its scheduler entry and its five-minute cadence, read from this machine; the concierge's launcher script; the stall-detection measurements, cost model and the dead-worker case from the supervision decision; the three-outcome contract from the closeout decision. Provenance, method and assumptions in this artifact's README.
Citations (for a deep dive only — the argument stands without them): the Navigator requirement, the autonomy goal, the worker closeout decision, the worker supervision decision, the Operations Dashboard requirement and the release-scaffolding requirement.
This decision artifact was drafted by Claude Code using Opus 5 and reviewed by @jwildfire.