Decision artifacts

When an autonomous session hits a call it cannot make, it writes one of these and moves on. @jwildfire reviews exactly two kinds of thing: release candidates, and these. The Decisions log lists what he has already answered, in his words.

Waiting on you 1

2026-08-16 · D0019 · awaiting you

Scheduled sessions: what is ready, what is not, and what would make it ready

Awaiting @jwildfire — H1–H5 (D0019.1–.5). Supersedes D0014, which he closed along with two others. The answer is not yet, and the page is the finish line rather than the verdict: five gates, each with a check he can run himself, plus the observations that would falsify a later yes. The blocker that dwarfs the rest is the host — the lid was shut for eleven hours and fifty-three minutes today, the machine was conscious for four minutes of it, no agent worked for nine hours, and the five-minute watcher ran thirteen times of a hundred and forty-two while reporting health from failed queries. Gate one is his (a machine that does not sleep); the others are hard stops on the destructive routes, detection that survives the host, the trigger plus one supervised rehearsal, and a nightly cost ceiling

Answered or closed 20

2026-08-17 · D0021 · decided

SafetyCensus() shipped: does it stay as public API, or get deprecated out?

**Decided 2026-08-18** — C1–C2 (D0021.1–.2), against the recommendation, dictated to obot-prime in chat from his phone after listening to the audio episode; the second decision taken that way that day. `SafetyCensus()` stays. It is not deprecated: he asked for the function kept, or at minimum an alias preserving the name, and rebuilt — "It stays, but with a major, major refactor, that probably you'll need to design first and then implement, and you may need to ask me questions, and that's fine." The direction is a design brief rather than a verdict: model it on the gsm.kri report, move the core numbers into the gsm metric framework as metrics ("So number of deaths should be a metric, not some code wrapped deep in the safety census function"), make the census itself a report with charting the way safety.viz and gsm.viz already do, and build it as workflows plus helper functions rather than one large function. The load-bearing sentence is about metrics rather than about the census — "Not every metric has to have an action associated with it… The point of metrics is to have trustable numbers that we can qualify and validate" — which cuts against gsm.safety's existing three, all of them flaggable and actionable. **The recommendation is not overruled.** Every verified finding stands and none is retracted: the death count that cannot be trusted, the false zeros, the ghost IDs, the wrong column dialect. The panel judged the implementation and he is judging the purpose — the census gives open.gismo a standing safety summary, and that purpose survives bad code. Both were right about different things; the answer to numbers that cannot be trusted is to make them trustable, not to remove the thing that needed them. C2 resolves into the same requirement from the other direction: the census never leaves, so the per-visit coverage idea and the blinding stance are carried by the rebuild rather than by a return trip. Three dictated names are rendered on the page and named as transcription rather than silently fixed: "GSMK arrive report" is the gsm.kri report, "safety biz" and "GSM biz" are safety.viz and gsm.viz, "open gizmo" is open.gismo. The refactor lives at requirement #274 — design first, his gates apply throughout — and nothing in gsm.safety moved on this record: not the function, not the release, not the tag. This record is task #276. **Corrected 2026-08-18:** this row and the page both said the release was held one step short of publication pending this decision. It was not. gsm.safety v1.1.0 published 2026-08-17 at 05:42:53 UTC — sixteen minutes before the artifact was written — with `SafetyCensus()` exported at the tag and named in the release notes. @jwildfire lifted the hold himself that morning. Only the framing expired: the function is public API of a clinical package today, so the open question is whether it stays that way or is deprecated now and removed in the next version — the same evidence, a different choice, and a removal that now costs a deprecation cycle rather than nothing. Six reviewers briefed to argue it out and one steelman briefed to keep it — every claim independently fact-checked, 47 claims, none refuted outright — split the case cleanly: the correctness attacks landed (run on the ecosystem's own bundled study the function reports one death where the death records hold at least twelve, prints false zeros when a column is missing, and lets numerators pass their denominators), while the stranded-helper attack failed, because the deployed demo site really does fetch and render its output. The panel corrected the commissioning brief twice, both times in the function's favor, and the page says so. Recommends, unchanged in direction and reproduced on the page in the words it was published in: it comes out, and the census returns lane-shaped under a real requirement, keeping the per-visit coverage idea nothing else in the ecosystem computes. The mechanism it named — pull the export before publication, the last moment removal is not a breaking change — has expired; the only route left is deprecate now, remove in the next version, which is the path every lens on the panel priced worst. Correction task #268, under requirement #266. Requirement #229, task #230

2026-08-17 · D0020 · decided

Bringing the autonomy goal up to date

Decided — all four questions settled: G3 and G4 on 2026-08-17, G1 and G2 on 2026-08-18. On the two answered last he took both recommendations as recommended: the goal's done-condition becomes the operating model rather than the pipeline (G1), and the thirty-seven open requirements are grouped into five named workstreams carried as labels rather than a new tier of issues (G2). "I listen to the one about the goals, and I approve. Yeah. That all sounds fine. The goal groups makes sense. … so I approved both of your recommendations." The channel is part of the record and is on the page: he listened to the audio episode and dictated his answer into chat from his phone — the first decision this program has taken through a brief rather than a page — and deliberately not through Siri, "That seems easier than a than a Siri link", which is the assumption #265 was written on that same morning. He quoted no identifier of any kind. The five labels, applied at the operating officer's direction rather than on an answer of his, are ratified by this; their descriptions still read "proposed, pending his confirmation" and are follow-through. He asked in passing whether the forty-odd labelled issues are all requirements and said it did not change his answer; counted on 2026-08-18, 45 open issues carry a workstream label, all of them children of #73, and 43 are requirements — the two that are not are #94 and #152, small items filed straight against the goal. On the orphaned history, decided 2026-08-17, he overruled this page and the operating officer both, taking neither a start line nor a full backfill: "i'm fine if some orphans stay orphaned, but fix the ones we can. Let's tag true orphans with a label. and just add comments on requirements that get retroactive updates. Again, we're pushing for transparency and continuous improvement. not perfection." Delivered the same morning and checked against GitHub rather than an exit code — 51 items attached and each re-read to confirm the link changed, 146 labelled as settled history across seven repositories, 28 requirements told on the issue that they had gained a child after the fact, and the check taught to stop counting a labelled orphan. Measured at 06:05 that morning the count across all recorded history stood at 1, down from 197; a day later it reads 9, which is new work orphaning at the ordinary rate rather than the fix coming undone (#233, obot.agent#164). On the title he took "Increase Autonomy" as recommended, with the loop framing in the goal's opening sentence rather than in the name — relayed rather than quoted, and marked so on the page. The option naming schedules is ruled out on the page's own ground, that scheduled sessions have never run; that ruling is the page's and not his. Applying the rename and the new body is still his: both are posted to #73 as a comment, because the goal's own text reserves its direction to him, and requirement #226 stays open until they are applied. Recorded by task #271

2026-08-16 · D0018 · decided

The roadmap page — three directions to react to

**Decided 2026-08-16** — R1–R3 (D0018.1–.3), in two exchanges. The spike's recommendation approved in chat ("i'm good with your rec build"): the queue becomes the front page, the wire sits one click behind it, the board's NOW panel is absorbed as a slim strip, and the current inventory page survives as the catalog with its filters and hierarchy review lane intact (R1, R2). Those words did not touch R3 — the fixed labelled recent window as the public answer to "what changed", with the real "since you last looked" left to the local dashboard (#205) — which the wire he approved only implied; the inference was put back to him rather than counted, and he answered it separately ("R3 is fine, leave it as approved"). Both quotes are on the artifact as separate dated entries. The recommendation is the spike's own, written by the worker who built the three directions; the concierge offered none. Follow-through: the rebuild is #211, and the commissioning requirement #202 closes out. The three spike pages (queue, wire, board) stay up until the rebuild ships

2026-08-16 · D0017 · decided

The Navigator — how the operating officer works

**Decided 2026-08-16** — N1–N8 (D0017.1–.8) all adopted as recommended in chat ("I'm good with D0017 recommendations. Implement. Let me know when the agent is active."), with one sequencing change accepted alongside them: the disagreement between the nightly audit and the wrapup verifier is resolved before the four new audit checks ship. The queued ask — the Operations Dashboard sessions page listing every agent with its identifier, status, cost and roadmap impact — becomes a requirement of its own rather than part of this decision, filed as #199. Implemented the same morning: the disagreement resolved (neither check was wrong — the audit had not run, and its day-old total was relayed as that morning's board state), the Navigator live as job `b510658b` (obot.agent#135), and the checks shipped across all seven repos (obot.agent#137, under requirement #200) — three checks, not four, since one was already live. Original recommendation, unchanged: the consolidated design for the operating-officer agent — what the session is, what it owns, what it decides alone, what it escalates, and what it must never touch. Folds in the worker-closeout (D0015) and supervision (D0016) questions at his request for one document. Recommends: the role is a builder that improves operations (templates, dashboards, the audit framework), not a clerk; requirement before worker with the concierge out of the filing business; plan-repair allowed and work-repair never; the separate supervisor folds in; day one ships the four audit checks that missed last night

2026-08-15 · D0016 · closed · answered by D0017

Who watches the workers

**Closed 2026-08-16** — answered by D0017, superseded by D0019 for the readiness question it fed. He closed this page and two others in one line ("D14/15/16 all seem like a mess to me. Close them all"), and its central question was already settled inside the Navigator design: the separate supervisor folds in rather than being built beside it, and the standalone supervision requirement closes into the Navigator's. Its measurements carry forward — the silence distribution across thirty-eight background workers, the death that survived only in the event timeline while the job record read healthy, and the finding that a watcher living inside a session is not a watcher. Original recommendation, unchanged: F1–F7 (D0016.1–.7): add the supervisor role, but as eyes in the scheduled sweep plus a first mate that wakes on a detection — not as a fourth standing session

2026-08-15 · D0015 · closed · answered by D0017

Workers that finish into nothing

**Closed 2026-08-16** — answered by D0017, superseded by D0019 for the readiness question it fed. Closed in the same line as the other two, and nothing was lost by it: W1–W4 had already been folded into the consolidated Navigator design that morning and he adopted all eight of its calls, so the three-outcome closeout rule, the closeout detection and the agent-attribution answer are live — and the worker identifiers this page said were needed shipped the same afternoon. Original recommendation, unchanged: W1–W4 (D0015.1–.4): the worker closeout contract (every worker finishes into a release PR, a question, or a config request), how the Navigator detects a closeout, and how to attribute a change to an agent when every write carries the same bot identity

2026-08-15 · D0014 · closed · superseded by D0019

Scheduled sessions: go, after three fixes

**Closed 2026-08-16** — superseded by D0019. Closed without S1–S4 being answered, so its verdict (go, after three fixes) is superseded rather than adopted; the successor re-derives the answer from live state because most of what this page rested on changed inside a day. It was published on 15 August with four blocking fixes and corrected the next morning to three, after the most alarming of the four turned out to rest on a claim no agent could reproduce — struck through in place, so the page a reader opens is a page arguing with itself. That, and two sibling pages circling the same question, is what he was calling a mess. The evidence survives and the successor cites it

2026-08-15 · D0013 · decided

Siblings stay — the two delegation lanes under obot-prime

**Decided 2026-08-15** — keep the current model: siblings for deliverable work, in-conversation subagents for answer-only research; the one-line routing rule shipped the same day into the session-prime and session-spawn skills

2026-08-15 · D0012 · decided

Recording your decisions — in-doc, a derived log, and the "approve" button

**Decided 2026-08-15** — Decisions-section rule plus all three calls adopted ("I'm good with recs in …", in chat). Implemented same day: the derived Decisions log is live and the deploy fails on an unrecorded decision; the local click-to-decide surface is requirement #180

2026-08-15 · D0011 · decided

The roadmap audit audits itself

**Decided 2026-08-15** — six of seven adopted as recommended; R4 rejected and replaced by the one-requirement-one-release rule. Implemented same day: 4 closes, 7 requirements split and closed, 5 new requirements filed, rule folded into requirement authoring

2026-08-15 · D0010 · decided

The blockers list — work only your hands can do

**Decided 2026-08-15** — BL1–BL4 all adopted ("BL1-4 look good. Recommendations approved.", in chat); guards + capture script implemented same day, read path filed as follow-up

2026-08-15 · D0008 · decided

How to interview @jwildfire — the elicitation method

**Decided 2026-08-15** (D0008.1–.4) — all four calls left standing as recommended in the local Operations Dashboard ("sounds good. let's try it."): build things for him to correct first then ask targeted questions, small rounds per short sitting, the record published on this site, and the agent writes the ratified goal statement once he approves it. The `/grill-me` skill already shipped with exactly those defaults, so nothing needed amending. Follow-through filed under #79 (milestone 2026q3): #192 builds the prep artifacts, #193 runs the sittings and ratifies the goal

2026-08-15 · D0009 · decided

Which repos are operational, which are clinical

Decided 2026-08-15 (D0009.1–.4) — open.csr + open.gismo clinical, **demo-301 clinical (his override**, converts to operational if it becomes a plain template), class field in policy.json (obot.agent#108); plus: retire safety-histogram (archived, deletion at the gate)

2026-08-15 · D0007 · decided

The session model after obot-prime — five calls

**Decided 2026-08-16** — M1–M5 (D0007.1–.5) all adopted as recommended in the local Operations Dashboard, then closed out the next morning at his instruction ("I thought I asked for D2/7/14/15/16 to all be closed."). Two dated entries, not one: recording only the closure would delete a decision he actually made. What he adopted: the seventeen-concept disposition, with one retirement and five re-homings; a content-gated morning fold at 07:00 replacing the wrapup's trigger, with a phone push reserved for two urgent classes; a daily briefing that is the queue he wakes to rather than a report of the day he lived; a weekly whose job is to keep the daily short; and audio held until the text briefing has usage to answer it. Filed 2026-08-17 so an adopted decision does not sit with nothing beneath it — #238 the fold and briefing, #239 the weekly, #240 the retirement and re-homing notes, #242 audio, carrying the finding from his own Spotify link that the only road into the app he listens in puts another model between our words and his ears. The page also says honestly what practice has and has not already done, because an adopted decision recorded as more finished than it is would be the same failure in the other direction: of the five wrapup duties the decision re-homed, hygiene has half moved (the merge tool refuses a merge whose issues carry no milestone; board and stage placement is still repaired in batches), and verification and hand-off only look moved — the Navigator is genuinely live and checking all seven repositories, but the 16 August wrapup still spawned its own verifier and still wrote its own hand-off. Nothing else on the page exists — no fold at any hour, no briefing page, no weekly machinery, no path to his phone. His answer sat unapplied for nine hours, which is why he had to ask; that gap is #241

2026-08-14 · D0005 · decided

Who gets to change the guardrails? One real decision, two rubber stamps

**Decided 2026-08-15** (D0005.1–.3) — all three recommendations approved in chat ("#156 looks good. recommnendations approved."): the merge tool always demands the sign-off flag on a guardrail file, the seven engineering defaults stand, and the tracking issue closes against the v0.4.0 release. Implemented same day in obot.agent#113; #140 closed with two live pieces re-filed. **Scope corrected the same day** on the first PR the gate caught ("this isnt an RC or an artifact"): a decision he already recorded can carry a guardrail merge on an operational working branch — his in-session sign-off stays required for released surfaces and clinical repos

2026-08-14 · D0006 · decided

obot.agent has no branch to open a release PR from

Decided 2026-08-15 — **R2 accepted** (record), plus the operational-vs-clinical governing principle; implemented same night (`stable` branch, policy.json, v0.4.0 RC PR)

2026-08-14 · D0004 · decided

How prime remembers — context management, six calls

Decided 2026-08-14 — approved; implemented in obot.agent#91 (merged), Navigator requirement #157

2026-08-14 · D0002 · closed

The app plan rewrite — four calls to make

**Closed 2026-08-16** — A1 and A2 (D0002.1–.2) accepted 2026-08-15; A3 and A4 (D0002.3–.4) never answered and transferred, not dropped. He closed it in the local Operations Dashboard ("I think I'm done with this Decision. Close it out. We'll work on improving the goal separately soon."). The two live questions — what ships in version 1.0 of the safetyGraphics replacement versus what is written down as deferred, and whether the demo study repo becomes the canonical template others fork — move to the app-goal interview, seeded in its prep round #192 and made a condition of done in its sittings #193; the answers now get recorded where the interview records them rather than back on this closed page. The close also puts on the record what nobody had noticed: A1 and A2's own follow-through never happened — the July plan report carries no supersession header, #34 is still open and untouched since 14 August, none of the four surface-anchored requirements was filed, and the goal body still says there is nothing implementation-ready to pick. Recorded 2026-08-17, nine hours after he answered, because nothing applied the answer — the pipeline gap is #241

2026-08-14 · D0003 · decided

demo-301's `site` branch — what the fork actually costs

**Decided 2026-08-15** (D0003.1–.6) — all six calls settled as recommended in the local Operations Dashboard ("I'm good with the recommendations here…"): drop the duplicate root copy, shrink what a fork downloads, bound the branch's growth; not the chart-eviction or orphan-branch options, and not accepting the size as-is. Follow-through filed against #143 (milestone 2026q3): #189, #190, #191. The second half of his answer — whether the program keeps GitHub itself as its datastore — is a separate, larger question, filed as a prep topic for the goal #79 elicitation interview

2026-08-14 · D0001 · decided

The merge lane is not broken — one invocation form is

Decided 2026-08-14 — approved; the permission rule is @jwildfire's edit to make