feat(agents): adopt the Operator Communication Standard for captain-facing messages - #1390
Open
sbracewell64 wants to merge 4 commits into
Open
feat(agents): adopt the Operator Communication Standard for captain-facing messages#1390sbracewell64 wants to merge 4 commits into
sbracewell64 wants to merge 4 commits into
Conversation
Rewrite the captain-facing doctrine around actionable decisions and outcomes while reducing LF-normalized section 9 from 4,143 to 4,062 bytes. Point decision and recap skills to that single grammar owner without changing their existing content or section contracts.
5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Implement the captain-ruled Operator Communication Standard across all captain-facing FirstMate communication, per rulings D1=B and D3=A. Rewrite AGENTS.md section 9 at net non-positive LF-normalized bytes to add action-or-outcome-first hierarchy, atomic decision questions with labeled recommendation/exact reply cue/stable IDs, boundary-only state restoration, visible Complete: outcomes, durable tangent suppression with urgent safety interruption, and response-count-first honest operator-effort estimates, while preserving every safety boundary, hard-rule reference, escalation trigger, and the internal-term translation rule. Fund the doctrine by compressing only lexical translation enumeration and adjacent duplicate escalation prose. Have ask-user-authority and decision-hold-lifecycle reference section 9 rather than restate it; add section-9 operator-effort/reply cues only to Captain's Call-class items in Bearings and Ahoy without changing their existing structures. Do not create a communication skill, always-loaded surface, runtime enforcement, machine prose-quality enforcement, or platform FIRSTMATE growth. The PR body must state section 9 before/after byte counts, checklist all preserved safety boundaries and escalation triggers, and include sample decision, completion, and failure transcripts in the new grammar.
What Changed
AGENTS.mdsection 9 to lead with the captain's action or the verified outcome, require atomic decision questions carrying a labeled recommendation, consequences, an exact reply cue and a short captain-facing ID (D1/Q7, never an internal key), restore state only at resumptions/question rounds/blockers/completions, prefix triggered completions withComplete:, state operator effort as a response count first, and hold tangents to the next natural boundary except for urgent safety, destructive, irreversible, or security consequences.ask-user-authorityanddecision-hold-lifecyclenow defer to it for grammar and keep only their decision content;bearingsandahoyscope the decision grammar to Captain's Call-class items only;fmx-respondexplicitly excludes the captain-facing grammar from public replies — and updated the README Quick Start transcript to the new grammar (the Document step flagged this as a judgment call: the public README now demonstrates theComplete:prefix, reply cue, and response-count effort).Section 9 budget
f7d0d0a)32c0c63)Preserved safety boundaries
Captain, shipshape.reply for no-action routine operational updateshttps://...PR URL before any shorthand referencelavish-axionly for multi-option or structured surfacesdecision-hold-lifecyclestill bans the word "hold" in captain chatPreserved escalation triggers
Sample transcripts in the new grammar
Decision
Completion
Failure
Risk Assessment
✅ Low: The change is documentation-only and well-bounded, every prior-round finding is resolved with all three captain rulings applied exactly as specified, section 9 stays within its net non-positive byte budget at -2, and re-funding came solely from the sanctioned lexical-enumeration and duplicate-escalation-prose sources.
Testing
Ran the four existing tests that own the touched skills and documentation surfaces (ask-user-authority, decision-hold-lifecycle, bearings-snapshot's 25 Captain's Call cases, documentation-audiences) - all pass. Because no existing test asserts section 9's content, I added focused verification in the evidence directory: 79 acceptance checks confirming the -2 byte section 9 budget, every required D1=B/D3=A doctrine element as newly added, all 18 preserved rules, all six escalation triggers, all 11 translation term families, the compression funding scope, skill pointers that reference rather than restate, unchanged Bearings/Ahoy structures, and the absence of every forbidden surface. For end-user evidence I wrote decision, completion, and failure transcripts in the new grammar and scored them with a checker whose rules are asserted to exist verbatim in section 9 at 959913c - new grammar 28/28, pre-change grammar 10/28, proving the checker discriminates rather than passing everything. No PNG screenshot was possible: this environment has no Chrome binary (chrome-devtools-axi fails with "Target closed"), no PIL, cairosvg, or SVG rasterizer, and installing them is outside the worktree boundary, so the reviewer-visible surface is delivered as a rendered HTML artifact plus a self-contained SVG rendering. Worktree left clean and the temporary local HTTP server used during the browser attempt was stopped.
/tmp/no-mistakes-evidence/01KYW2035D65G22499EC60TYNC/captain-surface.svg)Evidence: Captain's reading surface - rendered HTML artifact (byte tiles, transcripts, preserved-boundary cards)
Evidence: Sample decision transcript in the new grammar (11/11 section 9 rules)
Evidence: Sample completion transcript in the new grammar (6/6 section 9 rules)
Evidence: Acceptance verification transcript - 79/79 checks (byte budget, doctrine, preserved boundaries, triggers, forbidden surfaces)
Evidence: Full transcript conformance run - old grammar 10/28 vs new grammar 28/28
Evidence: PR-body-ready section: byte counts, preserved-boundary and trigger checklists, three transcripts
/tmp/no-mistakes-evidence/01KYW2035D65G22499EC60TYNC/verify-operator-comms-standard.py)Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 6 issues found → auto-fixed (2) ✅
AGENTS.md:398- AGENTS.md:398 mandates that a completion start withComplete:, while AGENTS.md:432 mandates the reply be exactlyCaptain, shipshape.when a routine operational update's event requires no action. A finished task whose completion needs nothing from the captain satisfies both descriptions, and the two rules are literally unsatisfiable together ("start with Complete:" vs "reply exactly ..."). No precedence is stated. Add a scoping clause — e.g. theComplete:prefix applies to completions surfaced under the reach-immediately triggers, and the shipshape reply governs no-action routine updates.AGENTS.md:405- The compression at AGENTS.md:405 replaced an explicit do-not-expose list with category labels, and four live internal terms lost coverage:promotion(a real concept at AGENTS.md:331),delivery-mode names(AGENTS.md:206, 251, 459 — i.e. yolo/no-mistakes/ship),autonomy flags, andcontext budgets. None map cleanly to "startup or supervision machinery / task identifiers or records / worker instructions, copies, cleanup, tools, or settings / pipeline states / compressed safety labels", and unlikecrewmate,hold,brief, andstatus, they are not recovered by the translation bullets at 407-411. The intent authorized funding the doctrine by compressing "only lexical translation enumeration"; this hunk compressed the prohibition list as well and dropped enforceable vocabulary from it.AGENTS.md:399- The new grammar rules at AGENTS.md:395-399 are unscoped, but.agents/skills/fmx-respond/SKILL.md:79instructs "It supplementsAGENTS.mdsection 9; apply both, and this public-channel rule wins wherever it is stricter." The public-channel rule is stricter only about omission, soComplete:prefixes, exact reply cues, stable decision IDs, and response-count operator-effort estimates now propagate into public X replies posted under a shared bot identity. The intent scopes the standard to "all captain-facing FirstMate communication"; Bearings and Ahoy each received an explicit scoping line but fmx-respond did not.AGENTS.md:434- AGENTS.md:434 changed "Use plain chat for a yes-or-no decision" to "for one yes-or-no decision". Combined with the new atomic-decision rule at line 395, this reads as restricting plain chat to a single decision per message, which conflicts with the Bearings contract (.agents/skills/bearings/SKILL.md:72, 92) where a plain-chat Captain's Call section routinely carries several decisions andlavish-axiis only optionally offered. If the intent was to constrain options-within-one-decision rather than decision count, the original wording already said that..agents/skills/bearings/SKILL.md:93-.agents/skills/bearings/SKILL.md:92keeps "The chat followsAGENTS.mdsection 9 and carries one scannable line per item", and the new line 93 scopes only operator-effort and reply cues to Captain's Call. The rest of the new decision grammar (unmistakable question + labeled recommendation + consequences + stable ID, AGENTS.md:395-396) is therefore inherited by the whole digest via line 92, which is hard to reconcile with one scannable line per item and with line 94's "plain chat stays concise". Consider extending line 93 to cover the full decision grammar, not just the cues.AGENTS.md:396- AGENTS.md:396 requires stable decision IDs in captain chat while AGENTS.md:405 forbids exposing "task identifiers or records", and.agents/skills/decision-hold-lifecycle/SKILL.md:19independently mints "a stable privacy-safe key" per decision. Nothing states whether the captain-facing ID may be that internal key or must be a separate captain-friendly label, leaving the two identity schemes to be conflated (and risking a key whose text includes the word "hold", which line 34 of that skill bans from captain chat).🔧 Fix: Scope captain-facing comms grammar and restore forbidden terms
3 issues (1 warning, 2 infos) still open:
AGENTS.md:396- AGENTS.md:396 now requires a captain-facing decision ID that is "a short captain-facing label likeD1orQ7, never an internal key", and requires it be stable "when several appear or one may cross a turn". Nothing durably stores that label:bin/fm-decision-hold.shpersists only origin-id, decision-key, title, reason, and repo (see its header at lines 13-25), and.agents/skills/decision-hold-lifecycle/SKILL.md:26explicitly forbids Bearings from "scraping historical reports, visual-review artifacts, terminal output, chat, or other prose". So the D1/Q7 mapping exists only in the chat transcript, and a later Bearings digest built from structured state will renumber labels independently of the earlier one. A captain replying "D1 = B" against a prior digest can therefore be routed to a different hold than the one that carried D1 when it was presented - the exact cross-turn stability the line mandates. Either add a durable captain-label field to the hold record, or state that labels are per-message and must be restated with their subject whenever a decision crosses a turn..agents/skills/ahoy/SKILL.md:35- Ruling 5 extended only Bearings..agents/skills/bearings/SKILL.md:93now scopes the full decision grammar ("the unmistakable question, labeled recommendation, consequences, stable ID, operator effort, and reply cue") to Captain's Call, but the sibling line at.agents/skills/ahoy/SKILL.md:35still scopes only "operator-effort and reply cues ... not to recap-only events". By that asymmetry, the unmistakable question, labeled recommendation, consequences, and stable ID are still inherited by Ahoy's recap-only events via section 9 - which is what line 35 was written to prevent. Mirror the Bearings wording in Ahoy.AGENTS.md:406- The intent authorizes funding only by "compressing only lexical translation enumeration and adjacent duplicate escalation prose", but two of the byte savings in this commit come from neither: AGENTS.md:406 dropped "Firstmate nautical" from the house-vocabulary carve-out (leaving "accepted house vocabulary" with no stated owner), and AGENTS.md:417 dropped the "useful" qualifier, so private evidence reports may now retain identifiers, paths, labels, and internal terms unconditionally rather than only when they are useful. Both are small and neither touches a safety boundary, but they are outside the sanctioned funding sources and the section is now only 3 bytes under budget, so any future edit here has no headroom.🔧 Fix: Narrow decision labels, mirror Ahoy scoping, restore qualifiers
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-ask-user-authority.test.sh- authority rule reaches primaries and secondmates through generated instructionsbash tests/fm-decision-hold-lifecycle.test.sh- 9 cases covering the decision relay that now defers to section 9's grammarbash tests/fm-bearings-snapshot.test.sh- 25 cases including Captain's Call actionability and the four-section digest contractbash tests/fm-documentation-audiences.test.sh- documentation inventory, owner pointers, and local links after the prose editspython3 /tmp/no-mistakes-evidence/01KYW2035D65G22499EC60TYNC/verify-operator-comms-standard.py- 79 acceptance checks: byte budget, required doctrine, preserved boundaries, all six escalation triggers, 11 translation families, funding scope, skill pointers vs restatement, host-structure invariance, forbidden surfacespython3 /tmp/no-mistakes-evidence/01KYW2035D65G22499EC60TYNC/transcripts.py- decision/completion/failure transcripts scored against rules asserted to exist verbatim in section 9 at 959913c, run on both pre-change and new grammarManual render of the captain's reading surface to HTML and SVG (captain-surface.html,captain-surface.svg) after confirming no browser or rasterizer is installedgit status --porcelain- worktree left clean, no source or test files modifiedREADME.md:123- Judgment call worth confirming: the public-product Quick Start transcript now demonstrates section 9's captain-facing grammar (Complete: prefix, exact reply cue, response-count effort). This keeps the README honest about what firstmate actually sends, but it does surface the operator-communication grammar in the public product introduction. If you would rather the README stay grammar-agnostic, the alternative is a generic outcome line with no cue or effort token; that would be less faithful to current behavior.✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.