feat(bin): add harness adapter interface - #7
Closed
sbracewell64 wants to merge 33 commits into
Closed
Conversation
* feat: add verified pi-signed adapter * no-mistakes(review): Correct pi-signed maintainer verification date * no-mistakes(review): Correct remaining pi-signed verification dates * no-mistakes(review): Preserve authoritative pi-signed runtime identity * no-mistakes(document): Document pi-signed shared adapter semantics * no-mistakes: apply CI fixes
* fix(pi): rearm watcher across same-process session transitions Pi emits session_shutdown for ordinary /new, /resume, and /fork replacement as well as terminal quit. The primary watcher extension latched a module-level stopping flag on every shutdown, so a replacement session in the same process could not arm monitoring until Pi restarted. Own arm authority per session generation so only the active live generation may start, stop, or rearm the child. Replacement sessions can arm again without restarting Pi, stale prior-generation callbacks cannot mutate the active cycle, and real quit still blocks late rearm. * no-mistakes(review): Preserve Pi generation isolation and exit cleanup * no-mistakes(document): Correct Pi watcher transition documentation
* Consume quota-axi pace signals in dispatch profile array selection. Add quota-array-dispatch as the single owner of the pace-aware candidate choice, keep AGENTS.md to the intake boundary and load trigger, and cover the acceptance cases with sanitized schemaVersion 3 fixtures. * no-mistakes(review): Stop and report genuine quota dispatch ties * no-mistakes(document): Document quota pace freshness and uncertainty
…nguid#1171) * fix(grok): adapt Stop continuation to runtime capability * no-mistakes(review): Reject ambiguous Grok Stop payloads * no-mistakes(review): Reject duplicate Grok fields and accept spaced tmux sessions * no-mistakes(review): Enforce exact tmux cleanup selectors * no-mistakes(test): Fix historical tmux fixture and validate Grok Stop * no-mistakes: apply CI fixes
* fix(brief): make DOD scaffolding parse-safe on stock macOS Bash 3.2 fm-brief.sh built each Definition-of-done block and the not-enabled Herdr declaration with `VAR=$(cat <<EOF ... EOF)`. On Bash 3.2 (macOS /bin/bash) the lexer scans for the command substitution's closing `)` textually and tracks quote state through the heredoc body, so a single apostrophe, unbalanced quote, or unbalanced paren in that prose breaks parsing of the whole script. Every ship-brief scaffold (no-mistakes, direct-PR, local-only) failed with `unexpected EOF while looking for matching )`. Bash 4+ parses it fine, so the breakage stayed invisible everywhere except stock macOS. Replace all four command-substitution heredocs with `IFS= read -r -d '' VAR <<EOF || true`. That removes the `$(...)` wrapper and the entire defect class regardless of future prose, and preserves the variable expansion the direct-PR and local-only bodies need. `read` keeps the heredoc's trailing newline that `$(...)` used to strip, so trim one newline to keep every generated brief byte-identical to prior output. Guard the structure, not one historical phrase: a new test rejects any heredoc nested in a command substitution anywhere in fm-brief.sh, where the old assertion pinned a single apostrophe phrase and so missed the reintroduction. Extend the stock-macOS Bash CI job from parsing one script to the whole maintained shell surface (bin/*.sh, bin/backends/*.sh, tests/*.sh), matching bin/fm-lint.sh's canonical file set so parse scope and lint scope cannot drift apart. * no-mistakes(review): Captain: harden Bash structure and inventory guards * no-mistakes(document): Align stock macOS Bash contributor checks * no-mistakes(lint): Suppress deliberate SC2016 literal fixture warnings
* fix(test): pin teardown tmux baseline to historical kill selectors merge-base HEAD main collapses to HEAD after the exact-selector change lands on the default branch, so the old teardown fixture was accidentally exercising current exact targets. Resolve a content-historical permissive tmux adapter from first-parent history and force that post-squash topology inside the conformance case so main and feature branches keep the same old-vs-new contract. * no-mistakes(lint): Suppress intentional literal-pattern ShellCheck warnings
…id#1197) Cut the runtime skill to the compact pace-aware selection procedure plus minimum owner pointers. Keep every distinct decision rule and move expanded acceptance scenarios to deterministic fixture ownership assertions. Size: 170/1374/10187 -> 63/544/4068 (about 63%/60%/60% reduction).
…1219) * Inherit config/backend into secondmate homes with deliberate-override preservation Add backend to the shared inheritable config allowlist so launch, locked bootstrap, and config-push converge a primary pin into secondmate homes as each home local future-spawn default. Track last-inherited bytes in a private state provenance marker so deliberate per-home overrides survive present and absent primary convergence, keep --backend and FM_BACKEND stronger, and extend the existing inheritance tests plus docs and skill claims. * no-mistakes(review): Preserve equal unprovenanced backend overrides * no-mistakes(review): Preserve symlink overrides and verify spawn precedence * no-mistakes(review): Snapshot backend inheritance for consistent provenance * no-mistakes(review): Simplify backend inheritance to primary-authoritative convergence * no-mistakes(document): Document inherited backend override preservation * fix: restore primary-authoritative backend inheritance after document regression The document step reintroduced provenance and deliberate per-home override semantics after review had simplified config/backend to plain primary-authoritative allowlist membership. Restore the primary-always-wins path: present overwrites, absent removes, no provenance marker, and docs/tests match that contract. * no-mistakes(review): Add divergent backend precedence regression fixtures * no-mistakes(document): Document backend inheritance contract
* fix(pi): remove Calm's exclusive Pi upper-version ceiling
tests/fm-calm-pi-extension.test.sh gated on a closed PI_COMPAT_VERSIONS
allowlist ("0.81.1 0.82.0") that refused any other installed Pi, and docs
described that range as "supported" rather than verified evidence. The
Calm CHANGELOG shows no API introduced at either version, so there is no
evidence for a real minimum; the presentation adapters already probe the
exact method they patch rather than checking a version.
Replace the allowlist with dated version evidence that never rejects a
newer Pi, and make each presentation adapter degrade independently with
a diagnostic if a future Pi removes its API, instead of the whole Calm
extension failing to load. Rewrite the feasibility doc's "Pi 0.81.1
through 0.82.0" phrasing to state it as verified evidence, not a
ceiling.
* no-mistakes(review): Probe missing Calm adapter exports safely
* no-mistakes(document): Document Calm's unbounded Pi compatibility
…enguid#1204) * fix(guard): allow session-local todo tools in the primary The delegation-shape guard denied TaskCreate and TaskUpdate because their normalized names contain the `task` stem. Those tools write only the harness's session-local todo list, which has no executor: it spawns no agent, allocates no worktree, registers no schedule, and starts nothing that outlives the session. That is not the unaccounted work the guard exists to stop, so the stem match was a false positive, and the deny text told the primary to run bin/fm-brief.sh and bin/fm-spawn.sh to create a todo entry. Add a separately-reasoned PLAN_ONLY_TOOLS exact-name exclusion rather than widening OBSERVE_ONLY_TOOLS, whose documented contract is tools that only observe or stop existing work. Both lists stay exact-name so neither can widen by substring. Tests cover the two allowed names and six near-miss names that a substring or shortened-stem widening would release; both mutations were watched red. * no-mistakes(review): drop session-local todo tools from recommended deny list * no-mistakes: apply CI fixes
…claude pid (kunchenguid#1206) * fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid fm_harness_ancestry_pid() previously returned the first ancestor process whose command matched a verified harness name. Claude Code's Stop hook fires as a bg-spare worker several levels below the session's actual lock-owning claude process (hook shell -> claude bg-spare -> claude bg-pty-host -> claude -> claude(lock)), so the first match was the bg-spare worker, not the lock owner. fm_session_lock_owned_by_self() then never matched state/.lock, and the Claude Stop auto-arm silently treated its own primary session as an unrelated live owner and never armed the watcher. The walk now keeps going past a claude-named match, looking for a still more ancestral claude-named match, and stops the instant a non-match follows an already-found match (bounding it to a contiguous run rather than the literal ancestry top, so an unrelated claude-named process further up the real process tree is never mistaken for part of this session's own nested chain). Every other harness keeps the original first-match-wins behavior, since e.g. Pi's shared signed-wrapper ancestry actually holds the session at the inner engine pid, not an outer wrapper pid. Hop limit raised from 8 to 16 to cover the deeper bg-spare chain. * no-mistakes(review): Add nested-claude-ancestry regression test; fix nudge doc depth claim * no-mistakes: apply CI fixes
* fix: confirm watcher startup on MSYS * no-mistakes(review): gate MSYS arm ready timeout, cache uname, harden locale test * no-mistakes(review): validate OpenCode ready timeout, make uname cache internal
…d#1195) * fix(spawn): forward firstmate's CLAUDE_CONFIG_DIR to claude crewmates Crewmate panes are created by a long-lived tmux/herdr daemon that does not inherit firstmate's current environment. When firstmate runs under a non-default CLAUDE_CONFIG_DIR (for example a work-vs-personal subscription split), a bare `claude` in the crewmate pane fell back to the default ~/.claude store and launched unauthenticated, blocking the crewmate before it could do any work. fm-spawn now prefixes the claude launch with firstmate's own resolved CLAUDE_CONFIG_DIR when set, so the crewmate uses the same credential/config store firstmate is authenticated with. An unset value is the single-store default and adds no prefix; non-claude harnesses are unaffected. Adds three tests in fm-spawn-dispatch-profile.test.sh (forwarded-when-set, omitted-when-unset, non-claude-ignored) and pins CLAUDE_CONFIG_DIR in the test helper so launch assertions no longer depend on the developer's environment. * no-mistakes: apply CI fixes
…guid#1233) * fix: preserve dispatch harness identity * no-mistakes(review): Fix Grok counterfactual tuple validation * no-mistakes(document): Scope dispatch authentication to selected tuple * fix: restore dispatch instruction budget * no-mistakes(review): Scope dispatch authentication after candidate selection
* fix(bin): handle dash-leading harness process names (#2) * fix: handle dash-leading harness process names * no-mistakes(review): Make dash-leading harness regression hermetic * fix: preserve secondmate reply routes across relative homes Resolve relative home, data, and state inputs before durable charter generation, and fail when caller-relative directories cannot be resolved. Use absolute paths at the related spawn, AFK daemon, and X-mode cross-process handoffs so later processes cannot reinterpret them from another working directory. * no-mistakes(review): Preserve absolute overrides and normalize relative durable paths * no-mistakes(review): Normalize relative home before deriving durable paths * no-mistakes(document): Document relative durable-path normalization * no-mistakes(review): Captain: Ignore inherited CDPATH during relative path normalization * no-mistakes(lint): Fix empty CDPATH assignments for ShellCheck
* Add internal status skill * no-mistakes(document): register /status skill in documentation-audiences inventory * no-mistakes(lint): replace grep|wc -l with grep -c in status skill test * test: silence literal status skill patterns * Refactor bearings default to chat-only --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
* docs: add captain-approved project operation exception to hard rule 1 Firstmate stays read-only over projects by default, but when the captain clearly approves a concrete project operation and scope in the moment, firstmate may perform exactly that approved operation with its own tools. The approval is never inferred, broadened, or standing, and it does not relax the existing force, discard, unlanded-work, or merge-authority boundaries. * no-mistakes(review): Clarify captain-approved project operation boundaries * no-mistakes(document): Clarify captain-approved project operation scope * docs: cover directories and preserve the operation-or-scope alternative Widen the captain-approved project operation exception in AGENTS.md to files or directories, and restore the explicit operation-or-scope alternative that a prior pipeline auto-fix had collapsed into "and". Rework project-management SKILL.md's Remove section, which previously told firstmate to refuse project removal until a guarded helper existed; that helper was never built, so the text directly contradicted the new instruction-only exception. It now points at the exception plus the existing removal preflight it still requires unchanged. Update the one instruction-owners test assertion that hard-coded the sentence removed above, so the suite tracks current, not obsolete, text. * docs: add captain-approved project operation exception to hard rule 1 Firstmate stays read-only over projects by default, but when the captain clearly approves a concrete project operation and scope in the moment, firstmate may perform exactly that approved operation with its own tools. The approval is never inferred, broadened, or standing, and it does not relax the existing force, discard, unlanded-work, or merge-authority boundaries. * no-mistakes(review): Clarify captain-approved project operation boundaries * no-mistakes(document): Clarify captain-approved project operation scope * docs: cover directories and preserve the operation-or-scope alternative Widen the captain-approved project operation exception in AGENTS.md to files or directories, and restore the explicit operation-or-scope alternative that a prior pipeline auto-fix had collapsed into "and". Rework project-management SKILL.md's Remove section, which previously told firstmate to refuse project removal until a guarded helper existed; that helper was never built, so the text directly contradicted the new instruction-only exception. It now points at the exception plus the existing removal preflight it still requires unchanged. Update the one instruction-owners test assertion that hard-coded the sentence removed above, so the suite tracks current, not obsolete, text. * no-mistakes(review): Align project removal preflight with approved exception * no-mistakes(document): Align project removal documentation with approved exception * fix: restore removal test byte-for-byte and preserve the default sentence tests/fm-instruction-owners.test.sh had been changed to assert different text; restore it byte-for-byte to origin/main. project-management SKILL.md's Remove section now keeps the exact default "Never issue a raw removal command from Firstmate." sentence that test still asserts, immediately followed by the already-approved captain-operation-or-scope exception, so the default and the exception both stay explicit and consistent. * no-mistakes(document): Align project-write boundary documentation
…henguid#1275) * Route project intake through secondmate scopes * no-mistakes(test): Guard all main-home project registry mutations * no-mistakes(document): Consolidate secondmate routing documentation * no-mistakes: apply CI fixes * Restore new-project routing scope * no-mistakes(document): Clarify secondmate routing for new-project intake * no-mistakes: apply CI fixes
…#1282) * test: remove source-content assertions * no-mistakes(review): Replace source assertions with runtime behavior coverage * no-mistakes(review): Isolate Kimi task temp runtime coverage * no-mistakes(document): Refresh test cleanup documentation * no-mistakes: apply CI fixes
sbracewell64
force-pushed
the
fm/harness-interface-build
branch
from
July 30, 2026 02:27
ef9f166 to
f2089a8
Compare
…#1286) * fix(watch): bound how long a busy pane may run with no completed turn A busy pane (backend busy state or the harness's rendered footer) was unconditional, unbounded proof of liveness in every escalation path, so a hung foreground tool call behind a busy signature could run for hours undetected (2026-07 hibit-agent-focus-nonsteal-r1 incident: a catastrophic- backtracking regex hung one bash call for 25h behind an unchanging "Working..." footer). FM_BUSY_TURN_MAX_SECS (default 3600s) now bounds how long a busy pane may run with no completed turn (state/<id>.turn-ended, or its spawn record before any turn has completed). Past the bound, busy_turn_over_age routes the pane through the existing wedge_timer_check, reusing the identical stale reason, escalation counter, and demand-deep-inspection marker for human inspection only - never an automatic interrupt, signal, or restart of the worker or its tool process. A completed turn resets the age. Reproduced end-to-end against the real installed Pi TUI: a foreground `sleep 999999` bash call with no timeout renders the actual busy footer, and two captures ~15s apart show the elapsed counter changing the pane hash while the same turn stays unfinished. Running the pre-fix watcher against the real captures showed it never starts a wedge timer no matter how long the pane stays busy; the fixed watcher starts and escalates the timer through the same mechanism, while the real hung process remained untouched and alive throughout. * no-mistakes(review): fix: parse enriched AFK stale reasons * no-mistakes(review): fix: preserve enriched wedges during AFK supervision * no-mistakes(review): fix: route all enriched AFK wedges * no-mistakes(document): Clarify busy-turn age supervision documentation
…unchenguid#1261) A name-by-name list of config/ entries silently stops ignoring any new or home-local file placed there, which makes the working tree read as dirty and blocks guarded sync paths that refuse to touch a dirty home. AGENTS.md already documents config/ as captain-private and gitignored as a category; this makes .gitignore match that contract.
…al coverage (kunchenguid#1304) The second assertion in fm-gitignore-config.test.sh (added by kunchenguid#1261) greps .gitignore for a specific spelling of the config/ ignore pattern. It fails on a semantically equivalent pattern like config/** and does not prove Git actually ignores anything, per the completed source-content-test audit. Replace it with a real git check-ignore control test on a generated unrelated path, and strengthen the existing directory-coverage test with generated unpredictable direct and nested config/ paths.
* fix(herdr): place workers in the launching agent's exact workspace Herdr enforces no workspace-label uniqueness, and spawn resolved its container by taking the FIRST workspace whose label matched the home label. With two workspaces both labeled "firstmate", a worker launched from the second one was created in the first, so it appeared in a different space than the Firstmate the captain was watching. Reproduced end to end on Herdr 0.7.5 protocol 17 by running the real bin/fm-spawn.sh inside a launcher pane in the second "firstmate" workspace: the worker landed in w1 while its launcher was in w2, with an unrelated third workspace focused throughout, which also rules out any dependence on the focused workspace. Placement now binds to the launching process's own Herdr identity. Herdr injects HERDR_PANE_ID, HERDR_SESSION, and HERDR_SOCKET_PATH into every process it manages a pane for, and fm_backend_herdr_launcher_identity resolves that pane's current owning tab and workspace live from Herdr, cross-checking the pane against its tab and confirming the workspace exists exactly once in the session. The injected HERDR_TAB_ID and HERDR_WORKSPACE_ID are creation-time snapshots and are deliberately not read as current identity. Labels are no longer placement authority. A claimed parent identity that is unreadable, contradictory, stale, or from another named session or Herdr server stops the spawn before any worker endpoint exists, rather than degrading to a label search. A launcher with no Herdr ancestry has no workspace to inherit and keeps the per-home labeled container, which must now resolve to exactly one workspace; two same-labeled candidates refuse instead of adopting either. A --secondmate launch keeps standing up that home's own workspace by design. With presentation spaces enabled, the projected child is created and bound under that same exact parent and anchors its ordering on it, so a duplicated home label no longer makes the layout ambiguous. Projection, focus restoration, restart binding, and quarantine rules are unchanged, and children are never collapsed into the parent. tmux, Zellij, cmux, Orca, and the away-mode daemon terminal were each inspected and are not affected: none resolves a container by searching mutable labels. tests/fm-backend-herdr-launcher-workspace-e2e.test.sh drives the real spawn and teardown against an isolated Herdr lab, with its headline case running fm-spawn.sh inside a real Herdr pane so the identity comes from Herdr's own injection. The refusal matrix and the ordering anchor are covered deterministically in tests/fm-backend-herdr.test.sh. Eight existing real-Herdr suites inherited the developer terminal's own Herdr pane into their isolated lab sessions, which the new cross-session check correctly refuses. tests/herdr-test-safety.sh now owns herdr_forget_inherited_pane and those suites call it, so what they assert no longer depends on where they were launched from. Two unrelated fixes found along the way. tests/fm-secondmate-harness.test.sh had the same class of environment leak through CLAUDECODE, which outranks PI_CODING_AGENT in bin/fm-harness.sh and made its pi-signed ancestry case resolve "claude" whenever the suite ran inside Claude Code. And fm-spawn.sh's usage() printed a fixed line range that had already been truncating its own help mid-sentence. * no-mistakes(review): Enforce exact Herdr launcher and projection identity * no-mistakes(document): Document exact Herdr launcher workspace placement
* feat(calm): replace Pi's working row with an animated ship while Calm is on While Calm is active and one logical agent run is under way, Calm now hides Pi's built-in working row and renders a small two-row SSHHIP-derived boat in its place. When Calm is off, Pi's stock working row is left untouched. The presentation uses only public Pi extension API: setWorkingVisible(false) plus a temporary setWidget() component whose render(width) owns the responsive geometry and whose timer requests a TUI render. Visibility follows agent_start through agent_settled, so the boat does not flicker between tool calls, automatic continuations, retries, or compaction inside the same run, and settle, abort, and failure all reach the same cleanup. fm-calm.ts stays the sole owner of the presentation choice and the only caller of setWorkingVisible(); the new lib owns the sprite geometry and widget. * no-mistakes(review): Guarded Calm-off lifecycle visibility writes; focused tests pass * no-mistakes(test): Fixed Calm E2E wait to include tmux scrollback * no-mistakes(document): Document Calm working boat behavior * no-mistakes: apply CI fixes * feat(calm): slow the Calm boat, animate blue water, and make the sail directional The boat now moves one column every 880ms while a bounded fixed-cell water phase advances every 220ms, so the water ripples several times between boat steps and the presentation reads as calm. One scheduler drives both clocks and disposing the widget stops them together; ticks rather than wall-clock timestamps drive every state change, so tests seek animation time exactly. Colors are standard ANSI foreground codes instead of theme lookups: blue for every water cell and yellow for the complete boat, each run closed with a default-foreground reset so nothing bleeds into padding or later frames. ANSI bytes never enter geometry, so visible width stays exact. The mainsail is directional and trails aft of the mast: <| travelling right and |> travelling left. Direction reverses the moment the boat lands on an endpoint, so the endpoint frame already shows the new heading and no frame at or after a bounce shows the previous sail. * test(calm): wait for the Ctrl+O expansion redraw this block asserts * docs(calm): record the revised working-presentation verification evidence * no-mistakes(document): Fix Calm feasibility document EOF whitespace
…henguid#1349) * fix(dispatch): scope candidate authentication to its own surface A locally expired timestamp in one credential store was reported to the captain as a sign-out, including for dispatch candidates that never read that store. A `harness=pi, model=xai/grok-*` candidate authenticates through Pi's own xAI credential, but the only Grok quota reading available was gated on the standalone Grok CLI's separate token, whose expiry clock drifts independently. The always-loaded intake rule then turned that unreadable quota into a mandatory captain escalation. Add `bin/fm-auth-preflight.sh` as the deterministic owner of the parts that must not depend on agent memory: it resolves a tuple's authentication surface from quota-axi's own emitted auth sources rather than from a harness or model name, so another harness's CLI can never gate a candidate that does not use it. A vendor CLI is launched only when the tuple's own harness owns the credential store under test and a non-destructive discovery command is registered for it, which today is `grok models` alone. That probe runs at most once with stdin closed and a hard timeout, reads its verdict from the first stdout line because the command exits 0 either way, treats unrecognized output as indeterminate, and never invokes login, logout, or the interactive TUI. Quota is read at most twice, and unknown headroom never makes a candidate ineligible on its own. Update the dispatch procedure to match: usable authentication with unmeasurable headroom stays eligible at lower preference with the unknown disclosed, and stop-and-report is reserved for unresolved authentication, an unresolved relationship, or malformed configuration. Record that Grok's `credits.remaining` is a prepaid balance rather than window headroom. Gate quota-axi at 0.1.16 in bootstrap, the first build reporting per-credential auth sources. A stale install previously passed the presence check silently, which is why a fix published two days earlier was still not in effect. Replace the orphaned quota-array-dispatch fixtures, which encoded a `provider: "xai"` shape the tool never emits and had no consumer, with fixtures shaped like real 0.1.16 output that the new suite drives the script against. The suite asserts the verdict and, separately, which vendor CLIs were launched, so a Pi/xAI candidate reaching the Grok CLI fails. Map `tests/fixtures/<dir>` to its consuming suite so a fixture change selects the right tests instead of refusing. * refactor(bootstrap): give the quota-axi floor one owner The floor was stated twice - once in bootstrap's gate and once inline in the auth preflight - so bumping it needed two edits that could drift. Move it to bin/fm-quota-axi-lib.sh alongside its rationale, matching the existing tasks-axi library, and derive the comparison from the constant so the number appears exactly once. Bootstrap turns a failing check into the operator diagnostic; the preflight refuses to emit an unscoped verdict. Map the new library to both consuming suites so a bump re-runs them, and record that any usable source means the surface authenticates. * no-mistakes(review): Captain: bound quota checks and removed Python dependency * no-mistakes(review): Captain: enforce conservative headroom and exact preflight retry * no-mistakes(review): Captain: preserve OpenCode eligibility without auth-surface guessing * no-mistakes(review): Captain: reject malformed OpenCode model relationships * no-mistakes(review): Captain: exempt verified unmodeled tuples from intake escalation * no-mistakes(document): Updated dispatch authentication documentation * no-mistakes: apply CI fixes
…nchenguid#1350) * feat(x-mode): reconcile promised public replies deterministically A promised final reply in an X or Discord thread was only kept while the primary remembered it. Compaction or restart erased that memory, so a typed public-followup obligation could sit at pending-work after its PR merged and the original thread never got its reply. Make the promise durable state instead: - bin/fm-public-followup-emit.sh reports a typed terminal work result (source home, work id, generation, outcome, safe deliverables, bounded public-safe text) into the owning home's private inbox. The event id is derived from that identity tuple, so duplicate reports and restart replay converge with no coordination, and nothing ever parses a free-form done: sentence. - bin/fm-public-followup.sh registers a commitment, reconciles events through tasks-axi public-followup, and runs the idempotent delivery sequence (begin-delivery with the payload hash, post, record the posted receipt or a typed error) against the stored platform and opaque thread binding. A delivery interrupted between post and receipt refuses rather than risk a second public reply. - Session start surfaces unresolved commitments from disk, the existing relay poll surfaces a new terminal-result set once, and teardown refuses while this home still owes a public reply for that exact work. tasks-axi public-followup remains the only owner of the obligation state machine, state/x-context/ the only owner of the private request context, and fm-x-reply.sh the only thing that posts. Its new optional --receipt-file is the one addition there, so a caller can record how many messages were sent. A home that never opted into the myfirstmate relay gates out on a single [ -f "$FM_HOME/.env" ] test: no tasks-axi call, no backlog or context scan, no output, and no artifact. Evidence in docs/verification/public-followup.md. * no-mistakes(review): Hardened public-followup reconciliation and ownership guards * no-mistakes(review): Hardened typed terminal cleanup and receipt reconciliation * no-mistakes(review): Automated typed-delivery cleanup and strict backlog validation * no-mistakes(review): Fail-closed parent resolution and registration-safe delivery * no-mistakes(review): Harden relay gating and validate secondmate bindings * no-mistakes(review): Use owner-aware single-gate teardown protection * no-mistakes(document): Correct public-followup documentation drift * no-mistakes(lint): Quote done literals to fix ShellCheck warnings * no-mistakes: apply CI fixes
sbracewell64
force-pushed
the
fm/harness-interface-build
branch
from
July 31, 2026 02:42
f2089a8 to
6a2952d
Compare
…chenguid#1327) * feat: add semantic busy-state contract owner and event writer One owner (bin/fm-busy-lib.sh) for the captain-approved semantic busy-state redesign: a per-task gen-bound record written only by bin/fm-busy-event.sh, per-harness trusted-source classification with explicit source attribution, busy/idle/unknown/dead semantics where missing, malformed, stale, or untrusted semantic data is unknown - never idle - and endpoint death is the only process-level override. The Grok-only rendered-tail fallback and the standalone-Kimi verification gate live behind the same classifier. * feat: arm busy-state at spawn and convert Pi to the semantic extension path fm-spawn arms the busy-state contract for converted adapters and seeds busy/fm-spawn (the launch brief is a submitted turn). The Pi/pi-signed per-task extension now reports agent_start -> busy and agent_settled -> idle confirmed by ctx.isIdle(), covering auto-retries, compaction retries, tool loops, and queued continuations, while turn_end stays a wake notification touch. Teardown removes the new record, gen sidecar, and lock. Live-verified on Pi 0.82.0: seed -> agent-start busy -> agent-settled idle with the marker still touched. * feat: convert OpenCode to the semantic session.status plugin path The per-task plugin (renamed .opencode/plugins/fm-busy-state.js) now classifies from OpenCode's semantic session.status events - busy and retry are active, idle is inactive - latched to the worker's own session so a subagent child session can never clear the worker's busy state. The session.idle marker touch stays a wake notification. Teardown removes both the new and the legacy plugin filenames. Live-verified on OpenCode 1.17.18 in a real TUI pane: seed -> session-busy -> session-status-idle. * feat: convert Claude to the full lifecycle hooks path The per-task settings.local.json now wires UserPromptSubmit -> busy and Stop, StopFailure, and SessionEnd -> idle, so API-error and shutdown turn ends can never strand a busy record; Stop keeps the turn-ended notification touch. A refused (stale-gen) event exits 0 and stays silent so Claude's own lifecycle is never broken. Live-verified on Claude Code 2.1.220: UserPromptSubmit fires for the argv launch prompt, Stop closes each turn, a mid-stream Escape interrupt fires no closing hook, and the firstmate-controlled idle/fm-interrupt clear resolves it. * feat: gate Codex busy state behind verified semantic sources The approved contract prefers Codex's app-server turn lifecycle with capability negotiation and sanctions its lifecycle hooks as the intermediate. Live probes on codex-cli 0.145.0 show neither is usable for a pane worker: the app-server daemon is unreachable for a TUI thread and refuses to start outside the managed standalone install, and firstmate-written project hooks never fired (interactive with directory trust granted, and exec, both with --dangerously-bypass-hook-trust) while global hooks fired in the same runs. Codex therefore classifies unknown codex-unverified behind an explicit probe rather than falling back to idle or footer text, and fm-spawn installs no unverified Codex wiring. * feat: gate standalone Kimi busy state on live verification Standalone Kimi has no installed binary here, so per the approved contract its semantic path stays guarded and it classifies unknown kimi-unverified rather than idle - and never from its locale-sensitive moon-phase spinner, which the redesign forbids inventing as a state source. The gate records the preferred source order (Wire prompt request lifetime, which brackets a turn and reports cancellation, then the documented hooks including Interrupt because Stop does not fire on interrupts) and the exact evidence required to open it. Arming without wiring would seed a busy record nothing could clear, so both land together behind the same gate. * feat: route busy consumers through the contract and drop the global OR The watcher, crew-state reader, and away-mode daemon now decide busy state through bin/fm-busy-lib.sh: only an exact busy verdict counts as working, and unknown never becomes working or a silent idle, so a crew whose semantic state is missing, malformed, stale, or unverified surfaces instead of being absorbed. Crew-state reports the producing source in its detail. The watcher's global OR regex default is gone; Grok keeps its isolated fallback inside the contract. The daemon's supervisor-pane reader stays rendered-text - that pane is not a recorded task - but is now scoped to firstmate's own detected harness instead of every vendor signature. Secondmate pending-reply observation is deliberately unchanged and documented as a delivery-confirmation signal, not task state. * docs: point busy-state documentation at the single contract owner Adds a maintainer-architecture section naming bin/fm-busy-lib.sh as the owner of what busy means, with per-adapter sources, the unknown-never-idle rule, the endpoint-death override, and the two rendered-text readers that deliberately stay outside the contract. Replaces the stale regex-first prose in architecture, tmux-backend, herdr-backend, and configuration; converts the harness-adapters per-harness rows from UI signatures to the semantic source each harness uses; and records the live verification evidence, including why Codex and standalone Kimi stay unknown. * fix: arm away-launch signal handlers before acquiring the lifecycle lock fm_afk_launch_main acquired its lock and only then installed the EXIT, INT, and TERM traps. A signal arriving in that window terminated the process by default action and left the lock directory behind, which blocks the next away-mode launch until the stale-owner reclaim path clears it. The release helper only removes a lock this process owns, so the handlers are now armed first. The accompanying test also killed the child whether or not the lock had appeared and sampled cleanup the instant wait returned; it now requires the lock, then allows a bounded settle, so it proves the guarantee instead of racing it. * test: align fleet, Kimi, lifecycle, and detection suites with the contract The fleet snapshot and wake-daemon lifecycle fixtures now prove a working crew through its own semantic busy-state record instead of rendered pane text, which is what those consumers read. The Kimi watcher test asserts the approved contract directly: a standalone Kimi task classifies unknown rather than matching its moon-phase spinner, while Grok's isolated fallback still classifies only Grok. The pi-signed detection cases clear ambient harness markers, fixing a pre-existing failure where the running session's own CLAUDECODE outranked the fixture's marker. * fix: stop teardown from deleting a project's own Codex hooks file An intermediate revision wired Codex through a firstmate-written <worktree>/.codex/hooks.json, and teardown removed it alongside the other generated wiring. The Codex wiring was dropped when its probes came back unverified, so that removal now targets a file firstmate never creates - and a project may legitimately track its own .codex/hooks.json, which teardown would then delete from a pooled worktree. * fix: keep busy-record parsing from disturbing its sourcing caller The record parser split fields with set -- under a temporary noglob, which clobbers a sourcing caller's positional parameters and restores glob expansion even when the caller had disabled it. The watcher, the daemon, and the crew-state reader all source this library, so it now reads fields with read -a, which never globs and never touches caller state. * docs: state exactly which Claude hook paths were reproduced live The busy-state record listed all four wired Claude hooks in the source column, which could read as a claim that every one fired during the pass. UserPromptSubmit and Stop did; StopFailure and SessionEnd are wired from hook names confirmed present in the installed binary, but the abnormal turn ends they cover were not reproduced. * test: let reset_fakes own the crew-state busy-text fixture lifecycle The Grok fallback case set FM_FAKE_BUSY_TEXT and cleared it inline, so the variable's lifetime was owned by one test rather than by the shared reset that every other fake already uses. * no-mistakes(review): Fix semantic busy-state lifecycle races * no-mistakes(review): Make busy-state retirement idempotent * no-mistakes(review): Enforce semantic state boundaries for status and injection * no-mistakes(review): Restore harness-scoped away-mode busy guard * no-mistakes(document): Refresh semantic busy-state documentation * no-mistakes: apply CI fixes
Extract the per-harness busy signature that bin/fm-tmux-lib.sh owned into a harness-axis adapter layer - bin/fm-harness-adapter.sh plus bin/harnesses/<name>.sh - with ZERO behavior change. The busy table is the most-consumed harness fact in the system: 8 call sites across 5 files. It lived in a BACKEND-named library, so bin/fm-crew-state.sh, bin/fm-supervise-daemon.sh, and bin/fm-pending-reply-lib.sh each had to source a tmux library purely to read harness knowledge, and bin/fm-crew-state.sh routes non-tmux backends through it explicitly. bin/fm-watch.sh reached it transitively through a sibling. After this change bin/fm-pending-reply-lib.sh and bin/fm-watch.sh need no tmux library at all, and bin/fm-tmux-lib.sh keeps only the tmux pane probe that consumes the signature. The dispatcher mirrors bin/fm-backend.sh: a known-name registry, per-adapter files that are each their own lint root, and thin per-op dispatch. It keeps TWO deliberately different name sets - FM_HARNESS_KNOWN for crewmate and secondmate launches, and the narrower FM_HARNESS_PRIMARY for the primary session. kimi is in the first and not the second, which is why there is no docs/supervision-protocols/kimi.md and no kimi arm in bin/fm-supervision-instructions.sh; collapsing the sets would silently promote kimi to a primary harness. Resolving the pi-signed -> pi alias once in the dispatcher retires 15 scattered `pi|pi-signed` case arms. Two deliberate departures from the backend template, both because harness adapters are pure data rather than command sequences. Adapters are loaded eagerly at source time, not lazily per call: the matcher reads stdin, so it always runs as a pipeline subshell and a lazy source could never populate a cross-call guard - it would re-read the adapter on every busy check. And each signature is a constant rather than a function, so the hot path resolves it with a variable read instead of a fork. Measured over 200 checks, the result is at parity with the pre-refactor code (1526us vs 1576us per check); the intermediate lazy-plus-command-substitution shape was 2x slower. Byte-identical behavior is proven, not asserted. tests/fm-harness-adapter.test.sh runs the PRE-refactor fm_busy_lines_match, extracted from the merge-base with main, and the refactored fm_harness_busy_match against the same capture corpus in separate processes, then diffs the verdict logs: 108 harness/capture pairs, byte-identical. It also diffs the resolved regex string for every harness name, the empty name, and an unregistered name. tests/lib.sh gains fm_test_base_ref as the single owner of old-vs-new baseline resolution, shared with tests/fm-backend.test.sh. It now picks the NEWEST merge-base among candidate refs rather than the first that resolves: a disposable task worktree is routinely cut from a pool whose local main is many commits stale, and pinning the baseline there made unrelated landed changes read as behavior differences. Contract docs updated in the same change: AGENTS.md section 4, CONTRIBUTING.md (whose harness-adapter bullet named the old busy-signature location), .agents/skills/harness-adapters/SKILL.md, bin/fm-lint.sh's roots and its canonical-file-set test, and bin/fm-test-run.sh's family map.
sbracewell64
force-pushed
the
fm/harness-interface-build
branch
from
July 31, 2026 04:30
6a2952d to
f3c2308
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Ship phase 3 of the firstmate harness-interface project: extract a harness adapter interface (bin/fm-harness-adapter.sh + bin/harnesses/.sh) around the per-harness busy signature that bin/fm-tmux-lib.sh used to own, with ZERO behavior change. This is PR A of an approved five-PR, method-sliced staging (A: busy signature; B: turn-end install/cleanup; C: process identity; D: launch command and flags; E: fold per-harness supervision prose into docs/supervision-protocols). Method-slicing was chosen over harness-slicing specifically to avoid a temporary dual code path where per-harness knowledge would live in two shapes across two PRs.
Deliberate framing: this imitates the maintainer's own adapter-layer template in upstream PR kunchenguid#183 (feat: a new interface, a byte-identical-behavior guarantee stated in the PR body, old-versus-new conformance tests as the evidence, and every contract doc updated in the same PR). Byte-identical behavior is a HARD GATE here, not an aspiration: no behavior change is intended anywhere in this diff.
Motivation (measured, from a prior architecture review): the busy table is the most-consumed harness fact in the system - 8 call sites across 5 files - and it lived in a BACKEND-named library. bin/fm-crew-state.sh, bin/fm-supervise-daemon.sh, and bin/fm-pending-reply-lib.sh each had to source a tmux library purely to read harness knowledge, and bin/fm-crew-state.sh routes NON-tmux backends through it explicitly. bin/fm-watch.sh reached it transitively through a sibling library. After this change bin/fm-pending-reply-lib.sh and bin/fm-watch.sh need no tmux library at all.
Decisions a reviewer reading only the diff would not know:
The registry deliberately keeps TWO different harness name sets. FM_HARNESS_KNOWN (7 names) is launchable as a crewmate/secondmate per docs/configuration.md; FM_HARNESS_PRIMARY (6 names) is supported for the primary session per README.md Requirements. kimi is deliberately in the first and NOT the second. That is exactly why there is no docs/supervision-protocols/kimi.md and no kimi arm in bin/fm-supervision-instructions.sh - those absences are correct, not drift. Do NOT suggest merging the two lists or adding a kimi arm: it would silently promote kimi to a primary harness and contradict what README advertises.
Two deliberate departures from the bin/fm-backend.sh template, both because harness adapters are pure data rather than command sequences. (a) Adapters load EAGERLY at source time, not lazily per call: fm_harness_busy_match reads stdin, so it always runs as a pipeline subshell, and a lazy source could never populate a cross-call guard - it would re-read the adapter on every busy check in the watcher poll loop. (b) Each signature is a CONSTANT, not a function, so the hot path resolves it with a variable read instead of a fork. Measured over 200 checks: this is at parity with pre-refactor code (1526us vs 1576us per check), while an intermediate lazy-plus-command-substitution shape measured 2x slower. The eager-load rationale is documented at the foot of the dispatcher.
The SC2034 'appears unused' disables in bin/harnesses/*.sh are deliberate: each adapter is linted as its own canonical root, so its consumer (the dispatcher) is out of shellcheck's scope. bin/fm-lint.sh's ROOTS and tests/fm-lint.test.sh's canonical file set were both extended for the new directory.
tests/lib.sh gains fm_test_base_ref as the single owner of old-vs-new baseline resolution, shared with the pre-existing tests/fm-backend.test.sh (whose private copy was removed). It now picks the NEWEST merge-base among candidate refs rather than the first that resolves. This is a real bug fix, not cosmetics: a disposable task worktree is routinely cut from a pool whose local main is many commits stale, and pinning the baseline there made unrelated already-landed changes read as behavior differences. Verified that tests/fm-backend.test.sh still passes with the corrected, newer baseline.
tests/fm-kimi-harness.test.sh's two location assertions were retargeted from grepping bin/fm-tmux-lib.sh's source text to asserting on the RESOLVED regex, so they follow the fact instead of its address. Its behavioral assertions were not changed.
Deliberately OUT of scope, do not flag as missing:
Evidence: tests/fm-harness-adapter.test.sh runs the PRE-refactor fm_busy_lines_match (extracted from the merge-base with main) and the refactored fm_harness_busy_match against the same capture corpus in separate processes, then diffs the verdict logs byte-for-byte - 108 harness/capture pairs identical - plus a byte-identical diff of the resolved regex string for every harness name, the empty name, and an unregistered name, and safety assertions that an unregistered harness never borrows another's signature and FM_BUSY_REGEX still overrides everything. A real-tmux end-to-end run on an isolated socket confirmed the capture-to-adapter-to-verdict path.
Note on the local test suite: seven test files fail on this machine (fm-backend-tmux-smoke, fm-calm-pi-extension, fm-watcher-lock, fm-pi-watch-extension, fm-turnend-guard, fm-secondmate-harness, fm-session-start). Each was verified to fail IDENTICALLY on a clean checkout of the base commit c0c0881 with none of these changes present - they are environmental to this machine (WSL2 tmux, node ESM, watcher beacon timing, and an ancestry test that detects the real claude process the session runs inside), not regressions from this work.
What Changed
bin/fm-harness-adapter.shplus per-harness data adapters underbin/harnesses/(claude, codex, grok, kimi, opencode, pi) — that now owns the per-harness busy signature previously hard-coded inbin/fm-tmux-lib.sh. All 8 call sites acrossbin/fm-crew-state.sh,bin/fm-supervise-daemon.sh,bin/fm-pending-reply-lib.sh,bin/fm-watch.sh, andbin/fm-tmux-lib.shresolve busy state through the adapter, sofm-pending-reply-lib.shandfm-watch.shno longer source any tmux library. Adapters load eagerly as constants (parity with pre-refactor timing), and the registry keeps the crewmate-launchable (FM_HARNESS_KNOWN) and primary-session (FM_HARNESS_PRIMARY) name sets deliberately separate.tests/fm-harness-adapter.test.shruns the pre-refactor matcher (extracted from the merge-base) and the newfm_harness_busy_matchagainst the same capture corpus and diffs verdicts byte-for-byte across 108 harness/capture pairs, plus resolved-regex diffs per harness and safety checks that unregistered harnesses borrow nothing andFM_BUSY_REGEXstill overrides.tests/lib.shgains a sharedfm_test_base_refbaseline resolver that picks the newest merge-base (fixing false diffs in worktrees cut from a stale local main), and the kimi test's location assertions now target the resolved regex instead of source-text greps.bin/fm-lint.shROOTS,tests/fm-lint.test.sh), the gotmp test fixtures, and docs (CONTRIBUTING.md,docs/configuration.md,docs/scripts.md,docs/tmux-backend.md, agent skills) were extended for the newbin/harnesses/directory; the review pipeline caught and auto-fixed a missed gotmp fixture symlink and a stale shellcheck-inventory doc line.Risk Assessment
✅ Low: The fix commit resolves both round-1 findings exactly as instructed (verified by re-running the fixture sourcing repro, which now loads fully under set -eu), and the underlying branch remains the deliberately mechanical, conformance-tested byte-identical extraction already verified in round 1.
Testing
Ran the four touched test suites (harness-adapter conformance, backend, kimi-harness, gotmp) — all pass, including the byte-identical gate diffing 108 pre-refactor vs refactored busy verdicts and every resolved regex string against the merge-base — and performed a manual real-tmux end-to-end check on an isolated socket proving the capture-to-adapter-to-verdict path yields correct busy/idle classifications; tests/fm-lint.test.sh was left to the lint step per the no-linters rule, and the working tree was left clean.
Evidence: Harness-adapter conformance test transcript (byte-identical gate)
ok - harness registry: the crewmate set and the narrower primary set stay distinct ok - harness registry: pi-signed resolves to pi once, so call sites drop their own alias arms ok - harness registry: every verified harness has a sourceable adapter with a non-empty signature ok - busy signatures: every harness's regex string is byte-identical to the pre-refactor constant ok - busy verdicts: 108 harness/capture pairs classify identically old vs new ok - safety: an unregistered harness is never classified busy by borrowing a verified signature ok - safety: claude's and kimi's scoped signatures stay out of the shared fallback ok - safety: FM_BUSY_REGEX still overrides every per-harness signature all fm-harness-adapter tests passedEvidence: Real-tmux end-to-end: capture-pane → adapter → verdict
=== real-tmux end-to-end: capture-pane -> fm-harness-adapter -> busy/idle verdict === tmux tmux 3.6, isolated socket: fm-adapter-e2e-36202 --- pane capture (claude busy spinner on screen) --- 1:✻ Thinking… (12s · esc to interrupt) verdict harness='claude' -> busy verdict harness='EMPTY' -> busy verdict harness='codex' -> busy --- pane capture (idle composer prompt on screen) --- 1:│ > │ verdict harness='claude' -> idle verdict harness='EMPTY' -> idle === done ===Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 2 issues found → auto-fixed ✅
tests/fm-gotmp.test.sh:60- tests/fm-gotmp.test.sh's fake bin/ trees (built at lines 53-61 and repeated at 158-162) symlink the real bin/fm-tmux-lib.sh and bin/fm-composer-lib.sh but not the new bin/fm-harness-adapter.sh or bin/harnesses/. fm-tmux-lib.sh now sources fm-harness-adapter.sh via dirname of BASH_SOURCE, which resolves to the fake tree, so the source fails (reproduced: 'No such file or directory', exit 1 under set -eu). The teardown tests currently pass only because every tmux-backend dispatch on their path is '|| true'-guarded, which suspends errexit and lets the library half-load with fm_harness_busy_match undefined and FM_HARNESS_BUSY_REGEX_DEFAULT unset — the fail-open state the adapter's fail-closed load guard exists to prevent. tests/fm-backend.test.sh's equivalent fixture was updated for the new sibling (copies fm-harness-adapter.sh and bin/harnesses/); this fixture was missed. Fix: symlink fm-harness-adapter.sh and the harnesses/ directory into both fake trees..agents/skills/firstmate-coding-guidelines/SKILL.md:97- The shellcheck inventory prose in .agents/skills/firstmate-coding-guidelines/SKILL.md line 97 still reads "bin/*.shandbin/backends/*.shmust passshellcheck" — this PR extended the canonical lint set with bin/harnesses/.sh (bin/fm-lint.sh ROOTS, tests/fm-lint.test.sh CANON, and CONTRIBUTING.md were all updated), so this doc line is now stale. Add bin/harnesses/.sh to it.🔧 Fix: link harness adapter into gotmp fixtures; update lint docs
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-harness-adapter.test.sh— registry name sets, pi-signed alias, byte-identical regex strings and 108 old-vs-new verdict pairs, unregistered-harness and FM_BUSY_REGEX safety propertiesbash tests/fm-backend.test.sh— backend conformance suite now consuming the shared fm_test_base_ref baseline resolverbash tests/fm-kimi-harness.test.sh— kimi suite with location assertions retargeted to the resolved regexbash tests/fm-gotmp.test.sh— gotmp fixtures with the adapter linked inManual real-tmux end-to-end on isolated socket: rendered claude busy spinner and idle composer prompt in a live pane, piped capture-pane output through fm_harness_busy_match, verified busy/idle verdictsSanity-checked that bin/fm-pending-reply-lib.sh and bin/fm-watch.sh no longer source a tmux library for busy matching✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.