Skip to content

fix(live-proof): provide cold-checkout prerequisites to planning - #1264

Merged
steipete merged 1 commit into
mainfrom
steipete/terminal-plan-prerequisites-20260827
Aug 27, 2026
Merged

fix(live-proof): provide cold-checkout prerequisites to planning#1264
steipete merged 1 commit into
mainfrom
steipete/terminal-plan-prerequisites-20260827

Conversation

@steipete

@steipete steipete commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Summary

Fix the missing prerequisite context behind the independently generated terminal-plan failure observed during PR1259, without changing that PR's landed command-preservation behavior.

The producer already had general build-sequencing guidance, but planning did not receive the effective live_test execution contract. The proof target is a separate cold checkout; the configured installation suppresses lifecycle scripts and does not inherit the controller's dist. The retained plan directly loaded two tests that import dist/clawsweeper.js, so both imports failed before the expectations could succeed. The full historical model read trace is unavailable, so this does not claim the model never encountered build information.

This patch supplies trusted effective profile context to the production prompt (including fallback resolution, actual normalized setup, install-script policy, cold-head checkout, and browser-only startup), strengthens prerequisite guidance, and shares the existing unchanged setup normalizer with execution. It does not guess or automatically repair generated commands.

No public decision schema, repository setup defaults, parser/driver command multiplicity, assertions, timeouts, or publication contracts change. Every declared entry/run still executes, including intentional repeats. OpenClaw Bay is unaffected: no lifecycle, status, public data-contract, or observer-surface change.

Rebase and review disposition

Current tested head: 213927808a1d14c11fb8507e45dcc061ceeade62
Base: 2129a78a502e4e6ed3dd0c521db5034743f1757a

The independently owned PR1261 landed while this PR's prior review was being published. This branch is now rebased onto it. The sole conflict was in shared imports and adjacent additions in test/review-prompt-context.test.ts; both owners' test blocks were retained. Runtime and changelog changes merged automatically. No additional runtime rewrite was needed.

A subsequent main advance, PR1265, required only changelog conflict resolution. All incoming notes were preserved and our exact entry was retained at the end of Changed to avoid repeated top-note collisions. Runtime/prompt changes merged automatically; proof was refreshed again for the final head above.

The earlier changelog finding was a target-scope false positive and was explicitly cleared by the review of the prior head, which rated proof, patch, and overall readiness A with no findings or rank-up work. This targets openclaw/clawsweeper, not openclaw/openclaw; the intended ClawSweeper Unreleased entry is retained. AGENTS.md explicitly scopes release-owned changelog handling to openclaw/openclaw, and the landed policy fix further clarifies target ownership.

That earlier review/CI is history, not approval of the rewritten head. The conflict-resolved candidate and committed range received fresh independent Codex reviews, and the actual generation/cold-checkout proof below was rerun for this exact head. Current-head CI and ClawSweeper review remain required landing gates.

Current-head controlled real-behavior proof

Claim: a real generated plan supplies the build prerequisite missing from configured setup, then executes the original diagnostic command from a cold current-head checkout without hidden prebuilds or changed execution semantics.

The exercised path was production buildReviewPrompt → actual constrained Codex generation using the full unchanged production schema → production decision parser → production executeLiveProof setup → real tmux. The request fixed the original two test files, test-name pattern, and expectations, supplied their real current-head source context, and prohibited tools. It did not instruct the generator to build or supply an expected plan. The inner tests are the selected diagnostic fixture; the runtime proof is the outer generation/setup/execution path, not unit-test success alone.

Generation completed at 2026-08-27T15:31:48.912Z. The unchanged scripts/e2e/terminal-proof-generate.mjs attested restricted read-only execution, denied outside reads and workspace writes, no writable paths, disabled child networking, and zero tool calls. When the host CLI changed, its strict version check correctly refused the new version before inference. The original official 0.150.0-alpha.13 CLI was then installed in task-private storage; no attestation, tool inventory, permission, or host configuration was weakened. No raw transcript or private model identifier is included.

The actual, unedited generated entry was:

pnpm run build && node --test --test-name-pattern='preserves every terminal command including exact entry repeats|terminal entry with expectations executes once' test/decision-parser.test.ts test/live-proof.test.ts

Its only steps were the two original expect_output assertions. Production parsing preserved the plan unchanged. Replay verified current source/compiled/target hashes and an exact command allowance before execution. The controller was built separately; each target was an independent cold clone at the tested head.

Observation Retained no-build plan Fresh generated plan
Initial dependencies / dist absent / absent absent / absent
After production pnpm install --ignore-scripts --frozen-lockfile dependencies present, dist absent dependencies present, dist absent
Declared / actual entry commands 1 / 1 1 / 1
Diagnostic test invocations 1 1
Actual terminal exit 1 0
Missing dist/clawsweeper.js imports both selected files none
Expectations not run after failure both observed, detail: ok
Overall verification failed passed
Compiled output after execution absent present
Dedicated tmux cleanup passed passed

Positive verification is bound to 213927808a1d14c11fb8507e45dcc061ceeade62 and completed at 2026-08-27T15:33:55.453Z:

$ tsc -p tsconfig.json
✔ decision parser preserves every terminal command including exact entry repeats
✔ terminal entry with expectations executes once and preserves long final output publicly
tests 2; pass 2; fail 0
both expectations: satisfied=true, detail=ok, present_at_start=false
drive_status=completed; overall_pass=true

This does not rely on the schema-v1 NOT OBSERVED successful-exit fallback. The negative replay fails during execution rather than the historical producer's expectation timeout; it proves the same prerequisite defect without claiming to explain historical status sampling.

Provider/environment: local macOS real tmux; proof IDs cold-negative-final-head and cold-positive-final-head; Node controller 24.20.0, target Node 26.7.0 on the allowlisted PATH; pinned pnpm 11.10.0; tmux 3.7c. Both runs used separate HOME/cache/tmp, the exact 17-key noncredential environment allowlist, and a dedicated task-local socket. Target isolation facts are operator telemetry, not kernel containment. Local --checkout semantics deliberately bypass live GitHub lookup. Fixture item number 1259 is only a diagnostic identifier; the actual controller/checkout head is this PR's tested head above.

Integrity/provenance:

production prompt SHA-256: 880db767ee21afd673ed724afd7d7eadf79d1ce072aca2a4d0a1b28184e87b84
production schema SHA-256: 758dd96182b5db7b9b4cb98164b4dbb9efd30d3cb7baf7590ad2159ef7b1aad9
full generated decision SHA-256: 3d81535811705a092a694adb1004b1822e2573fbdfc92aad513b51f91159e340
generator helper SHA-256: 2b12954b5afe14cb2d451a5457342ca1ad5f29703b2deb5426f162f80692f41e
replay helper SHA-256: 5f841c4ed29bb8b9347b21ce6411bdccf8c5190f881dfa57efae3d78de20292c
src/clawsweeper-review-runtime.ts SHA-256: 112f42271ca3cde415bc22192e0d40670d5a896f0e1c4b523c296c0b523d3985
src/repository-profiles.ts SHA-256: 22c74f749bcd136572fee4786e6e5a69ebb552ed2cf102818e91e8a7fb888604
prompts/review-item.md SHA-256: b0da4e8e2617f6d5cbf8b1067bafb7a380f878d41e41b064be81293368b7414a
compiled executor SHA-256: 26ae5c574e2bc1804644a7a7d60f4735e05e4f73df684e6844337271afd9cc33
compiled terminal driver SHA-256: d0bda02f1289fdca3614656dd1f33121ac251e31dac9c53c36d14ef2b99362d4

The committed investigation record and sanitized evidence retain the earlier initial-investigation snapshot. The current-head proof above is a separate fresh generation/replay, not relabeled old evidence. The changed profile/prompt hashes after the policy rebase are explicitly bound above. Raw diagnostics and the complete generated decision remain private; the relevant plan, observations, and integrity hashes are reproduced here.

Supporting checks and limits

All three builds and 166 focused tests passed after reconciliation, including both the landed release-policy tests and the new prerequisite-context tests. Fresh independent Codex reviews of the resolved staged candidate and committed range reported no accepted/actionable findings at their configured P0 threshold.

Earlier local full runs exposed pre-existing host Git-pruning and Homebrew/offline-fixture issues plus two timing-sensitive repair tests. These are documented in the historical investigation; no assertion, deadline, version shim, or persistent Git setting was changed to hide them. The prior head's full CI passed under the repository's declared Corepack bootstrap. This rewritten head must pass its own CI before landing; earlier green checks do not substitute.

The remaining automation risk is bounded and explicit: unusual repository build chains still depend on planner inspection. This patch supplies trusted facts rather than adding executor inference or rewriting. The proof does not guarantee unconstrained planner compliance, recording, publication, or outer review-workflow materialization. No live apply/close, manual workflow pause/dispatch, or PR1259 branch/landing mutation was performed.

@clawsweeper

clawsweeper Bot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

🦞👀
ClawSweeper picked this up.

Pull request received. I will update this pull request when review starts.

@clawsweeper clawsweeper Bot added P2 Normal priority bug or improvement with limited blast radius. proof: sufficient Contributor real behavior proof is sufficient. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. status: ⏳ waiting on author ClawSweeper has contributor-facing work open and is waiting for author action. labels Aug 27, 2026
@clawsweeper

clawsweeper Bot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Codex review: needs maintainer review before merge. Reviewed August 27, 2026, 11:42 AM ET / 15:42 UTC.

ClawSweeper review

What this changes

The PR supplies generated live-proof plans with trusted cold-checkout setup details, preserves executor command semantics, and adds focused tests and proof documentation.

Regression provenance

Possible regression — probable (reproduction; reviewed change). No predecessor PR is attributed.

Merge readiness

⚠️ Ready for maintainer review - 2 items remain

Keep this PR open for normal maintainer review: the introduced prompt context is narrowly sourced from the effective repository profile, and the exact-head cold-checkout proof supports the intended fix.

Priority: P2
Reviewed head: 213927808a1d14c11fb8507e45dcc061ceeade62

Review scores

Measure Result What it means
Overall readiness 🦞 diamond lobster (5/6) Strong exact-head controlled proof and focused source coverage support a narrowly scoped automation fix.
Proof confidence 🦞 diamond lobster (5/6) Sufficient (live_output): The PR body supplies exact-head after-fix production prompt generation, parsing, cold setup, and real-tmux execution with a recorded passing result; redact private host details in any future proof updates.
Patch quality 🦞 diamond lobster (5/6) No actionable review findings were identified.

Verification

Check Result Evidence
Real behavior Verified Sufficient (live_output): The PR body supplies exact-head after-fix production prompt generation, parsing, cold setup, and real-tmux execution with a recorded passing result; redact private host details in any future proof updates.
Evidence reviewed 5 items Trusted execution context: The introduced prompt builder derives setup, package manager, install-script policy, cold-checkout properties, and browser-only startup from the resolved repository profile rather than PR-authored GitHub context.
Focused regression coverage: The introduced tests assert that explicit and fallback profiles remain separate from untrusted item context and that normalized setup respects disabled, absent, and install-script-opt-in configurations.
Executor behavior preserved: The install-script normalizer was moved to a shared module and remains called by the executor before each configured setup command; no command replay or executor inference was added.
Findings None None.
Security None None.

Live Verification

Command: pnpm run build && node --test --test-name-pattern='preserves every terminal command including exact entry repeats|terminal entry with expectations executes once' test/decision-parser.test.ts test/live-proof.test.ts

Result: PASS (completed)

$ tsc -p tsconfig.json
✔ decision parser preserves every terminal command including exact entry repeats (8.697897ms)
✔ terminal entry with expectations executes once and preserves long final output publicly (60.043304ms)
ℹ tests 2
ℹ suites 0
ℹ pass 2
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 384.531746

Assertions:

  • PASS expect_output: preserves every terminal command including exact entry repeats
  • PASS expect_output: terminal entry with expectations executes once

How this fits together

ClawSweeper builds a review prompt from GitHub context and repository profiles, then executes declared proof plans in a separate cold checkout. This changes the trusted profile information available to the planner before it emits a terminal or browser plan.

flowchart LR
  A[Repository profile] --> B[Review prompt]
  C[GitHub item context] --> B
  B --> D[Generated proof plan]
  D --> E[Cold checkout setup]
  E --> F[Terminal or browser execution]
  F --> G[Verification result]
Loading

Before merge

  • Resolve merge risk (P1) - The controlled proof covers one terminal build chain; unusual repositories still rely on the planner inspecting their scripts and imports before it chooses prerequisites.
  • Complete next step (P2) - No discrete repair is identified; await exact-head check completion and ordinary maintainer review.
Agent review details

Security

None.

Review metrics

Metric Value Why it matters
Change footprint runtime +37 net, tests +111 net, docs/evidence +490 net, prompt +2 net Most of the 640 net lines preserve validation and audit context; the executable behavior change is concentrated in the prompt assembly and shared setup helper.
Controlled proof 1 declared/actual command; 2 observed expectations The submitted exact-head evidence directly checks that the generated plan adds the missing build prerequisite without replaying the diagnostic command.

Merge-risk options

Maintainer options:

  1. Land with the bounded planner responsibility (recommended)
    Accept the documented remaining variability after exact-head checks pass, because the executor remains declarative and the controlled cold-checkout proof validates the reported failure mode.
  2. Pause for broader planner coverage
    Defer landing if maintainers require evidence across additional repository build-chain shapes before changing the shared planning prompt.

Technical review

Best possible solution:

Retain the profile-derived cold-checkout contract and land it only after exact-head checks confirm the focused prompt and execution behavior.

Do we have a high-confidence way to reproduce the issue?

Yes. The retained no-build plan was replayed from a cold checkout after configured setup and failed to load the required compiled module; the final-head generated plan added the build and passed both assertions.

Is this the best way to solve the issue?

Yes. Supplying effective profile facts and an explicit cold-checkout contract fixes the information gap without making the executor guess or rewrite generated commands.

AGENTS.md: found and applied where relevant.

Codex review notes: model internal, reasoning high; reviewed against 2129a78a502e.

Labels

Label changes:

  • add rating: 🦞 diamond lobster: Overall readiness is 🦞 diamond lobster; proof is 🦞 diamond lobster and patch quality is 🦞 diamond lobster.
  • remove rating: 🦐 gold shrimp: Current PR rating is rating: 🦞 diamond lobster, so this older rating label is no longer current.

Label justifications:

  • P2: This is a bounded reliability fix for review automation rather than an urgent user-facing outage.
  • merge-risk: 🚨 automation: The PR changes trusted prompt input that determines executable live-proof plans.
  • rating: 🦞 diamond lobster: Overall readiness is 🦞 diamond lobster; proof is 🦞 diamond lobster and patch quality is 🦞 diamond lobster.
  • status: 👀 ready for maintainer look: ClawSweeper has no concrete contributor-facing blocker left for this PR. Sufficient (live_output): The PR body supplies exact-head after-fix production prompt generation, parsing, cold setup, and real-tmux execution with a recorded passing result; redact private host details in any future proof updates.
  • proof: sufficient: Contributor real behavior proof is sufficient. The PR body supplies exact-head after-fix production prompt generation, parsing, cold setup, and real-tmux execution with a recorded passing result; redact private host details in any future proof updates.

Evidence

What I checked:

  • Trusted execution context: The introduced prompt builder derives setup, package manager, install-script policy, cold-checkout properties, and browser-only startup from the resolved repository profile rather than PR-authored GitHub context. (src/clawsweeper-review-runtime.ts:465, 213927808a1d)
  • Focused regression coverage: The introduced tests assert that explicit and fallback profiles remain separate from untrusted item context and that normalized setup respects disabled, absent, and install-script-opt-in configurations. (test/review-prompt-context.test.ts:192, 213927808a1d)
  • Executor behavior preserved: The install-script normalizer was moved to a shared module and remains called by the executor before each configured setup command; no command replay or executor inference was added. (src/live-proof/setup.ts:1, 213927808a1d)
  • Exact-head real behavior proof: The PR body records a constrained generation through production parsing and real tmux on the final head, with a cold target, one declared/actual command, two observed expectations, and successful test output. (213927808a1d)
  • Feature history: Recent merged live-proof work includes the command-preservation fix and cold-build output work; the current PR commit is by the same recurring area contributor. (src/live-proof/execute.ts:34, c0af16349bfb)

Likely related people:

  • steipete: Authored the current prompt-context change and prior merged work on terminal-plan command preservation. (role: recurring live-proof contributor; confidence: high; commits: 213927808a1d, c0af16349bfb, a958131e8846; files: src/clawsweeper-review-runtime.ts, prompts/review-item.md, src/live-proof/execute.ts)
  • Vincent Koc: Recent merged work maintained cold-build behavior and final-result visibility in the same live-proof execution area. (role: recent adjacent contributor; confidence: medium; commits: f211e21fb89d, afe976209aa5; files: src/live-proof/execute.ts)

Rating scale

Score Internal tier Crab rank Meaning
6/6 S 🦀 challenger crab Exceptional readiness
5/6 A 🦞 diamond lobster Very strong readiness
4/6 B 🐚 platinum hermit Good normal PR; ordinary maintainer review
3/6 C 🦐 gold shrimp Useful, but confidence is limited
2/6 D 🦪 silver shellfish Proof or implementation needs work
1/6 F 🧂 unranked krab Not merge-ready
N/A NA 🌊 off-meta tidepool Rating does not apply

Overall follows the weaker of proof and patch quality.
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics.

Workflow

  • ClawSweeper keeps one durable marker-backed review comment per issue or PR.
  • Re-runs edit this comment so the latest verdict, findings, and automation markers stay together instead of adding duplicate bot comments.
  • A fresh review can be triggered by eligible @clawsweeper re-review comments, exact-item GitHub events, scheduled/background review runs, or manual workflow dispatch.
  • PR/issue authors and users with repository write access can comment @clawsweeper re-review or @clawsweeper re-run on an open PR or issue to request a fresh review only.
  • Maintainers can also comment @clawsweeper review to request a fresh review only.
  • Fresh-review commands do not start repair, autofix, rebase, CI repair, or automerge.
  • Maintainer-only repair and merge flows require explicit commands such as @clawsweeper autofix, @clawsweeper automerge, @clawsweeper fix ci, or @clawsweeper address review.
  • Maintainers can comment @clawsweeper explain to ask for more context, or @clawsweeper stop to stop active automation.

History

Review history (3 earlier review cycles)
  • reviewed 2026-08-27T13:55:50.776Z sha 6da6a9c :: needs changes before merge. :: [P2] Remove the release-owned changelog entry
  • reviewed 2026-08-27T14:17:58.781Z sha 6da6a9c :: needs maintainer review before merge. :: none
  • reviewed 2026-08-27T15:00:48.862Z sha 6da6a9c :: needs maintainer review before merge. :: none

@steipete

Copy link
Copy Markdown
Contributor Author

@clawsweeper re-review

The PR body now records the disposition of the changelog finding with exact target-policy citations. This targets openclaw/clawsweeper, not openclaw/openclaw: CONTRIBUTING.md lines 16–19 expressly distinguishes those repositories, and AGENTS.md lines 50–52 names the OpenClaw release-owned changelog. The one-line ClawSweeper entry is intentionally retained; the removal request is a target-scope false positive, not an accepted unresolved code finding.

Please review the unchanged head 6da6a9c9f69e2478065d2f9035fc82dcec68039e and the updated body. The current-head actual generation → cold setup → real tmux proof remains sufficient and unchanged, and the PR's own full CI is green.

@clawsweeper

clawsweeper Bot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

Re-review progress:

@clawsweeper clawsweeper Bot added merge-risk: 🚨 automation 🚨 Merging this PR could break CI, automerge, proof capture, label sync, or automation. rating: 🦞 diamond lobster Very strong PR readiness with only minor maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. and removed status: ⏳ waiting on author ClawSweeper has contributor-facing work open and is waiting for author action. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. rating: 🦞 diamond lobster Very strong PR readiness with only minor maintainer review expected. labels Aug 27, 2026
@steipete
steipete force-pushed the steipete/terminal-plan-prerequisites-20260827 branch from 6da6a9c to 2139278 Compare August 27, 2026 15:39
@steipete

Copy link
Copy Markdown
Contributor Author

@clawsweeper re-review

Please review updated head 213927808a1d14c11fb8507e45dcc061ceeade62 and the current PR body after reconciling the independently landed main changes. Both owners' tests and all changelog notes are preserved; all 166 focused tests pass. Fresh independent reviews and exact-head actual generation → cold setup → real tmux proof are complete, with both expectations observed and one test invocation. Earlier head reviews and CI do not substitute for this head's gates.

@clawsweeper

clawsweeper Bot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

Re-review progress:

@clawsweeper clawsweeper Bot added rating: 🦞 diamond lobster Very strong PR readiness with only minor maintainer review expected. and removed rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. labels Aug 27, 2026
@steipete
steipete merged commit 0e19056 into main Aug 27, 2026
14 of 15 checks passed
@steipete
steipete deleted the steipete/terminal-plan-prerequisites-20260827 branch August 27, 2026 16:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge-risk: 🚨 automation 🚨 Merging this PR could break CI, automerge, proof capture, label sync, or automation. P2 Normal priority bug or improvement with limited blast radius. proof: sufficient Contributor real behavior proof is sufficient. rating: 🦞 diamond lobster Very strong PR readiness with only minor maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant