fix(review): stop applying OpenClaw changelog rules to other repos - #1261
Conversation
|
🦞👀 Pull request received. I will update this pull request when review starts. |
|
Codex review: needs maintainer review before merge. Reviewed August 27, 2026, 10:10 AM ET / 14:10 UTC. ClawSweeper reviewWhat this changesThe PR moves the OpenClaw core changelog rule into its exact repository profile, changes shared review guidance to defer to the target repository, and adds prompt-selection coverage. Merge readinessKeep open for normal maintainer review: the patch correctly makes release-note guidance target-specific, preserves the OpenClaw core restriction, and has current-head real CLI proof with no actionable review findings. Priority: P2 Review scores
Verification
Live VerificationCommand: Result: PASS (completed) Assertions:
How this fits togetherClawSweeper builds each Codex review prompt from shared instructions plus the selected repository profile. That prompt drives structured review decisions and rendered reports for configured target repositories. flowchart LR
A[GitHub item] --> B[Review context]
B --> C[Shared review instructions]
C --> D[Selected repository profile]
D --> E[Codex review prompt]
E --> F[Structured decision]
F --> G[Rendered report]
Before merge
Agent review detailsSecurityNone. Review metrics
Root-cause clusterRelationship: Members:
Proposal only: this assessment does not dispatch repair, suppress jobs, mutate sibling items, close, or merge anything. Technical reviewBest possible solution: Keep the core-only restriction in the openclaw/openclaw profile and let every other target repository supply its own release-note policy. Do we have a high-confidence way to reproduce the issue? Yes, from source: current main injects the changelog rule from the shared template for every target, whereas the runtime already has exact normalized profile selection available for the repair. Is this the best way to solve the issue? Yes. Moving only the core-specific instruction to the existing exact-profile seam removes the wrong global scope while retaining the core rule and existing safeguards. AGENTS.md: found and applied where relevant. Codex review notes: model internal, reasoning high; reviewed against 0bd84d42bc04. LabelsLabel changes:
Label justifications:
EvidenceWhat I checked:
Likely related people:
Rating scale
Overall follows the weaker of proof and patch quality. Workflow
HistoryReview history (4 earlier review cycles)
|
dbbddba to
14214b6
Compare
|
@clawsweeper re-review The main body now begins with current-head proof for The prior stale-head proof concern is addressed. Please review the current head and final body; no source or safety guard was changed to obtain this result. |
|
🦞🧹 I asked ClawSweeper to review this item again. Re-review progress:
|
Current-Head Live Proof and Landing Status
Exact head:
14214b6f84a7be78785b24d0360080be3c368630. Base and observed live-review main:0bd84d42bc0487c32af2285006884d4f9b2f7763. The requested actual CLI live test is complete. This section supersedes the earlier head/CI/pending-live-test statements preserved as history below.The rebase resolved only an additive CHANGELOG conflict, preserving both entries. The six other task files remain byte-identical to the earlier reviewed head. Build, 73 focused tests and the retained automerge assertion pass on this head; the fresh committed-branch Codex autoreview is CLEAN at its default P0 scope. No policy, release/repair/close/merge guard, timeout, dependency or coverage threshold was weakened.
Actual user-facing CLI execution
The candidate built application was launched as
pnpm run review -- --local-only, targeting the liveopenclaw/clawsweeperPR1261. It created its normal fresh managed target checkout, hydrated real GitHub item/context data, passed the native read-only checkout inspection, called real Codex, parsed the generated decision and rendered the local Markdown report. This was not a unit test, internal-function-only smoke, fake GitHub transport or stub model.Observed execution: CLI exit 0;
review_status: complete;pull_head_shaexactly14214b6f84a7be78785b24d0360080be3c368630;review_cache_hit: false;review_sandbox: read-only;local_checkout_access: verified;review_terminal_failure: false;review_checkout_inspection_failed: false. The model stream ended withturn.completedand noerrororturn.failedevent. Recorded model elapsed time was 650,510 ms. Both the task checkout and managed target remained clean; the app removed its temporary review tree.The actual parsed result was
keep_open, high confidence, zero review findings,overallCorrectness: "patch is correct", and sufficient real-behavior proof. No incorrect ClawSweeper changelog-removal finding was raised. Its conclusion was: “Keep release-note restrictions in the owning repository profile, with shared guidance deferring to each target’s established policy and existing mutation safeguards unchanged.”keep_openalone was not used as proof: the parent inspected the actual findings, rationale, risks, parsed/rendered status and runtime completion.The generated review still mentioned the previous incomplete CLI attempt because that was the body supplied at invocation. That warning is now addressed by the subsequently observed outer CLI exit, successful native parse/render, completed report and cleanup, not by treating the raw model response as successful execution. This final body records the outcome for a fresh durable review.
The authoritative static/profile section before GitHub Context identified ClawSweeper, contained the shared own-policy/no-blanket-permission rule, and contained neither the old ambiguous restriction nor the core-only restriction. The real PR body quotes BEFORE text as historical evidence, so absence was not incorrectly asserted over untrusted body quotations.
Environment: local macOS arm64, Node 24.19.0, pinned pnpm 11.10.0, existing authenticated Codex/GitHub runtime. The normal read-model-unavailable fallback performed live GitHub polling. No container, deployment, scheduler rollout, live apply/close/repair or review publication is claimed by this local-only command. The ordinary model budget remained 1,200,000 ms, including the application's retry budget.
Reproducible command shape from the candidate checkout, using a fresh artifact directory and the configured approved Codex runtime (the recorded private model selector is intentionally omitted from this public receipt):
Declared constraints and failed attempts: Attempt 1 failed before inference because host
fetch.prune=trueremoved the managedorigin/mainref. Attempt 2 used only the process-local pruning setting above, completed real model inference with no findings, but failed the app's strict parser because the generatedliveProofPlan.entrywas multiline. Attempt 3 completed with the explicit format-only additional prompt shown above. No verdict was prescribed, JSON was not rewritten, and schema/parser validation and containment were not bypassed. Global Git pruning remained true. The first two attempts remain failed evidence; this is not a claim that the default unqualified CLI is reliable under every host setting or model output.This is one successful advisory live review sample with those declared inputs, not a stochastic reliability benchmark. The app's normal compaction retained 12,024 body characters in the model context; no claim is made that every byte of the full PR body was presented. The captured source revision was
34a57a34941417f4a59cdcf2d7c151c7240f06a4144e9eb0d2ad5a333e40306b; the final durable review must bind to the newer final-body revision, not reuse that local sample's binding. Source/build identity remained fixed through execution. Raw reasoning/transcripts are not included in this proof.aab0edb4496493080c6059da50436ef0453ee88b665ccab55c656e67dd6ce5814412169b9090b7835e5eddae9a3d7a06bdf23a234b6d2daec11b1c5dd9414b70727ec2d5c9579aa414d0ee708cca537cd2ef40dc24145ca71d0037dabeafdb1d5595ab43c92f992b298843d14fcba031d11d5d12ee20e142e1f5c32a0f30a62fFull local artifacts are retained; these hashes do not claim to publish those files. The observed fields, command, result and limits above are the durable public receipt.
Current-head policy matrix and CI
Separately, the parent executed the actual built production prompt assembler on this current commit for all 24 frozen contexts (21 construction cases plus three model-input contexts). Every newly generated prompt was byte-identical to the prior captured candidate. Core-only selection, target policy ownership and no-blanket-permission guidance passed. This was an executed current-head replay, not merely source comparison. The earlier six controlled model responses remain historical samples, not relabeled as new calls.
Exact-head CI run 33076175845 is successful, including full pnpm check, sparse repair build and Windows launcher. Attempt 1 failed one unchanged live-proof output assertion (3,781 passed, 1 failed, 8 skipped); its focused two-test file passed locally, then only the failed CI job was rerun. Both this failed attempt and the historical local full-suite failures below remain disclosed. No test or guard was changed to obtain green CI.
Finding disposition: the durable review's stale-head proof concern is addressed by the fresh current-head 24-context execution, exact-head CI and the completed actual CLI run above. There is no accepted code finding or outstanding Rank-up move from the reviewed result. No Bay change is required: this changes policy text, not status/data/telemetry/navigation/action contracts. Before landing, the latest durable ClawSweeper review must match this exact head and this final body; no stale review or proof label is being used as an override.
What Problem This Solves
Resolves a problem where ClawSweeper reviews apply
openclaw/openclaw's release-owned changelog policy to PRs in other repositories. The initial durable reviews of #1254 and #1255 requested removal of valid maintainer-authored ClawSweeper changelog entries. Both fresh reviews corrected the finding after the PR bodies supplied the explicit repository scope fromAGENTS.mdandCONTRIBUTING.md.The durable histories preserve the original findings at the same reviewed heads: 1254 review history, initially “Remove release-owned CHANGELOG entries”; 1255 review history, initially “Remove the release-owned changelog entry.” This change fixes policy ownership rather than removing the valid notes or requiring every PR body to defend its repository identity.
Why This Change Was Made
Inspection of the actual assembled prompts established that the complete shared template was included for every target before the selected repository profile was appended. The template said “For OpenClaw PR release-note review” without an exact owner/repo; the ClawSweeper profile did not counter that ambiguity.
Bounded history shows that c9b353df34af made the changelog release-owned while carrying forward the ambiguous “OpenClaw” repository name. b757ba62ae80 later documented the literal owner/repo distinction in CONTRIBUTING.md; the shared prompt still lacked it.
The existing
repositoryProfileFor(item.repo).promptNoteseam now owns the restriction in the built-inopenclaw/openclawprofile. Selection uses the normalized exact target identity supplied by the review workflow, not the organization, display name, author association, linked repository or PR body. The shared template defers to the target's own policy and explicitly says that being outside core does not grant contributors or workers permission to edit release-owned files. No new configuration, fallback, policy flag or runtime branch is introduced.The complete core restriction remains: normal PRs and repair/automerge/autofix work do not edit its release-owned changelog; missing changelog edits are not blockers; release context belongs in PR bodies/commit messages; attribution restrictions remain intact. All existing ClawSweeper entries are preserved. Repair, close, release and merge guards, schema, limits and exact-refresh behavior are outside this patch. In particular, this does not modify #1258 or its branch.
User Impact
Reviews receive an unambiguous target-specific release-note contract without inheriting another repository's restriction. This removes the deterministic source of the scope confusion; it does not promise that stochastic model reviews can never misclassify policy. Contributors still follow their own target repository's release-note rules.
OpenClaw Bay Impact
No Bay code change is needed. Review prose/findings may change as intended, but there is no record schema, queue lifecycle, status projection, public API, navigation or action-control change. Bay remains observer-only. The existing review-policy hash already includes the template and selected profile, so this patch changes normal policy identity without a manual version bump or fleet re-review dispatch.
Documentation Impact
The active target-repository reference now documents the exact profile ownership boundary, and the binding
AGENTS.mdsentence namesopenclaw/openclawliterally.CONTRIBUTING.mdalready specifies the correct scope and is unchanged. The reference remains owned by ClawSweeper maintainers;src/repository-profiles.ts,config/target-repositories.jsonand the production review assembler are its source of truth. Changes to target selection, profile prompt ownership or release-note policy require a documentation update. This patch adds the expected concise maintainer ClawSweeper changelog entry, not an OpenClaw release edit.Earlier Evidence — Pre-Rebase Head
Exact head:
dbbddba966f4cc3d17c28c4dc126a24f2356d6ef. Direct parent/tested base:71df3a1ce714d737e250008597075bb5eaeb2ac4. This is one focused seven-file commit on the refreshed base, including the independently landed terminal-proof and pinned-pnpm fixture changes. The old unrelated-history localmainwas preserved; only this new task branch was committed/rebased.Node 24.19.0, pinned pnpm 11.10.0: build, 73 focused tests, the preserved automerge prompt assertion, static checks, lint, and 12 changed-coverage tests passed. The new table-driven coverage exercises 56 combinations of target identity/role/review variant, including ClawSweeper versus core, configured ClawHub, both generic owner fallbacks, mixed-case normalization, misleading body/title/URL identity, missing changelog, issues, previous reviews and autogenerated PRs. The independent precommit and exact committed-branch Codex autoreviews are CLEAN at the helper's default P0 scope, with no accepted/actionable findings reported. Parent inspection also verified that repair, close/merge guards, schema, target config, package manifest and lockfile have no branch delta.
The initial full-check launch stopped before tests because nested pnpm selected host 11.24.0 instead of pinned 11.10.0. A task-local Corepack shim fixed that launcher mismatch without editing package/global configuration. The subsequent full check completed and failed: 3,735 passed / 42 failed / 9 skipped / 0 cancelled. All 42 failures were missing synthetic
origin/mainrefs in unchanged command/Git/target-validation tests. The host has globalfetch.prune=true; a focused failing Git fixture passed with a process-localfetch.prune=falseoverride, while the global setting remained true. The complete controlled rerun also failed: 3,773 passed / 4 failed / 9 skipped / 0 cancelled. It removed the missing-ref failures but hit four unchanged deadline-sensitive target-validation tests: install-managed workspace links, the shared setup deadline, the shared checkout-identity timeout, and reserved mutation-proof time. Those same four tests then passed 4/4 in one focused rerun with coverage instrumentation and unchanged timeouts. Both failed full runs remain failed evidence; no local full-green claim is made. No test, timeout, coverage threshold or production guard was changed.Hosted exact-head CI is green: CI run 33054612460, head
dbbddba966f4cc3d17c28c4dc126a24f2356d6ef, passed fullpnpm check, sparse repair build smoke and Windows Codex launcher. Parent verification independently matched both the run head and live PR head to that SHA. The hosted result does not relabel either failed local full run. Other checks passed;[code]smithwas skipped.This evidence update changes only the main PR body. It is not merge approval: any eventual landing still requires the latest durable ClawSweeper review to apply to the current head and body, with every accepted finding resolved. No merge or automerge was requested or executed.
Executed focused commands (Node 24 and the pinned pnpm launcher on PATH):
node --test --test-name-pattern='^review prompt keeps automerge opt-in from becoming generic manual review$' test/clawsweeper.test.tsThe controlled full-check command uses a process-local Git setting, not a repository/global write:
Real Behavior Proof
Claim and surface. The real built prompt assembler selects the release-note contract by target owner/repo, and real generated review responses preserve valid ClawSweeper maintainer notes without granting contributor permission in core or in another restrictive repository. This is prompt/context and model-boundary proof, not a live publication/apply/merge exercise.
Source. BEFORE is the exact base template and base profile. AFTER is the head template and head profile. The upstream terminal-proof prompt/schema changes are present on both sides, so this comparison does not conflate them with the release-note fix.
Actual BEFORE text, injected into every target:
Actual AFTER shared guidance:
Only the core profile now supplies:
The full moved paragraph—including missing-changelog, release-context and attribution rules—was checked for equality after only whitespace normalization and the explicit repository-name replacement. The full assembled prompt pairs differ only by the shared paragraph replacement and core-profile addition.
Executed construction. A disposable
git archiveof the exact base (no Git metadata or registered worktree) and the clean candidate worktree were compiled with the locked TypeScript compiler. Forty-eight fresh Node subprocesses invoked the actual compiledreviewPromptForTestexport, which delegates directly to productionbuildReviewPrompt: 21 construction prompts plus three model prompts per variant. The matrix covered initial PR, issue, follow-up and autogenerated variants across five targets, plus conflicting body/URL identity. BEFORE injected the ambiguous restriction into all 21; AFTER emitted the core block only in the four core-target variants and not in the other 17. Both sets used byte-identical frozen inputs. Source/build/input hashes were stable before and after capture, and the archived source diff exactly matched the actual seven-file branch diff.The capture program's entire invocation seam was:
The executed external generator accepted
(repository, frozen-input-root, base-SHA, head-SHA, new-output-directory), built the base snapshot, and launched that seam with cwd at the relevant source root. It never invoked a model or reimplemented the assembler. Its SHA-256 wasea9076f92a355238ffad2a9779d0faf0f140aac72baadd7fbe41f519e2886153. Full source/input/prompt snapshots and the generator are retained locally; this paragraph does not claim they are publicly downloadable. The source links, observed excerpts, outcome table and hashes here are the durable public proof receipt.Controlled generated reviews. Each scenario proposed only one accurate changelog bullet documenting an already-existing queue guarantee. The ClawSweeper author was a maintainer (
MEMBER); the two restrictive targets had ordinary contributor authors and no separate release authorization. The initial PR body contained no policy quotation or finding-disposition rescue. The full old ClawSweeper AGENTS.md/CONTRIBUTING.md files were supplied through the assembler'sadditionalPromptlocal-context seam; other repository policies and runtime evidence were explicitly synthetic. Those policy and input bytes were identical across BEFORE and AFTER, including the older ambiguous AGENTS wording—AFTER did not receive a clearer policy fixture.Six fresh authenticated Codex inferences used the exact assembled prompts and matching production decision schema: one per scenario per variant, no inference retries. Each was ephemeral, read-only, with user/project instruction auto-loading, shell/tools, plugins, hooks, browser, delegation and web search disabled. All six exited 0 within the unchanged 180-second per-call timeout, recorded zero tool events, and parsed through the real production decision parser. The model was not a stub or a keyword classifier. Environment: local macOS arm64, Node 24.19.0 for generation/parsing, pinned pnpm 11.10.0, existing CLI authentication; no secrets were included in context. No container, remote lease, production checkout inspection or live GitHub mutation is claimed.
openclaw/clawsweeper, valid maintainer noteopenclaw/openclaw, ordinary contributor editopenclaw/example-tool, its own restrictive policyAll six returned
keep_open, so that field alone was not used as proof. The assessment above checks findings, risks, explicit release decisions and rationale. In particular, the core model used a required release-owner decision rather than a finding; the other target retained a concrete P2 ownership finding.No classifier-reproduction claim. The ClawSweeper baseline already classified the controlled fixture correctly. Three earlier exploratory baseline probes also produced correct policy outcomes and remain retained separately. The production incident is evidenced by the durable histories linked above, not claimed to be reproduced by these fixtures. The deterministic before/after result is removal of the cross-repository prompt injection; these sampled generated responses demonstrate the intended policy distinctions, not a measured reliability improvement.
Selected artifact identity. These are hashes of actual captured bytes, not expected-output fixtures:
67f5c29c24d2340a35a5a2596c2acb44c4152825781c7ae32649c45664290c727d5399457dd9af003031fde48ddcae46a2ccefb5fd0cadf59223e6907757a558d9cb2bb3926ec9cd4361de011464e569122406765f223efcd03319ec1476ae9d22c74f749bcd136572fee4786e6e5a69ebb552ed2cf102818e91e8a7fb8886044a658d090488acc0d04d08c465eceb103c922e4d541c1fcdfe8677e50eec9e03fba166d744d1ccdf281c84846e6549cddcdafbcddb18f2fb3fef509686710e895f9328ee1958ad3410407d701becea317c0c4ae650385afe85563516447c61e5de6f5f1c92df7f9bf077f7f6513c428ca8f62a5b8856f8ceebbd33cbd656d416The production schema is identical across the final pair:
758dd96182b5db7b9b4cb98164b4dbb9efd30d3cb7baf7590ad2159ef7b1aad9. Full-index source-diff SHA-256:b44e2bfdba16ace1360fd1cb1cdb2dfd3dff408bad3c471f8006ffaaf98dba27.Limits and failed attempts. This is controlled fixture context, not a replay of the original production review hydration or discovery order. Supplying complete policy files in the Maintainer Request section increases policy salience relative to ordinary repository discovery. One generated response per final cell cannot guarantee future stochastic behavior, and the synthetic SHAs in fixture evidence are not source provenance. No production rollout, durable re-review, mass invalidation, release, merge, repair or close was executed. Source-snapshot setup first aborted on pnpm's attempted dependency reconciliation; a later external verifier assertion incorrectly assumed the core profile was one line. Both failed attempts were retained and corrected in the external harness; neither changed tracked source, dependencies, tests or gates. The separate full-suite environment failures remain disclosed above.