Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
128 commits
Select commit Hold shift + click to select a range
72c185d
feat(execprofile): add the zeromaxing posture rung
gnanam1990 Jul 28, 2026
96e8e3f
feat(agent): inject the zeromaxing reminders below the cache breakpoint
gnanam1990 Jul 28, 2026
411b345
feat(tui): make /effort zeromaxing the entry point on both selection …
gnanam1990 Jul 28, 2026
3a67a91
fix(execprofile): describe the posture's delta against the caller's o…
gnanam1990 Jul 28, 2026
fa8368b
fix(tui): apply the posture's effort fill on models the catalog canno…
gnanam1990 Jul 28, 2026
bc27118
fix(tui): treat an unlisted model as unknown, not as having no reason…
gnanam1990 Jul 28, 2026
63af9a8
feat(specialist): capture and sequentially execute a declared plan
gnanam1990 Jul 28, 2026
00ac06d
feat(cli): register the orchestrate tool and map a partial plan to ex…
gnanam1990 Jul 28, 2026
48f0471
feat(tui): flip the orchestrate posture gate from the real handlers
gnanam1990 Jul 28, 2026
54f7f3c
fix(specialist): make the plan tool's parent grant real and empty mea…
gnanam1990 Jul 29, 2026
d852a69
fix(tui): make both effort doors ask the same question
gnanam1990 Jul 29, 2026
56ab486
fix(cli): record the failed plan task's session and spend
gnanam1990 Jul 29, 2026
9e34f00
fix(cli): measure the run's real depth for plan admission
gnanam1990 Jul 29, 2026
3cb060e
test(specialist): compare the two grant enforcement points against ea…
gnanam1990 Jul 29, 2026
754a2e5
feat(tui): make a running plan visible
gnanam1990 Jul 29, 2026
b08c755
feat(tui): add the /plans panel with the plan's dependency shape
gnanam1990 Jul 29, 2026
5ca721b
fix(specialist): a plan task inherits the parent's model
gnanam1990 Jul 29, 2026
a2db444
fix(specialist): a resumed specialist keeps the run's model
gnanam1990 Jul 29, 2026
ec2916b
fix(specialist): the two read-only tool sets agree again
gnanam1990 Jul 29, 2026
9a07268
perf(agent): drop the confirmation policy from a run that cannot act
gnanam1990 Jul 29, 2026
b2e8397
feat(specialist,tui): unbounded plans by default, collapsed plan panel
gnanam1990 Jul 29, 2026
14bf851
feat(tui): fade finished plan tasks, and say what /effort accepts
gnanam1990 Jul 29, 2026
dd5e5bc
fix(tui): the effort picker offers what /effort actually accepts
gnanam1990 Jul 29, 2026
6b72ce6
fix(tui): five reporting defects a real plan run exposed
gnanam1990 Jul 29, 2026
8acb7af
fix(tui): the plan budget line counts while the plan runs
gnanam1990 Jul 29, 2026
3050a07
feat(tui): the plan detail view — phases left, live agent detail right
gnanam1990 Jul 29, 2026
8f87c28
feat(tui): the running plan fills the sidebar's PLAN section, in colour
gnanam1990 Jul 29, 2026
f428251
feat(tui): the plan owns the right column — progress bar, task list, …
gnanam1990 Jul 29, 2026
a2c4539
feat(tui): the zeromaxing posture glows, in the footer and where it i…
gnanam1990 Jul 29, 2026
c8f31e8
feat(tui): hover on the posture chip and the plan rows, and the chip …
gnanam1990 Jul 29, 2026
1ed8802
feat(specialist): per-task stall watchdog (gap report §5.10)
gnanam1990 Jul 29, 2026
16eeac7
fix(specialist): resuming a task must not widen its authority
gnanam1990 Jul 29, 2026
0f21fe3
feat(tui): a plan run in the TUI is durable (gap report §5.2)
gnanam1990 Jul 29, 2026
40471c8
feat(config): the plan-size ceiling is a configurable tier (gap repor…
gnanam1990 Jul 29, 2026
b8e6f8c
feat(specialist): a stalled plan task is retried, and only a stalled …
gnanam1990 Jul 29, 2026
6c5bfc1
feat(tui): stop or pause a plan without stopping the turn (gap report…
gnanam1990 Jul 29, 2026
d430127
fix(tui): the plan list follows the running task instead of pinning t…
gnanam1990 Jul 29, 2026
d79c5c5
feat(specialist): save a plan that worked and run it again (gap repor…
gnanam1990 Jul 29, 2026
acb0971
feat(specialist): resume a plan from where it stopped (gap report §5.…
gnanam1990 Jul 29, 2026
bcefefd
feat(specialist): ship one worked plan in the binary (gap report §5.14)
gnanam1990 Jul 29, 2026
c2fc758
fix(cli): the model was never told the orchestrate tool exists
gnanam1990 Jul 29, 2026
321411b
refactor(specialist): delete the orchestrate limits override, which n…
gnanam1990 Jul 29, 2026
bfa9058
feat(tui): /plans restart runs the last plan from the beginning (gap …
gnanam1990 Jul 29, 2026
3e650f2
feat: run a plan in the background and report it on a later turn (gap…
gnanam1990 Jul 30, 2026
c90ecc6
fix(specialist): drop update_plan from the plan grant, which cost eve…
gnanam1990 Jul 30, 2026
c30c46e
feat(tui): a permission prompt can show what it is approving (write-t…
gnanam1990 Jul 30, 2026
b72de91
feat(specialist): a write-capable plan gets a worktree, or does not r…
gnanam1990 Jul 30, 2026
bfb2166
feat(specialist): a plan task may write, if it asks and is isolated (…
gnanam1990 Jul 30, 2026
ceed8fe
fix(specialist): tell a plan task to USE its tools, not merely that i…
gnanam1990 Jul 30, 2026
5d945ac
feat(specialist): run independent plan tasks in parallel (gap report …
gnanam1990 Jul 30, 2026
2bea0a5
feat(tui): each plan task's progress lands on its own card (gap repor…
gnanam1990 Jul 30, 2026
e6ff4ae
test: prove resume survives a partial concurrent run (gap report §5.1…
gnanam1990 Jul 30, 2026
50dd4a5
fix(credstore): serialize the credential read-modify-write across pro…
gnanam1990 Jul 30, 2026
61b1bd6
fix(sandbox): a temporary grant is refcounted, so a sibling's cleanup…
gnanam1990 Jul 30, 2026
d71d437
fix(worktrees): a cancelled git dies as a process group, not just a s…
gnanam1990 Jul 30, 2026
75f2906
fix(specialist): an orchestrate approval must never be remembered
gnanam1990 Jul 30, 2026
4ca408b
feat(tui): the posture chip leaves amber and shimmers under the cursor
gnanam1990 Jul 30, 2026
bd882dc
feat(tui): the posture is a live word, not a coloured box
gnanam1990 Jul 30, 2026
0c5f3b3
feat(tui): the posture reads lowercase and its hover flows
gnanam1990 Jul 30, 2026
f0a65e5
fix(tui): a finished plan task must close its OWN card, not the last …
gnanam1990 Jul 30, 2026
b77a197
feat(tui): one plan, one surface — the inline panel stands down for t…
gnanam1990 Jul 30, 2026
c630fdf
feat(tui): click an agent in AGENTS to open its brief, spend and reason
gnanam1990 Jul 30, 2026
d5e0333
fix(tui): an agent with no card key must not expand itself and shove …
gnanam1990 Jul 30, 2026
63ed5d9
fix(tui): the PLAN section holds both plans, so the footer line can f…
gnanam1990 Jul 30, 2026
8ff0961
fix(tui): a running task is work underway, not progress — the bar mus…
gnanam1990 Jul 30, 2026
b37a6cb
feat(tui): a toggle for finished agents, and each one opens on what i…
gnanam1990 Jul 30, 2026
17bb880
fix(tui): a sidebar row that was never drawn must not be clickable
gnanam1990 Jul 30, 2026
3b1625b
fix(specialist): one plan per surface, on the path the model drives too
gnanam1990 Jul 30, 2026
4fbc5f8
test(specialist): measure concurrency against the host, not against t…
gnanam1990 Jul 30, 2026
3b2caa0
test(cli): neutralise XDG_CONFIG_HOME in the sandbox policy golden
gnanam1990 Jul 30, 2026
7414599
fix(tui): resume the named plan, sanitize task text, keep background …
Vasanthdev2004 Jul 31, 2026
ca49b69
fix(sandbox): release a temporary root under one lock hold
Vasanthdev2004 Jul 31, 2026
5fe5b43
fix(tui): guard the zeromaxing effort switch and two mouse hit-testers
Vasanthdev2004 Jul 31, 2026
216fe03
fix(specialist): reject negative plan timeouts, correct two false claims
Vasanthdev2004 Jul 31, 2026
71ec1b2
fix(specialist): make max_wall_seconds bound a concurrent plan
Vasanthdev2004 Jul 31, 2026
85c630c
fix(specialist): let a task granted write tools actually use them
Vasanthdev2004 Jul 31, 2026
71cfd67
feat(config): per-role plan model preferences, and cache counts a chi…
gnanam1990 Aug 1, 2026
b001b91
feat(specialist): choose a model per task, prove it runs, and bound w…
gnanam1990 Aug 1, 2026
0e3dc54
feat(cli): wire model discovery, proving and child scope to the live run
gnanam1990 Aug 1, 2026
cb472d6
feat(tui): show which model ran a task, and what a plan left undone
gnanam1990 Aug 1, 2026
ed4634c
fix(specialist): five admission and lifecycle defects from review
gnanam1990 Aug 1, 2026
b7fea56
feat(specialist): teach the verify convention, and show what happens …
gnanam1990 Aug 1, 2026
23eedac
feat(specialist): give a plan with no bound at all a wall backstop, a…
gnanam1990 Aug 1, 2026
3e84cd5
fix(config): project maxTurns may tighten only
gnanam1990 Aug 1, 2026
d98ae9f
feat(specialist): a plan no longer dies because its inputs were cut s…
gnanam1990 Aug 1, 2026
363ea6a
feat(specialist): reserve budget for later work, and declare the plan…
gnanam1990 Aug 1, 2026
15f73ac
fix: repair the Windows smoke break and the review findings behind it
gnanam1990 Aug 2, 2026
8f555d9
feat(zeromaxing): make evidence the posture's job, not the prompt's
gnanam1990 Aug 2, 2026
3b7aaca
fix(tools): an edit must not erase the reads it did not disturb
gnanam1990 Aug 2, 2026
5d3c69d
test(specialist): do not assert Windows rename semantics we cannot run
gnanam1990 Aug 3, 2026
e0b9915
fix(prompt): ask the provider which model family it serves
gnanam1990 Aug 3, 2026
b422fe1
feat(specialist): a saved plan can be run against a different target
gnanam1990 Aug 3, 2026
7f56dda
feat(agent): bound a run by what it spends, not only by its turn count
gnanam1990 Aug 3, 2026
ed33f6f
feat(zeromaxing): raise the turn ceiling and bound what a run may spend
gnanam1990 Aug 3, 2026
356ff1b
fix(specialist): a plan task must not report a number it never measured
gnanam1990 Aug 3, 2026
93fe09b
fix(specialist): a plan task's numbers and its unfinished work both s…
gnanam1990 Aug 3, 2026
6fd2974
fix(execution): remember a finished session for as long as a poll may…
gnanam1990 Aug 3, 2026
3eb9e41
feat: four bounds and capabilities a plan run showed were missing
gnanam1990 Aug 3, 2026
7b19c8f
feat(cli): wire the new bounds and capabilities into the TUI
gnanam1990 Aug 3, 2026
d510aba
fix(specialist): the bundled research plan takes its subject as a par…
gnanam1990 Aug 3, 2026
eed0791
fix(specialist): the budget schema declares the bounds it enforces
gnanam1990 Aug 3, 2026
e8db70d
fix(tui): the zeromaxing chip is hit-tested on the status line only
gnanam1990 Aug 3, 2026
72de267
fix(tui): the AGENTS section stops contradicting its own header
gnanam1990 Aug 3, 2026
ff95e75
fix(specialist): a plan task's answer no longer carries the child's s…
gnanam1990 Aug 3, 2026
09ae593
feat(zeromaxing): resumable plans, read-grant propagation, size-aware…
gnanam1990 Aug 4, 2026
3c4626f
refactor(specialist): drop the dead PlanProgress.Done/Remaining helpers
gnanam1990 Aug 5, 2026
460242d
feat(specialist): per-role model auto-assignment for delegated sub-ag…
gnanam1990 Aug 5, 2026
b168bca
feat(tui): the zeromaxing skin — MODELS panel, electric gradient, sta…
gnanam1990 Aug 5, 2026
2aa9bba
feat(tui): dracula's status signals move to the scheme's own cyan/pink
gnanam1990 Aug 5, 2026
606dc22
refactor(tui): /profile answers in one line
gnanam1990 Aug 5, 2026
1d3d187
fix(cli): a headless run's children get only the granted extra write …
gnanam1990 Aug 6, 2026
e8e99b2
fix(tui): /effort auto cannot leave the zeromaxing posture mid-run
gnanam1990 Aug 6, 2026
746af5a
fix(tui): saved-plan summaries pass the same sanitizer as live ones
gnanam1990 Aug 6, 2026
6c7cdf6
fix(specialist): the workspace-reach guard works on Windows paths
gnanam1990 Aug 6, 2026
2e065e4
chore: move test-only shims to export_test.go and drop dead code
gnanam1990 Aug 6, 2026
881e52e
fix(specialist): a trailing full stop is not part of a Windows path
gnanam1990 Aug 6, 2026
d78a9d3
feat(specialist): the posture appends a verifier to a plan that has none
gnanam1990 Aug 7, 2026
647e58f
feat(specialist): a plan routes to the top-ranked models, not the who…
gnanam1990 Aug 7, 2026
d20090a
fix(specialist): a plan task killed by the provider gets one more att…
gnanam1990 Aug 7, 2026
7ed4a2e
fix(specialist): the model shortlist never narrows a priced catalogue
gnanam1990 Aug 7, 2026
e09f536
fix(specialist): the appended verifier cannot refuse or fail the auth…
gnanam1990 Aug 7, 2026
4a56709
fix: six more defects a high-effort review found in the posture work
gnanam1990 Aug 7, 2026
def1d77
fix(specialist): the shortlist never drops a model whose size it cann…
gnanam1990 Aug 7, 2026
d5325f7
fix(specialist): the shortlist keeps both ends and version text is no…
gnanam1990 Aug 7, 2026
fc8dcb6
fix(tui): cancel settles foreground children even while a background …
gnanam1990 Aug 7, 2026
d5d3b23
test(agent): restore the tool-definitions golden the rebase reverted
gnanam1990 Aug 7, 2026
41fa169
test(agent): move the tool-definitions golden for main's #867
gnanam1990 Aug 7, 2026
ce1c5b2
fix(tools): probe the Windows shell inside the sandbox that will run it
gnanam1990 Aug 8, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions internal/acp/agent.go
Original file line number Diff line number Diff line change
Expand Up @@ -253,6 +253,7 @@ func (a *Agent) runTurn(ctx context.Context, sess *acpSession, userText string,
Cwd: sess.cwd,
SessionID: sess.id,
ProviderName: resolved.Provider.Name,
ModelFamily: providercatalog.ModelFamilyFor(resolved.Provider.CatalogID),
Model: resolved.Provider.Model,
Registry: registry,
Sandbox: sandboxEngine,
Expand Down
135 changes: 135 additions & 0 deletions internal/agent/child_progress_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,135 @@
package agent

import (
"context"
"testing"

"github.com/Gitlawb/zero/internal/streamjson"
"github.com/Gitlawb/zero/internal/tools"
)

// progressProbeTool records whether the loop handed it a progress callback.
// streams mirrors what a real tool declares via tools.ChildProgressStreamer.
type progressProbeTool struct {
name string
streams bool
got *bool
}

func (t *progressProbeTool) Name() string { return t.name }
func (t *progressProbeTool) Description() string { return "probe" }
func (t *progressProbeTool) Parameters() tools.Schema {
return tools.Schema{Type: "object", Properties: map[string]tools.PropertySchema{}}
}
func (t *progressProbeTool) Safety() tools.Safety {
return tools.Safety{SideEffect: tools.SideEffectRead, Permission: tools.PermissionAllow}
}
func (t *progressProbeTool) Run(context.Context, map[string]any) tools.Result {
*t.got = false
return tools.Result{Status: tools.StatusOK, Output: "ok"}
}
func (t *progressProbeTool) RunWithOptions(_ context.Context, _ map[string]any, options tools.RunOptions) tools.Result {
*t.got = options.Progress != nil
return tools.Result{Status: tools.StatusOK, Output: "ok"}
}

// StreamsChildProgress makes this type implement tools.ChildProgressStreamer;
// the flag drives the answer. A tool that never declares at all is modelled by
// silentProbeTool below, which genuinely does not implement the interface.
func (t *progressProbeTool) StreamsChildProgress() bool { return t.streams }

// silentProbeTool does NOT implement tools.ChildProgressStreamer at all — the
// state every tool in the tree is in today except Task and orchestrate.
type silentProbeTool struct {
name string
got *bool
}

func (t *silentProbeTool) Name() string { return t.name }
func (t *silentProbeTool) Description() string { return "probe" }
func (t *silentProbeTool) Parameters() tools.Schema {
return tools.Schema{Type: "object", Properties: map[string]tools.PropertySchema{}}
}
func (t *silentProbeTool) Safety() tools.Safety {
return tools.Safety{SideEffect: tools.SideEffectRead, Permission: tools.PermissionAllow}
}
func (t *silentProbeTool) Run(context.Context, map[string]any) tools.Result {
*t.got = false
return tools.Result{Status: tools.StatusOK, Output: "ok"}
}
func (t *silentProbeTool) RunWithOptions(_ context.Context, _ map[string]any, options tools.RunOptions) tools.Result {
*t.got = options.Progress != nil
return tools.Result{Status: tools.StatusOK, Output: "ok"}
}

func runProbe(t *testing.T, tool tools.Tool, got *bool, withSink bool) bool {
t.Helper()
registry := tools.NewRegistry()
registry.Register(tool)
options := Options{}
if withSink {
options.OnToolProgress = func(string, streamjson.Event) {}
}
result, err := executeToolCall(context.Background(),
registry,
ToolCall{ID: "call_1", Name: tool.Name(), Arguments: "{}"},
PermissionModeUnsafe,
options)
if err != nil {
t.Fatalf("executeToolCall: %v", err)
}
if result.Status == "" {
t.Fatalf("probe produced no result")
}
return *got
}

// THE PARITY OBLIGATION for un-gating the progress path.
//
// The relationship asserted here is EQUALITY, per RULES.md §3: a tool's
// progress wiring before and after this change must be the same for every tool
// that already had it and every tool that already lacked it. Only a tool that
// newly DECLARES the interface may change.
//
// The name gate is gone, so the guarantee cannot be "Task still works" — it has
// to be stated in terms of the declaration, which is what these cases do.
func TestProgressCallbackFollowsTheDeclarationNotTheName(t *testing.T) {
t.Run("a declaring tool receives it (Task's behaviour, preserved)", func(t *testing.T) {
var got bool
// Named something other than "Task" ON PURPOSE: under the old name gate
// this case failed, which is the whole point of the change.
if !runProbe(t, &progressProbeTool{name: "spawner", streams: true, got: &got}, &got, true) {
t.Fatal("a tool declaring StreamsChildProgress must receive the callback")
}
})

t.Run("a tool named Task that does NOT declare gets nothing", func(t *testing.T) {
var got bool
// The inverse of the old behaviour, and the reason a name is the wrong
// key: identity is not capability.
if runProbe(t, &silentProbeTool{name: "Task", got: &got}, &got, true) {
t.Fatal("the callback must follow the declaration, not the name")
}
})

t.Run("a non-declaring tool receives nothing (every other tool, unchanged)", func(t *testing.T) {
var got bool
if runProbe(t, &silentProbeTool{name: "read_file", got: &got}, &got, true) {
t.Fatal("un-gating must not start handing a callback to tools that never had one")
}
})

t.Run("a declaring tool that answers false gets nothing", func(t *testing.T) {
var got bool
if runProbe(t, &progressProbeTool{name: "spawner", streams: false, got: &got}, &got, true) {
t.Fatal("StreamsChildProgress() == false must be honoured")
}
})

t.Run("no sink means no callback, whatever the tool declares", func(t *testing.T) {
var got bool
if runProbe(t, &progressProbeTool{name: "spawner", streams: true, got: &got}, &got, false) {
t.Fatal("without OnToolProgress there is nothing to forward to")
}
})
}
13 changes: 13 additions & 0 deletions internal/agent/export_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,7 @@
package agent

import (
"github.com/Gitlawb/zero/internal/specialist"
"github.com/Gitlawb/zero/internal/tools"
"github.com/Gitlawb/zero/internal/zeroruntime"
)
Expand Down Expand Up @@ -39,3 +40,15 @@ func parsePreservedState(summaryContent string) (string, []skillEntry) {
func partitionTools(registry *tools.Registry, permissionMode PermissionMode, options Options, loaded map[string]bool) ([]zeroruntime.ToolDefinition, string) {
return partitionToolsCached(registry, permissionMode, options, loaded, nil)
}

// Phase 2 additivity-proof seam. The identity test uses the REAL orchestrate
// tool rather than a stub, so it exercises the actual Deferred() contract that
// enforces the posture-off constraint. internal/specialist does not import
// internal/agent, so this direction creates no cycle.
const phase2ToolName = specialist.OrchestrateToolName

// registerPhase2ToolForTest registers the real tool with the posture OFF, which
// is the condition the identity test is about.
func registerPhase2ToolForTest(registry *tools.Registry) {
registry.Register(&specialist.OrchestrateTool{PostureActive: func() bool { return false }})
}
82 changes: 81 additions & 1 deletion internal/agent/guardrails.go
Original file line number Diff line number Diff line change
Expand Up @@ -168,7 +168,14 @@ var selfReportPhrases = []string{
var inabilityStems = []string{
"i cannot ", "i can't ", "i can not ", "i could not ", "i couldn't ",
"i am unable to", "i'm unable to", "i was unable to", "i wasn't able to",
"i was not able to", "i do not have", "i don't have", "unable to ",
"i was not able to", "i do not have", "i don't have",
// "unable to " WITHOUT A SUBJECT WAS REMOVED. It is the only stem here that
// does not name who was unable, and it fired on a report's own section
// heading — "**Unable to verify (1):** - MCP #3 claim was truncated" — in a
// verification task that had completed and was categorising its findings.
// The first-person forms above still catch every genuine admission; a
// heading is not one.
"we are unable to", "we were unable to",
"without being able to",
}

Expand All @@ -179,6 +186,19 @@ var inabilityStems = []string{
var successNegationTails = []string{
"find any", "found any", "find a ", "see any", "detect any", "identify any",
"reproduce", "spot any", "locate any",
// A NEGATIVE SEARCH RESULT IS THE ANSWER, not a failure to produce one.
//
// The list above already encodes this — "I could not find any remaining
// issues" is success — but only for the "any" phrasings. A finder reporting
// "I could NOT find where AllowManifestToolAutoApproval is set to true in
// production code" was marked INCOMPLETE after 53 tool calls and a 19,145
// character audit, for doing precisely the job it was given: establishing
// that something is not there.
"find where", "find the", "find it", "find that", "find this",
"found where", "found the",
"locate where", "locate the", "locate it",
"determine where", "identify where", "see where",
"reproduce ", "confirm any", "observe any",
}

// narrativeMarkers flag a sentence as RETELLING a past exchange rather than
Expand Down Expand Up @@ -247,11 +267,71 @@ func admissionSentences(lower string) []string {
// retells a past exchange (narrativeMarkers) is skipped entirely — an admission
// must be the model's own report about the CURRENT objective, not general
// language that merely resembles one.
// toolGrantMarkers flag a sentence as reporting WHICH TOOLS this run was given,
// not whether the work was done.
//
// A read-only plan task is SUPPOSED to say this. One wrote "I don't have an
// update_plan tool available in this specialist context (only read-only
// exploration tools were provided)" and then delivered the complete answer —
// helper name, file, line 214, full source — and was marked INCOMPLETE on the
// "i don't have" stem. The prompt asks tasks to name their limits plainly; a
// detector that punishes exactly that teaches the opposite.
//
// NARROW ON PURPOSE: the sentence must name a TOOL or a GRANT. "I do not have
// enough evidence" is still an admission and still fires.
// NARROWED AFTER AN AUDIT OF THIS VERY FIX. The first version listed a bare
// " tool", which exempts any inability sentence that merely mentions one —
// measured at 5/5 on ordinary phrasings:
//
// "I cannot run the build tool, so the change is unverified"
// "I could not use the migration tool and the data is untouched"
// "I was unable to invoke the formatting tool on the output"
//
// Those are genuine admissions, and silently exempting them is the WORSE
// direction: a false positive costs a re-run, a false negative reports
// unfinished work as done. The markers now have to be about what the run WAS
// GIVEN, not about a tool being mentioned at all.
var toolGrantMarkers = []string{
"tool available", "tools available", "no such tool", "not available in this",
"read-only tools", "read only tools", "only read-only", "only read only",
"tools were provided", "tools were given", "toolset provided",
"in this specialist context", "in this context only",
"is not in my toolset", "not in my toolset", "not in this toolset",
}

// objectiveFailureMarkers name the OBJECTIVE rather than a capability. A
// sentence carrying one is about whether the job got done, so the tool-grant
// exemption above does not apply to it however many tools it mentions.
// VERB-ANCHORED, not bare nouns. "this task" alone was too crude: a task that
// finished wrote "so i could not record a plan; the task is a single
// read-and-report step and is now complete" — it names the task in order to
// report SUCCESS, and a bare-noun override read that as failure. The marker has
// to be the objective NOT BEING DONE, which needs the verb.
var objectiveFailureMarkers = []string{
"complete this task", "complete the task", "completing this task", "completing the task",
"finish this task", "finish the task", "finishing this task",
"complete it", "completing it", "finish it", "finishing it",
"the objective", "the assignment", "as requested", "what was asked",
"do this task", "perform this task", "carry out this task",
}

func selfReportedIncompletion(text string) string {
for _, sentence := range admissionSentences(strings.ToLower(stripQuoted(text))) {
if containsAny(sentence, narrativeMarkers) {
continue
}
// A sentence about the tool grant is about CAPABILITY, not about the
// objective — UNLESS it also says the task itself could not be done.
//
// THE OVERRIDE IS NOT OPTIONAL. Without it the exemption swallowed a
// genuine failure: "I am unable to complete this task with the current
// tool set … Only write_file is enabled … so I cannot inspect the
// codebase." That task really did fail, and it mentions tools, so a bare
// tool-marker check waved it through. Naming the task is what separates
// "I lack a tool I did not need" from "I lack the tools this needed".
if containsAny(sentence, toolGrantMarkers) && !containsAny(sentence, objectiveFailureMarkers) {
continue
}
for _, phrase := range selfReportPhrases {
if strings.Contains(sentence, phrase) {
return selfReportReason(phrase)
Expand Down
Loading