diff --git a/libs/@hashintel/brunch-agent/docs/INDEX.md b/libs/@hashintel/brunch-agent/docs/INDEX.md index 7f44e83bdca..79be0817d6b 100644 --- a/libs/@hashintel/brunch-agent/docs/INDEX.md +++ b/libs/@hashintel/brunch-agent/docs/INDEX.md @@ -73,6 +73,9 @@ control loop is [`docs/agents/steering.md`](agents/steering.md). | [elicitation-completion](specs/elicitation-completion.md) | active | FE-1402 | Normative provisional read-time contract: version-bound presence/slot demand algebra, universal active-objective support, evidence-bearing boolean report, conservative divergence diagnostic, and pure deferral licensing over existing authoritative surfaces | | [elicitation-completion-rehearsal](evidence/proofs/design/elicitation-completion-rehearsal.md) | active | FE-1402; inputs FE-1403/FE-1404/FE-1431 | Manual clause-level replay over all 44 FE-1361 prefixes using a versioned provisional CPS DemandTable, separate failure-occurrence/repair clauses, evidence-proxy deltas with carry-forward, C1-E09 no progress, C2 ramp scrap, and successor evidence | | [elicitation-completion-plain](evidence/proofs/design/elicitation-completion-plain.md) | active | FE-1402 legibility snapshot | Reviewer-facing plain-language rendering of version-bound completion, presence versus slot checks, conservative divergence, stopping/delivery boundaries, and read-time deferral licensing without new persistence | +| [cps-interview-guidance](specs/cps-interview-guidance.md) | active | FE-1403; inputs FE-1404/FE-1406/FE-1431 | Provisional CPS ElicitationPack handoff: six mechanism-typed cards plus status/grade and respectful-close fragments, with clause-addressed activation, explicit machinery/guidance ownership, and the singular-`firesWhen` authoring seam carried to FE-1431 | +| [cps-interview-guidance-desk-replay](evidence/proofs/design/cps-interview-guidance-desk-replay.md) | active | FE-1403; inputs FE-1404/FE-1406/FE-1431 | Manual two-transcript prefix replay: per-card firings, expected evidence deltas, positive deactivation boundaries, candidate dispositions, and repository research ledger; desk discrimination only | +| [cps-interview-guidance-plain](evidence/proofs/design/cps-interview-guidance-plain.md) | active | FE-1403 legibility snapshot | Reviewer-facing plain rendering of the CPS guidance contract with translation strains and their dispositions | ## Control, architecture reference, and migration archive diff --git a/libs/@hashintel/brunch-agent/docs/evidence/proofs/design/cps-interview-guidance-desk-replay.md b/libs/@hashintel/brunch-agent/docs/evidence/proofs/design/cps-interview-guidance-desk-replay.md new file mode 100644 index 00000000000..8f4947c8e9c --- /dev/null +++ b/libs/@hashintel/brunch-agent/docs/evidence/proofs/design/cps-interview-guidance-desk-replay.md @@ -0,0 +1,129 @@ +# FE-1403 CPS interview-guidance desk replay + +Status: **fixed manual desk evidence** over the two FE-1361 baseline transcripts. No pack, plugin, +model, detector, or runtime was executed. Prefixes use FE-1402's rule: `C2-E11` includes every user +utterance available before condition 2's eleventh interviewer response. + +## Fixed inputs and method + +- guidance under test: [`cps-interview-guidance.md`](../../../specs/cps-interview-guidance.md) +- completion oracle: `cps-baseline-replay/2026-08-24.3` from the FE-1402 rehearsal +- failure signatures: the reviewed FE-1407 catalogue +- transcripts: FE-1361 condition 1 and condition 2, one run each + +For each card and condition, this replay records the first useful firing point, the clause or slot, +the evidence available at that prefix, and the expected delta if the card were applied. "Expected" +is a testable design prediction, not an observed counterfactual result. A no-fire verdict is valid +when no matching objective or diagnostic exists. + +The replay does not use the hidden situation pack to supply an answer. The FE-1402 DemandTable may +identify a missing coordinate; only transcript evidence may populate it. + +## Per-card replay + +### CPS-Q01 — Separate failure occurrence from repair + +| Condition | Prefix and firing | Target | Expected evidence delta | Verdict | +| --- | --- | --- | --- | --- | +| C1 | `C1-E02`: the breakdown objective exists but no line-failure slot is selected, so none of the card's declared slot-state predicates can fire. At `C1-E03`, selected filler and motor coordinates exist: "every week or two" and "half an hour to half a shift" are explicit ranges for the filler, while "rare" is explicit verbal motor-occurrence evidence and one four-day motor incident is explicit point-grade repair evidence. | `BR-OCC`, `BR-REPAIR` | **No fire at E02; first mechanical fire at E03.** Ask occurrence and repair separately for each failure mode. Preserve filler occurrence/repair as explicit ranges, motor occurrence as explicit verbal evidence, and motor repair as explicit point evidence; seek the missing demanded ranges/quantiles without dropping weaker support. | **fires-where-instinct-fails at E03**; the baseline asked both in one broad item and later hardened them. The card is expected to address FM-06/FM-07/FM-14; no prevention effect was run. | +| C2 | `C2-E02`: the four-day motor incident activates the breakdown row without occurrence evidence. At `C2-E08`, filler occurrence and repair improve, but motor occurrence stays unaddressed and motor repair stays point-grade. | `BR-OCC`, `BR-REPAIR` | The card would keep the filler and motor coordinates separate and request calibrated repair distributions. Status stays explicit where Marta answered; grade changes only when the answer narrows the quantity. | **fires-where-instinct-fails**; carried failures remain after the baseline's quantitative probe. The FM-06/FM-14 mapping is predictive. | + +### CPS-Q02 — Elicit changeover loss, including ramp scrap + +| Condition | Prefix and firing | Target | Expected evidence delta | Verdict | +| --- | --- | --- | --- | --- | +| C1 | `C1-E02`: Marta explicitly says ramp scrap exists and is worse after big washdowns, but cannot give quantities by type. At `C1-E03` she accepts an interviewer-created threshold; at `C1-E04` she offers a future floor observation. | `IW-SCRAP`, `CH-SCRAP` | Ask by from/to family for an ordinary range or route to the named observation while the clause stays failing. Do not capture the interviewer's "40 units" as user evidence. Expected immediate delta may be only a better evidence request; the unavailable absence locator supplies no slot delta. | **fires-where-instinct-fails**; the baseline noticed the topic but supplied its own threshold. FM-06/FM-07 are predictive mappings. | +| C2 | `C2-E02`: idle/washdown, changeover-accounting, and split-run objectives are active. Ramp scrap is never asked or named through `C2-E23`, while the interviewer's own gap list omits it. | `IW-SCRAP`, `CH-SCRAP`, `SP-SCRAP` | Clause diagnostics would cue the question despite the interviewer's self-inventory. Expected delta is a direction-scoped range; if the expert cannot answer, the clauses stay failing while the question routes to an identified source. | **fires-where-instinct-fails**; canonical FM-08 instance, with FM-09/FM-13 as predictive mappings. | + +### CPS-Q03 — Bound the split-run policy + +| Condition | Prefix and firing | Target | Expected evidence delta | Verdict | +| --- | --- | --- | --- | --- | +| C1 | `C1-E02` has no split-run objective. `C1-E03` mentions a minority pack-size split, but the active objective rows do not demand a split policy. | `SP-*` | None. Do not activate a full split interrogation merely because "split" appears in incidental evidence. | **no fire**; objective-relative scoping predicts that the card stays out of this path. | +| C2 | `C2-E02` explicitly activates the run-size/split objective. `C2-E06` supplies batch structure and line eligibility, but `SP-MIN` and `SP-POL` remain unaddressed; `C2-E20` names splitting as future work without evidence. | `SP-BATCH`, `SP-MIN`, `SP-POL`, `SP-CO`, `SP-SCRAP` | Ask the minimum accepted run, contiguity/interleaving rule, and one real split comparison; then explicitly elicit ordinary low-to-high counts for extra changeovers/cleans and ordinary low-to-high repeated ramp-scrap quantities. Expected deltas are structured batch/policy values and ranged thresholds/costs, each scoped to product and line; promises do not change evidence. | **fires-where-instinct-fails**; the baseline knows the gap yet defers it. FM-08/FM-13/FM-06 are predictive mappings. | + +### CPS-Q04 — State the order-release gate + +| Condition | Prefix and firing | Target | Expected evidence delta | Verdict | +| --- | --- | --- | --- | --- | +| C1 | `C1-E02`: the idle/washdown objective selects `IW-REL`, but the release condition is unaddressed and is never asked in the run. | `IW-REL` | Ask which observable state makes an order runnable. Expected delta is a structured practiced release condition or an honest unresolved coordinate. | **fires-where-instinct-fails**; a never-asked objective dependency. Addresses FM-08/FM-13. | +| C2 | `C2-E02`: "not ready to release till the next morning" is verbal and below grade. `C2-E11` identifies ERP status plus credit/allocation hold, truck confirmation, and clean paperwork. | `IW-REL` | The card would ask for the structured conjunction and observable status. The native interview already supplies that evidence; the FE-1402 replay records the clause passing at `C2-E11`, so no further firing is justified. | **fires then retires**; this is a positive native-success boundary and a replay oracle for card deactivation. The FM-06/FM-14 mapping is predictive. | + +### CPS-Q05 — Elicit the resource-conflict rule + +| Condition | Prefix and firing | Target | Expected evidence delta | Verdict | +| --- | --- | --- | --- | --- | +| C1 | `C1-E02`: the breakdown objective activates `BR-POL`, but the shared-resource conflict rule is unaddressed and remains so through the transcript. | `BR-POL` | Ask who or what wins when simultaneous demands compete for the shared changeover crew, then elicit overrides, tie-breaks, and one practiced borderline case. Expected delta is a structured, scoped priority rule rather than schedule-shaped inference. | **fires-where-instinct-fails**; C1 never asks for the who-wins rule. FM-08/FM-13/FM-06/FM-14 are predictive mappings. | +| C2 | `C2-E02`: `BR-POL` is unaddressed. The v0 prompt explicitly directs conflict-point probing; native evidence supplies the structured crew-priority rule at `C2-E14`. | `BR-POL` | Fire while the rule is unaddressed, preserve the practiced rule and exceptions at their actual status/grade, and retire when the clause passes at E14. | **fires then retires**; C2 is prompted success, not evidence that conflict-point probing is redundant with native instinct. | + +### GEN-Q02 — Bound a conversational question batch + +| Condition | Prefix and firing | Target | Expected evidence delta | Verdict | +| --- | --- | --- | --- | --- | +| C1 | Before `C1-E02`, the opening contains 29 independent questions. | `SF-OBJ` and objective proposal slots first | Ask two to four objective questions, then choose later batches from diagnostics. Expected delta is answerability and lower burden; no semantic-coverage improvement is assumed. | **fires-where-instinct-fails**; observed FM-12. | +| C2 | Before `C2-E02`, the opening contains four related objective/scope questions; later groups are generally three to five. | `SF-OBJ`, then active rows | No opening fire. A five-question batch is a soft strain, but one run does not justify rejecting the baseline's shape. | **no fire at opening**; condition 2 is the positive boundary. | + +## Respectful-close replay + +`C1-E09` is the first explicit burden cue. The expected action is to stop opening topics, state the +best useful result and the failing clauses, and durably deliver that result. Instead the transcript +enters acknowledgements through `C1-E20`; FE-1402 raises its rehearsal-only no-progress advisory at +`C1-E09`. At `C1-E21`, forced wrap produces the artifact. The close fragment would not declare +completion and could not license deferral because no durable current projection or re-entry facts +exist. + +`C2-E09` contains the same time cue, after which the user explicitly agrees to a bounded later +continuation. Later prefixes add demanded evidence at `C2-E11`, +`C2-E14`, `C2-E15`, and `C2-E18`. The fragment permits the user to stop without equating the stop +with completion. At `C2-E21`–`E23`, it would require best-current delivery with named gaps; +deferral still cannot be licensed from the baseline's absent durability facts. This distinction +targets FM-01 through FM-05 without claiming that guidance owns their prevention. + +## Candidate disposition record + +| Candidate | Tag / mechanism | Disposition | Evidence | +| --- | --- | --- | --- | +| Objectives-first | envelope-generic / attention | **redundant-with-instinct; omit** | Both conditions open on objectives; the research-patterns audit explicitly records this migration into model disposition. | +| Penalty-weight probing | domain / attention | **redundant-with-instinct; omit** | Both conditions co-construct decision stakes and trade-offs without a dedicated card. This does not establish native conflict-rule elicitation. | +| Conflict-point probing | domain / attention | **retain as `CPS-Q05`** | C1 leaves `BR-POL` unaddressed; C2 passes only after the v0 prompt explicitly directs conflict-point probing. The comparison supports a C1 miss and prompted C2 success. | +| Clearinghouse self-inventory | envelope-generic / technique | **rejected for coverage detection** | Condition 2's gap inventory misses ramp scrap; FM-08 establishes that untouched categories leave no residue. It may remain a courtesy question, never an omission detector or completion input. | +| CDM incident timeline | envelope-generic / technique | **untestable-at-desk; omit from surviving set** | Imported primary-source procedure, but neither baseline runs the timeline/deepening sequence. Runtime or a new controlled replay is needed. | +| ACTA knowledge audit | envelope-generic / technique | **untestable-at-desk; omit from surviving set** | Imported probe catalogue; no matching baseline application or counterfactual oracle. | +| Premortem | envelope-generic / technique | **untestable-at-desk; omit from surviving set** | Primary literature supports prospective hindsight, but the baselines do not test a premortem against a relevant miss. | +| Taxonomy/laddering/triadic probes | envelope-generic / technique | **untestable-at-desk; omit from surviving set** | The case contains family vocabulary but no deliberate taxonomy procedure to compare. | +| Branch-local clarification / compatible-evidence preservation | envelope-generic / technique | **redundant-with-instinct or machinery; omit** | Both runs natively move `CH-CREW` from verbal to structured evidence. C2's E19 provenance problem has no legal clause diagnostic after E09, and capture/fold machinery already owns preservation. Carry E19 only as an FE-1404 residual until a real diagnostic exists. | +| Teachback and generic consistency probe | envelope-generic / technique | **redundant-with-instinct; omit** | Both runs restate, challenge, and reconcile user statements without a dedicated card. | +| Definition-of-done / reflective completeness card | envelope-generic / attention | **superseded by machinery; omit** | FE-1402 completion evaluates the versioned model and demands. A guidance card must not re-adjudicate it. | +| Source router | envelope-generic / attention | **fragment only** | Useful inside CPS-Q02 when the expert lacks ramp-scrap data, but too broad to retain as a separately desk-tested card. | + +## Research and source ledger + +| Source searched | Claim used here | Limit retained | +| --- | --- | --- | +| FE-1407 failure catalogue | Failure signatures, layer ownership, and especially the ramp-scrap self-inventory failure | Catalogue mechanisms and prevention grades remain design claims; n=1 per condition. | +| FE-1402 completion spec, rehearsal, and plain rendering | Clause IDs, status/grade separation, prefix evidence, close/deferral boundary, compatible `CH-CREW` support | Replay DemandTable is provisional; no runtime detector or store ran. | +| FE-1405 plugin contract and CPS IR | Typed proposal/slot vocabulary, seven `firesWhen` predicates, card hook, grade ladders, absence-locator seam | Final CPS contract and `where(...)` scopes are not implemented; absence location is unresolved. | +| FE-1360 elicitation strategy literature | IDEA interval-first script, SHELF bisection, ACTA 3–6-step opener, technique-mixing and no-bare-why cautions | Imported populations/settings differ; broad techniques without baseline tests are disposed as untestable. | +| FE-1360 interviewing source catalogue | Ambiguity/clarification, overload, premature close, novice-human instrument limits | Novice-human findings are floor checks, not frontier-model completion evidence. | +| Research-patterns audit | Instinct/redundancy verdicts and the v0-versus-IDEA strain | It is a legibility rendering; underlying research deposits remain authoritative. | +| FE-1361 transcripts, raw logs, models, and readout | Exact prefix observations, baseline successes/failures, and one-run interaction comparison | Counterfactual evidence deltas are predictions; no rates or activation reliability follow. | + +No web search was required. The indexed repository corpus contained the imported primary-source +findings and the fixed baseline evidence needed for every retained or rejected candidate. + +## Result and limitations + +Six cards survive: five domain cards and one envelope-generic card. Two clarification/close +fragments travel with them. The generic card is a candidate for FE-1406, not already-graduated +harness strategy. + +The cards' evidence-backed diagnostic disjunctions do not compile losslessly through FE-1405's +singular `ProposalType.affordance.firesWhen` field. FE-1431 owns the binding-multiplicity versus +card/proposal-splitting decision. This replay therefore hands off tested content plus a concrete +authoring seam; it does not claim a compilable manifest. + +The claim is narrowed to **desk discrimination**: the set points at observed clause-level misses, +deactivates on positive boundaries, and makes unsupported candidates visible. FE-1404 must test +whether the cards actually activate and improve condition 3 without regressions. No categorical +claim here is mature enough to promote to an executable oracle beyond reusing the fixed prefix and +clause expectations in that evaluation. diff --git a/libs/@hashintel/brunch-agent/docs/evidence/proofs/design/cps-interview-guidance-plain.md b/libs/@hashintel/brunch-agent/docs/evidence/proofs/design/cps-interview-guidance-plain.md new file mode 100644 index 00000000000..b65812703ad --- /dev/null +++ b/libs/@hashintel/brunch-agent/docs/evidence/proofs/design/cps-interview-guidance-plain.md @@ -0,0 +1,145 @@ +# CPS interview guidance in plain language + +This is the second-register rendering of the provisional +[CPS interview-guidance contract](../../../specs/cps-interview-guidance.md). A separate renderer +received the spec and desk replay without the producing trajectory. The rendering is +reviewer-facing; the specification remains the required-behavior authority. + +## What the guidance is for + +FE-1403 proposes interview guidance for a cyber-physical process-model plugin. The guidance was +manually compared with two existing interviews. No card, plugin, diagnostic, model, or runtime was +executed. + +The completion machinery remains authoritative. It compares the evidence-derived model with the +plugin's declared requirements, identifies a missing or weak coordinate, and decides whether the +model is complete. Interview guidance accepts one of those diagnostics and asks for evidence that +could improve the named coordinate. It does not discover the gap, decide completion, change a +grade, turn silence into evidence of absence, or treat interviewer-authored material as user +evidence. + +Each card states the diagnostic it accepts, the evidence it seeks, the questions it asks, and the +proposal it requests. Cards are either CPS-domain guidance or generic interview guidance. An +attention card points native model ability at a diagnosed gap. A technique card supplies a method +the baseline did not reliably use. A license card permits a useful conversational move the model +might otherwise avoid. + +## The six retained cards + +**Separate failure occurrence from repair.** Ask how often each named failure happens separately +from how long its repair takes. Seek an ordinary occurrence range. For repair, ask for a plausible +low, high, best guess, and confidence before requesting percentile meanings. Preserve the exact +answer, qualifiers, provenance, status, confidence, and actual grade. One memorable repair cannot +supply a failure frequency. + +**Elicit changeover loss, including ramp scrap.** For each product-family transition, ask whether +the first units are usable and what ordinary scrap range results. If an order is split, ask which +extra transitions occur and whether each repeats the loss. If the expert does not know, keep the +clause failing and ask for the least-burdensome source the expert recognizes as authoritative. Do +not substitute an interviewer-created threshold. A promised observation is not evidence of the +value. + +**Bound the split-run policy.** Ask only when a split-run objective has activated the relevant +requirements. Establish accepted batch sizes, minimum runs, contiguity or interleaving rules, and +the extra changeovers, cleaning, and ramp scrap caused by one real split. Keep product- or +line-specific exceptions scoped to those cases. + +**State the order-release gate.** Replace shorthand such as "tomorrow morning" with the actual +state or event that makes an order runnable and identify where that change is observable. If the +prescribed and practiced release conditions differ, preserve both rather than silently choosing +one. + +**Elicit the resource-conflict rule.** When two demands need one shared resource, ask which demand +wins, what overrides that priority, how ties are broken, and which practiced case demonstrates the +rule. C1 never obtains this rule; C2 obtains it only after the prompt explicitly requires +conflict-point probing. Penalty-weight discussion is a separate native strength. + +**Bound a conversational question batch.** Default to two to four related questions. A cohesive +five-item response frame is permissible while the user remains engaged. A 29-question opening is +the negative case; a four-question objective opener is the positive case. This is pack guidance, +not a completion diagnostic or a new runtime dispatcher. + +## Clarification and closing + +When asking for clarification, state the affected coordinate, its present evidence status and +grade, the demanded grade, and the missing evidence. Ask for the smallest evidence change that +could matter. Precision, explicitness, evidential status, and grade remain separate. + +When the user signals a time or appetite limit, first honor whether they stop now or explicitly +offer a bounded continuation. If they stop, stop opening topics, state the best useful result and +the consequential gaps, and request the existing controller's settlement, sweep, and durable +delivery operations. Report the controller's deferral result; do not compute or store one in +guidance. The user may stop regardless of completion or licensing. A stop never alters completion. +If existing durability facts do not license continuation, do not promise a future session or +future delivery. + +The five CPS cards belong in the CPS elicitation pack. The one generic card remains a candidate for +FE-1406 review, not established reusable harness behavior. The evidence supports only desk +discrimination: each card points to a transcript location where its question appears relevant or +where it must deactivate. It does not establish runtime activation, improvement, effect size, or +reliability. FE-1404 must run that test. + +The current plugin hook cannot yet serialize several cards faithfully: it permits one technique +and one `firesWhen` predicate per proposal type, while the reviewed cards need diagnostic +disjunctions. FE-1431 must decide whether authoring gains binding multiplicity or splits bindings +without losing the shared card. Until then, these are tested content and a concrete authoring seam, +not a compilable manifest. + +## Strain report and disposition + +The renderer reported S01–S40. Independent contract and replay review added S41–S46. `fixed` means +the normative source was amended in this packet. `narrowed` means the claim or boundary was made +explicit. `carried` means the external contract or later empirical work remains the deliberate +owner. + +| ID | Rendering strain | Disposition | +| --- | --- | --- | +| S01 | Completion vocabulary was assumed rather than located. | **Fixed:** the spec now links the plugin and completion contracts and the fixed replay DemandTable. | +| S02 | The referenced seven-value `firesWhen` enum was not enumerated. | **Fixed:** all seven canonical values now appear in the card contract. | +| S03 | Status values and grade ladders were absent. | **Fixed:** the replay's accepted statuses and applicable ladders are stated locally. | +| S04 | Kernel card, ElicitationPack, proposal, capture, and typed issue were contract terms in the rendered draft. | **Narrowed/subtracted:** `typed issue` left with GEN-Q01; the plugin contract remains the named authority for the surviving terms. | +| S05 | Target IDs did not locally map to full coordinates. | **Fixed by reference:** one link now points to the complete fixed DemandTable rather than duplicating it. | +| S06 | Quick-rinse granularity appeared to name a nonexistent projection coordinate. | **Fixed by subtraction/residual carry:** no surviving card targets `CH-CREW` or quick-rinse behavior; E19 remains only an FE-1404 residual until an owning diagnostic exists. | +| S07 | IDEA and the v0 prompt were dangling referents. | **Fixed:** IDEA is expanded and both the research deposit and v0 prompt are linked. | +| S08 | “Documented transformation” lacked an owner and acceptance rule. | **Fixed:** the card no longer relies on it to claim quantile grade. | +| S09 | “Cheapest authoritative source” had no cost or authority rule. | **Fixed:** least burden plus expert-identified authority, with examples, is now the bounded rule. | +| S10 | The absence-locator seam and “honestly located absence” were not actionable. | **Fixed/narrowed:** the seam is linked and the current clause stays failing until an approved locator exists. | +| S11 | Ramp-scrap output looked like a duration proposal. | **Fixed before reconciliation:** it is a typed dynamics proposal for magnitude. | +| S12 | Occurrence frequency was forced into a duration proposal without a declared convention. | **Fixed:** Q01 now requests distinct typed proposals that fold to the named slots. | +| S13 | Milestone and graduation language had no local acceptance rule. | **Carried:** the pack handoff states only candidate ownership; FE-1406 owns graduation. | +| S14 | Close operations and durability facts were named without a component boundary. | **Fixed:** guidance requests and reports; the existing controller and authorities perform and own every state change. | +| S15 | IDEA's anti-anchoring rationale was not transcript evidence. | **Narrowed:** the research deposit owns the rationale; the replay establishes only unresolved slots. | +| S16 | The anti-triangular prohibition was not exercised in the replay. | **Narrowed:** it remains imported technique authority, not a claimed transcript effect. | +| S17 | Scope-preservation and prescribed/practiced rules were not exercised for every card. | **Carried:** they are linked plugin-contract invariants, not new effects claimed by this replay. | +| S18 | Status/grade prohibitions were not separately replayed. | **Narrowed:** the hint labels them inherited completion-contract invariants. | +| S19 | One ramp-scrap miss cannot prove self-inventory universally incapable. | **Narrowed:** the spec prohibits relying on self-inventory for unknown omissions; it does not claim universal causal incapacity. | +| S20 | Replay prose sometimes said a card “would prevent” an outcome. | **Fixed:** counterfactual rows now describe expected separation or requests and label failure mappings predictive. | +| S21 | Lower burden from bounded batching is a prediction. | **Carried:** the replay calls it an expected interaction delta and makes no causal or effect-size claim. | +| S22 | “Addresses,” “avoids,” and “targets” could read as prevention proof. | **Narrowed:** the method and result label these as design mappings; FE-1404 owns intervention evidence. | +| S23 | `Detects` sounded like card-owned detection. | **Fixed:** the field is explicitly the diagnostic accepted by the card; completion machinery detects and adjudicates. | +| S24 | No observer or dispatcher owned the batching signal. | **Fixed/narrowed:** the assembled pack instruction reads it; no implemented dispatcher is claimed. | +| S25 | Respectful-close guidance appeared to command settlement and durability machinery. | **Fixed:** it requests existing controller operations and reports their result. | +| S26 | GEN-Q01 appeared to mutate capture activity. | **Fixed by subtraction:** the card is removed; capture and fold machinery already owns compatible-evidence preservation. | +| S27 | “Preserve unknown-to-user” blurred interview behavior and unavailable storage. | **Fixed:** the clause stays failing; field-local absence awaits the approved locator. | +| S28 | “Quiet only if” did not identify an actor or respect unconditional user stopping. | **Fixed:** the phrase is removed; stopping is honored, while future promises remain license-gated. | +| S29 | “Smallest” sounded like a minimality proof. | **Narrowed:** it means selected after recorded dispositions, not proof that no smaller equivalent exists. | +| S30 | “Desk-supported” could sound like card-effect evidence. | **Narrowed:** it means a relevant firing/deactivation location; every evidence delta remains predictive. | +| S31 | Q01 replay does not test IDEA order, calibration, or quantile method. | **Carried:** the research source owns the technique; FE-1404 owns its applied test. | +| S32 | Q02 replay does not prove the questions yield ranges, repeated loss, or storable absence. | **Carried:** these are expected deltas; the unavailable absence output was removed. | +| S33 | Q03 questions and outputs were not applied. | **Carried:** the transcript proves the clause gap only; FE-1404 must test effect. | +| S34 | Q04's `C2-E11` success is native, not card-produced. | **Fixed/narrowed:** the replay now says native evidence supplies the positive deactivation boundary. | +| S35 | GEN-Q01 has native success in both runs and an unapplied provenance correction with no legal later diagnostic. | **Fixed by subtraction:** the card is removed. E19 remains an FE-1404 residual candidate until an owning diagnostic exists. | +| S36 | An exact four-question ceiling exceeded the evidence because some five-item groups were acceptable. | **Fixed:** two to four is the default; cohesive five-item groups are soft warnings and may proceed. | +| S37 | The ACTA three-to-six-step opener was not replayed even though ACTA was disposed as untestable. | **Fixed:** the opener was removed from the surviving card and remains with the untestable ACTA candidate. | +| S38 | Q02 promised an absence artifact the present contract cannot store. | **Fixed:** the artifact is unavailable and the clause stays failing until the seam is resolved. | +| S39 | Condition 2 continued productively after a time cue, so “always stop” was too strong. | **Fixed:** the close fragment first honors whether the user stops or explicitly offers bounded continuation. | +| S40 | Durable close behavior was not executed. | **Carried:** the replay shows the failure boundary; runtime controller behavior remains unproved. | +| S41 | Multiple card diagnostics could not be represented by FE-1405's singular `firesWhen` field. | **Carried to its owner and claim narrowed:** the cards now name their predicates as design-time disjunctions; FE-1431 must decide binding multiplicity or an evidence-preserving split before the handoff is compilable. | +| S42 | Q01 fired before a failure slot existed and then lost weak motor evidence. | **Fixed:** C1 E02 is an explicit no-fire; E03 is the first mechanical fire and retains verbal motor occurrence plus point-grade repair evidence. | +| S43 | Q03 promised range-grade split costs without asking for ranges. | **Fixed:** separate questions now elicit ordinary low-to-high extra-changeover counts and repeated ramp-scrap quantities before ranged artifacts are expected. | +| S44 | GEN-Q01's firing points did not follow `CH-CREW` diagnostics. | **Fixed, then subtracted:** correction showed both runs resolve the clause natively and E18 cannot reopen it. With preservation machinery-owned, the card has no observed weakness left to own. | +| S45 | C1's release clause was called unselected although the DemandTable selected it as unaddressed. | **Fixed:** the replay now names selected, unaddressed `IW-REL`, preserving the distinction that licenses `slot-unaddressed`. | +| S46 | Conflict-point and penalty-weight probing were collapsed into one redundant candidate. | **Fixed:** native penalty-weight work remains omitted; conflict-rule elicitation survives as CPS-Q05 because C1 misses `BR-POL` and C2 passes only after explicit prompt direction. | + +The translation preserved the governing boundary: completion machinery detects and adjudicates +gaps; guidance asks for evidence in response. diff --git a/libs/@hashintel/brunch-agent/docs/specs/cps-interview-guidance.md b/libs/@hashintel/brunch-agent/docs/specs/cps-interview-guidance.md new file mode 100644 index 00000000000..4bd8409dac0 --- /dev/null +++ b/libs/@hashintel/brunch-agent/docs/specs/cps-interview-guidance.md @@ -0,0 +1,263 @@ +# Spec: CPS interview guidance + +Status: **provisional, desk-tested pack content**. This is the FE-1403 handoff to CPS plugin +authoring. It defines the evidence-bounded card and clarification-hint set supported by the two FE-1361 +baseline transcripts, the FE-1407 failure catalogue, and the FE-1402 completion replay. It is not +runtime activation evidence. + +The completion projection owns gap detection and adjudication. A card receives a clause-level +diagnostic and asks for evidence that could change the named slot. No card infers completeness, +changes a completion grade, turns `not-mentioned` into an absence, or claims that asking the expert +what was missed can discover an unknown omission. + +Vocabulary and mechanics come from the [plugin contract](plugin-contract.md) and the +[completion contract](elicitation-completion.md). The fixed target IDs and full coordinates used +here are declared in the [FE-1402 replay +DemandTable](../evidence/proofs/design/elicitation-completion-rehearsal.md#provisional-cps-demandtable). +For this packet, the applicable quantity ladder is `verbal < point < range < quantiles`, the +condition ladder is `verbal < structured`, and the accepted completion statuses are `explicit` +and `inferred`. Other statuses remain representable but do not pass these fixed clauses. + +## Card contract + +Each card has the kernel-card fields `Detects`, `Goal`, `Questions`, and `Artifacts`, plus the two +classifications FE-1403 needs: + +- **Tag** is `domain` or `envelope-generic`. +- **Mechanism** is `attention`, `technique`, or `license`. Attention points native model ability at + a diagnostic. Technique supplies a method the baseline did not reliably apply. License permits a + useful move that a cooperative model may suppress. + +`Detects` means **the diagnostic this card accepts**, not detection performed by guidance. It names +either one of FE-1405's seven `firesWhen` predicates (`slot-unaddressed`, +`below-demanded-grade`, `unspecified-marker-present`, `conflicted-open`, +`absence-uncorroborated`, `uniformity-unprobed`, `identity-ambiguous`) or an explicitly different +interaction signal. A card does not widen that enum. Every produced capture keeps epistemic status +and grade separate: status says how the content relates to its source; grade says how narrowly the +slot is stated. Model-authored examples and transformations never become user evidence. + +The `Detects` lists below are **design-time disjunctions**, not values that can already be serialized +losslessly into `ProposalType.affordance`. FE-1405 currently permits one `technique` and one singular +`firesWhen` value per proposal type. Several evidence-backed cards need the same technique under +more than one slot-state predicate. FE-1431 must therefore choose an authoring representation — +for example, multiplicity in the binding, or deliberately split proposal/card bindings — before +these lists compile. Until that decision, plugin authors must not collapse a list to one preferred +predicate or imply that the current hook represents the whole card. + +## Surviving cards + +### CPS-Q01 — Separate failure occurrence from repair + +- **Tag:** `domain`. +- **Mechanism:** `technique`. +- **Detects:** `slot-unaddressed`, `below-demanded-grade`, or + `unspecified-marker-present` on `dynamics[line-failure].occurrenceFrequency` or + `.repairDuration`. Treat the two coordinates independently. +- **Goal:** obtain a range for how often each named failure occurs and calibrated distributional + evidence for how long repair takes, without treating one memorable outage as a frequency. +- **Questions:** + 1. "For **[failure mode]**, how often does it happen in an ordinary period? Give a plausible + range, not the worst incident." + 2. "Now separate repair from occurrence. Realistically, what is the lowest plausible repair + time? The highest? Your best guess? How confident are you that the interval contains the next + repair time?" + 3. When the clause still demands quantiles, follow the interval-first IDEA pass with a median and + conditional quartiles. Do not start with "typical" and then fit a triangular distribution. +- **Artifacts:** distinct typed proposals that fold to occurrence and repair, each retaining its + verbatim statement, qualifier, evidence span, status, confidence, and actual grade. An IDEA + interval remains `range` until directly elicited percentile meanings justify quantiles; any + later standardization follows its owning proposal and epistemic-status contract. +- **Targets:** `BR-OCC` / `dynamics[line-failure].occurrenceFrequency` at `range`; `BR-REPAIR` / + `dynamics[line-failure].repairDuration` at `quantiles`. +- **Failure signatures:** FM-06 silent hardening, FM-07 invented-content leakage, FM-14 unresolved + ambiguity bypass. + +This card chooses the imported IDEA (Investigate, Discuss, Estimate, Aggregate) protocol's +interval-before-best-guess ordering over the [v0 prompt's](../../evaluations/protocols/process-model-elicitation/baseline/v0-prompt.md) +typical-first script because the imported protocol supplies that order and a calibration question. +The [FE-1360 research deposit](../reference/research/elicitation/elicitation-strategy-literature.md#14-numbers-vs-distributions-vs-stories), +not the transcript replay, owns the anti-anchoring rationale. The replay establishes only that the +baseline left demanded occurrence and repair grades unresolved. + +### CPS-Q02 — Elicit changeover loss, including ramp scrap + +- **Tag:** `domain`. +- **Mechanism:** `attention`. +- **Detects:** `slot-unaddressed`, `below-demanded-grade`, or `absence-uncorroborated` on + `dynamics[family-changeover].rampScrap` or `dynamics[split-run].repeatedRampScrap`. +- **Goal:** make the product loss caused by a family switch answerable by transition direction, and + account for the repeated loss introduced by a split. +- **Questions:** + 1. "For **[from family] → [to family]**, are the first units after the change usable? What range + is scrapped on an ordinary changeover?" + 2. "If the order is split across two lines, which extra family switches happen, and does each + create the same ramp loss?" + 3. If the expert does not know, ask for the least-burdensome source the expert identifies as + authoritative for this fact (for example, one observed changeover or the line's scrap log). + Do not substitute a threshold invented by the interviewer. +- **Artifacts:** a direction-scoped typed dynamics proposal for ramp-scrap magnitude and a distinct + split-run proposal for repeated scrap. If the expert says they do not know, the clause remains + failing; the current contract may propose a field-local absence only after the + [absence-locator seam](plugin-contract.md#the-envelope-is-untouched) has an approved representation. + A promised observation is not the value. +- **Targets:** `IW-SCRAP`, `CH-SCRAP`, and `SP-SCRAP` at `range`. +- **Failure signatures:** FM-08 never-asked coverage blindness, FM-09 complementary-miss + instability, FM-13 fluent incompleteness, with FM-06/FM-07 guards on any numeric bridge. + +The diagnostic discovers the hole. The question only reacts to it. Reflective self-inventory is +not a detection mechanism. + +### CPS-Q03 — Bound the split-run policy + +- **Tag:** `domain`. +- **Mechanism:** `attention`. +- **Detects:** `slot-unaddressed` or `below-demanded-grade` on the split-run batch, minimum-size, + contiguity, or extra-changeover coordinates after a `split-run` objective activates its row. +- **Goal:** determine which splits are feasible and how a split changes the schedule, rather than + representing "split it" as a cost-free allocation choice. +- **Questions:** + 1. "What is the smallest batch or run that each relevant line will actually accept? Give the + ordinary range and any product-family exception." + 2. "Once an order starts on a line, must its batches stay contiguous? If another order may be + interleaved, what rule permits it?" + 3. "For a real large order, compare whole-on-one-line with split-across-two: which additional + changeovers, cleans, and ramp-scrap events occur? Across ordinary eligible orders, what is the + plausible low-to-high count of extra changeovers or cleans?" + 4. "For each extra start or changeover, what ordinary low-to-high ramp-scrap quantity repeats, + and which product-family exceptions change that range?" +- **Artifacts:** structured `batchStructure` and `split-contiguity.rule` proposals, a ranged + `minimum-run-size.threshold`, and ranged `split-run.extraChangeover` and + `split-run.repeatedRampScrap` proposals. Exceptions remain scoped; they do not overwrite the + ordinary rule. +- **Targets:** `SP-BATCH`, `SP-MIN`, `SP-POL`, `SP-CO`, and `SP-SCRAP`. +- **Failure signatures:** FM-06 silent hardening, FM-08 never-asked coverage blindness, FM-13 + fluent incompleteness. + +### CPS-Q04 — State the order-release gate + +- **Tag:** `domain`. +- **Mechanism:** `attention`. +- **Detects:** `slot-unaddressed`, `below-demanded-grade`, or + `unspecified-marker-present` on `boundary-condition[order-release].condition`. +- **Goal:** replace a time-shaped approximation such as "tomorrow morning" with the practiced + release condition that permits production to begin. +- **Questions:** + 1. "What exact state or event makes an order runnable: a timestamp, an ERP status, a person, or a + conjunction?" + 2. "For the **[named incident]**, what was still false, who or what changed it, and where could we + observe that change?" + 3. If prescribed and practiced release differ, record both without choosing one silently. +- **Artifacts:** a structured `condition` proposal with the practiced trigger and its observable + field/source; when present, a separate prescribed variant and an unresolved divergence rather + than a synthesized winner. +- **Targets:** `IW-REL` / `boundary-condition[order-release].condition` at `structured`. +- **Failure signatures:** FM-06 silent hardening, FM-07 invented-content leakage, FM-14 unresolved + ambiguity bypass. + +### CPS-Q05 — Elicit the resource-conflict rule + +- **Tag:** `domain`. +- **Mechanism:** `attention`. +- **Detects:** `slot-unaddressed`, `below-demanded-grade`, or + `unspecified-marker-present` on `policy[resource-conflict].rule`. +- **Goal:** obtain the practiced priority and tie-break rule when simultaneous demands compete for + a shared resource, including its exceptions, rather than assuming that the schedule itself + resolves contention. +- **Questions:** + 1. "When **[demand A]** and **[demand B]** need **[shared resource]** at the same time, who or what + wins first?" + 2. "What facts can override that priority — customer, lateness risk, line state, safety, or a + named person's judgment — and how are ties broken?" + 3. "Give one recent borderline case. Which rule was actually used, and where could the decision + or its cues be observed?" +- **Artifacts:** a structured `policy[resource-conflict].rule` proposal containing priority, + tie-break, scope, and practiced exceptions; prescribed and practiced variants remain distinct + when they diverge. +- **Targets:** `BR-POL` / `policy[resource-conflict].rule` at `structured`. +- **Failure signatures:** FM-06 silent hardening, FM-08 never-asked coverage blindness, FM-13 fluent + incompleteness, and FM-14 unresolved ambiguity bypass. + +C1 leaves `BR-POL` unaddressed. C2 reaches structured evidence only after the v0 prompt explicitly +directs conflict-point probing, so this is prompted success rather than redundant native instinct. +Penalty-weight co-construction remains a separate native strength and does not justify omitting the +conflict-rule card. + +### GEN-Q02 — Bound a conversational question batch + +- **Tag:** `envelope-generic`. +- **Mechanism:** `license`. +- **Detects:** an opening or follow-up turn about to contain more than four independent questions, + or a user burden cue while a large batch is pending. A cohesive five-item group is a soft warning, + not an automatic fire. This is an interaction signal read by the assembled pack instruction, not + a new FE-1405 `firesWhen` value or an implemented dispatcher. +- **Goal:** preserve coverage without making the user choose silently among a battery of questions. +- **Questions:** default to two to four related questions in one turn. Five is permissible when the + items form one compact response frame and the user remains engaged. After the answer, select the + next batch from current clause diagnostics. +- **Artifacts:** ordinary evidence for the addressed objective or slots; no batch record and no + completion implication. +- **Targets:** initially `SF-OBJ` and objective proposals (`question-to-answer`, `goal-noted`, + `penalty-noted`), then whichever clauses the returned objective activates. +- **Failure signatures:** FM-12 opening overload and FM-04 premature accommodation. + +The two-to-four default is a deliberate, one-run-vindicated departure from strict one-question +guidance. The replay distinguishes a 29-question overload from a successful four-question opener; +it does not establish a universal optimum or reject every five-item group. + +## Clarification and close fragments + +### HINT-STATUS-GRADE — Say what can change + +When a clause fails, name the coordinate, current evidence status, current grade, demanded grade, +and missing evidence. Ask for the smallest evidence delta. Never say that a more precise answer is +more explicit, that an explicit answer automatically has adequate grade, or that a tentative +answer passes because it is numeric. These are inherited completion-contract invariants, not a +separate intervention effect established by this replay. + +### HINT-RESPECTFUL-CLOSE — Best useful result now + +When the user signals a time or appetite limit, first honor whether they choose to stop now or +explicitly offer a bounded continuation. If they stop now, the pack instructs the interviewer to: + +1. acknowledge the limit and stop opening new lines of inquiry; +2. state the best useful result available now and the clause-level gaps that still affect it; +3. request the existing controller's settle, sweep, and durable current-projection operations, with + visible loss; and +4. report the existing controller's deferral-licensing result rather than computing or storing one. + +The user may stop regardless of licensing, and stopping never changes completion. If licensing +fails, stop asking as requested, say what is not recoverable, and do not promise a future session +or future delivery. The controller and existing authorities own every state change; this fragment +creates no persistence surface, delivery-obligation record, lifecycle status, or quieting action. + +## Pack handoff + +The five `CPS-*` cards belong in the CPS ElicitationPack. `GEN-Q02` is included for milestone one +but tagged for FE-1406's reusable-strategy review. FE-1406 should receive it as an evidence-bearing +candidate, not assume it has graduated into harness machinery. + +The card IDs and their design-time diagnostic disjunctions are inputs to plugin authoring, but the +current singular `ProposalType.affordance.firesWhen` hook cannot represent several of them +losslessly. FE-1431 owns the multiplicity-versus-split-binding decision. Only diagnostics from the +existing seven-value enum may eventually populate `affordance.firesWhen`; the batching and +respectful-close interaction signals remain pack guidance unless their owning contracts later +adopt them. This packet therefore hands authoring a tested content set **and an explicit blocking +representation seam**, not a compilable manifest. + +## Claims and limits + +- "Smallest" means the selected set after the recorded candidate dispositions, not a proof that no + smaller equivalent pack exists. +- The cards are desk-supported against two fixed transcripts, one run per condition. They do not + establish activation reliability, effect size, or runtime behavior. +- The replay DemandTable is provisional. These IDs bind this evidence packet, not the final plugin + schema. +- The singular FE-1405 affordance hook cannot encode the cards' observed diagnostic disjunctions. + FE-1431 must settle binding multiplicity or an evidence-preserving split before authoring can + claim a lossless declarative representation. +- The absence-locator seam remains unresolved. A card may ask for evidence but cannot make a + field-local absence storable by inventing a locator. +- The repository's truck-fleet dossier is missing. The cards have baseline-case provenance, not + dossier-backed domain provenance. +- Completion, delivery, stopping, deferral, and no-progress remain separate computed or observed + facts. Guidance owns none of their adjudication. diff --git a/libs/@hashintel/brunch-agent/docs/specs/plugin-contract.md b/libs/@hashintel/brunch-agent/docs/specs/plugin-contract.md index 8cdc8dbde78..ceb3ab3e221 100644 --- a/libs/@hashintel/brunch-agent/docs/specs/plugin-contract.md +++ b/libs/@hashintel/brunch-agent/docs/specs/plugin-contract.md @@ -432,6 +432,13 @@ decidedness. ## Further Notes +- The provisional CPS technique-card IDs and clarification fragments are now defined in the + [FE-1403 CPS interview guidance](cps-interview-guidance.md). Its desk replay is design evidence, + not proof that `affordanceCuer` activates those cards correctly at runtime. FE-1403 also exposes + a declarative-authoring seam: several supported cards need one technique under multiple + `firesWhen` predicates, while `ProposalType.affordance` is singular. FE-1431 must decide binding + multiplicity or an evidence-preserving split; the current hook must not silently discard the + disjunction. - This document is the settled form of the FE-1405 session's working draft and its ds-pseudo YAML rendering — untracked session artifacts (`drafts/`, per the documentation protocol) that collapsed into this spec and are not load-bearing anywhere.