ADC Pool Level 01 输入绑定契约(contract-only;仅绑定 raw 轴,不授权执行) - #58
Merged
Conversation
…ct-only) Discharges PR #57's BLOCK-02 by binding Level 01's two raw axes and its linkage evidence to externally stored artefacts that already carry their own APPROVE. Nothing is executed: no context, target, pair, disposition or ranking is produced, and no new enumeration run is authorised. Correction to BLOCK-02 as written in PR #57. It said the only context enumeration came from the quarantined 2026-08-04 run. That was inaccurate. The 2026-08-02 enumeration was approved by PR #29, whose record explicitly authorises using its output as input to a further task, and the target-level evidence extraction was approved by PR #31. So an un-quarantined input chain exists, and the correct statement is that the 2026-08-04 artefacts must not be used, not that no usable context exists. I had generalised "quarantined" into "no usable input". Consequence: no new enumeration run is needed, so the previously stated "two contracts and two runs" reduces to one contract and one Level 01 execution. PR #57's approved text is not rewritten; the correction lives here. Three hard constraints measured from the bound artefacts: LOCK-02 admits at most one eligible context. The nine contexts are one canonical_c0 at confidence 0.93, seven not_calibrated derived strategies and one benchmark_only subgroup. PR #28's own prohibition on promoting a derived strategy to canonical fact is inherited as an outcome ceiling, so uncalibrated sources are forced to DEFER, enforced by test. The existing disposition column cannot be inherited as LOCK-01 output. Its values are benchmark/candidate/hold, produced under different criteria, and all 41 rows read not_scored_in_enumeration_run / not_assessed. Linkage evidence is target-level and disease-level: 292 units over 41 genes and 7 dimensions with no context column, all machine_extracted_requires_human_review, and only 2 of 20 expert-review batches complete. Disease-level evidence therefore cannot establish subgroup-specific linkage, and no_known_linkage_after_complete_search is unavailable because the search scope is not closed. DECISION-02 is recorded for adjudication: unreviewed machine-extracted evidence satisfies LOCK-03 existence, but such pairs must carry a review-status column and must not advance to Level 02 until expert review passes. Predicted shape is written in advance so a small pool is not misread as failure: one eligible context, active pool bounded by 41 pairs, 328 pairs on hold, and the real bottleneck is the remaining 18 expert-review batches. 264 tests pass. No change under src/, genmodules/, extensions/, docs/architecture/ or AGENTS.md, and no Gate added. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ection Both blockers accepted and fixed in this PR. A third error surfaced while fixing them and is corrected here too. Blocker 1: LOCK-01 was not bound to executable evidence. The contract only said it must be derived independently, which left how it is derived to the executor, so the eligible/hold/killed split over 41 targets was not reproducible. I first checked whether the approved layer contains protein-level surface evidence at all, since the alternative was to downgrade this PR to binding the raw target axis only. It does: surface_reachability carries 32 transmembrane_segment_count annotations with supporting direction, and 9 not_available. lock_01_derivation now fixes one source, one dimension, one join key and one decisive field, bars seven fields from the derivation including the prior disposition and both gate labels, whitelists a single locator for RETAIN, and defers on missing, conflicting and RNA-derived evidence. not_surface_target and identity_unresolved are declared unavailable this run: the approved layer contains no row asserting that a target is not a surface protein, and no identity-resolution field, so nothing may be excluded. 32 + 9 = 41 with no discretion and no exclusion. Blocker 2: the 36-row to 9-context projection was not frozen. The arithmetic was fine but the semantics were not unique. clinical_context_projection now fixes identity to indication_id alone, derives the context ref from it alone, requires six context-level fields to be constant within a group, collapses the four endpoint roles into unlocked metadata, fixes sort and dedupe keys, and routes conflicts and missing roles to undefined_context DEFER. Tests implement the declared rules over a synthetic fixture and prove order independence, duplicate tolerance, conflict routing and full provenance. Third error, self-caught by measurement: all 41 crc_prevalence units are unknown/not_available, so binding LOCK-03 to that dimension alone would have guaranteed an empty active pool. The source document's Lock 3 already lists existing CRC clinical targeting evidence as a valid linkage form, so binding only expression evidence was a misreading. LOCK-03 now accepts two bases, and a test asserts at least one is non-vacuous and that each basis's vacuity matches its measured count, so an all-empty binding fails loudly instead of silently producing an empty pool. The predicted shape is corrected from a misleading upper bound to exact values with a reconciliation test: 369 raw pairs, 1 eligible context, 32 eligible targets, 32 in the Eligible Universe Index, 27 active and 5 hold. The contract now also requires the result report to state that all 27 active pairs rest on ADC precedent alone, with no CRC expression evidence at all. 277 tests pass. Ten mutations caught and rolled back exactly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…nding only Both scientific-semantic blockers accepted. Together they reduce every possible RETAIN to zero, so this PR no longer authorises executing Level 01. Blocker 1: transmembrane segment count does not establish an extracellularly accessible protein form. It shows transmembrane topology only, not plasma membrane localization, not an extracellular domain, not epitope accessibility, and it does not exclude ER, Golgi or mitochondrial membrane proteins. Calling 32 targets eligible over-promoted the evidence. The source statement itself says it does not prove tumor-cell surface exposure; I read that line and still promoted it. eligible_surface_target now requires plasma-membrane localization, an extracellular domain or topology, and protein-level provenance, all three, all protein-level. Transmembrane-only defers, no annotation defers, and organelle or conflicting localization defers. Measured: plasma membrane, extracellular, localization, signal peptide and GPI all have zero occurrences in the approved units, so eligible becomes 0. Blocker 2: a pan-cancer ADC precedent does not establish CRC linkage. It shows ADC modality precedent only, and indication_fit is a catalog-derived label that holds for all 41 rows, so it cannot substitute for source-level CRC evidence. LB-precedent now requires the source to name a CRC/colorectal indication, or CRC cell line, PDO, PDX or animal-model targeting evidence, with the indication and source locator recorded. Other-cancer precedent is retained as target/modality metadata and holds. Measured source-level CRC units: 0. Measurement trap recorded: all 33 supporting adc_precedent statements contain "CRC", but only inside the disclaimer "precedent does not establish CRC efficacy or a safe therapeutic window". Counting on that substring is what produced last round's false 33/33, and the binding now forbids that criterion. Consequence: target eligible 32 -> 0, Eligible Universe Index 32 -> 0, active 27 -> 0. Executing would emit an empty index and an empty snapshot, which has no candidate value and invites being misread as "screening complete". Per the reviewer's own fallback from the previous round, the binding is downgraded to raw_axis_binding_only with authorises_level_01_execution false, and two gaps are registered: EVGAP-01 needs a controlled target-surface localization extraction, EVGAP-02 needs a controlled CRC-specific linkage extraction. Each needs its own contract-only PR and APPROVE. The "at least one non-vacuous basis" test would have blocked submitting an honest binding, so the invariant is inverted: vacuity is judged on qualifying rather than supporting units, and when nothing qualifies the binding must not authorise execution. 281 tests pass. Ten mutations caught and rolled back exactly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
VAL-B07 still required 32 eligible + 9 hold + 0 killed while the body of the contract, lock_01_derivation.coverage and predicted_result_shape had all moved to 0 eligible + 41 hold + 0 killed. One machine-readable contract stated two mutually exclusive results, so validation had no authoritative value. This was a miss in the round-2 revision. VAL-B07's prose now reads 0 eligible + 41 hold + 0 killed and notes that the 41 hold split as 32 from L1-02, transmembrane topology only, plus 9 from L1-03, no annotation at all. A structured expected_target_eligibility field makes the counts machine-comparable instead of requiring prose parsing, and the rule now declares validates: evidence_insufficient_binding_state with authorises_result_generation false, so it cannot be read as licensing a Level 01 result. Two tests lock this down. One asserts VAL-B07, the derivation coverage and the predicted shape agree key by key, and that the prose does not contradict the structured counts. The other asserts that while execution is unauthorised no validation rule may require eligible targets and every non-vacuous LOCK-01 rule must defer. Swept the repository per the acceptance criterion: no execution or validation requirement anywhere still treats the 32 transmembrane-only targets as eligible. The remaining mentions of 32 are descriptive - the hold breakdown, the contract describing the old mapping as an error, and the round-1 handoff record, which now carries a supersession marker pointing at the round-2 section. 283 tests pass. Five mutations caught and rolled back exactly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
状态
第三轮的唯一阻断(
VAL-B07旧计数残留)已修复。 请复审。fc87a62(前三轮为8f5c85d、8ac045e、6720edb)Ran 283 tests全部通过(基线 251 + 新增 32)scope_of_authorisation: raw_axis_binding_only,authorises_level_01_execution: falselogs/worklog.md这个 PR 是什么
把 Level 01 的 raw clinical context 轴与 raw target 轴绑定到 PR #29/#31 已批准的产物,冻结 LOCK-01 推导、LOCK-02 状态上限、LOCK-03 linkage 依据与 clinical context 投影,并登记两个阻断执行的证据缺口。
不执行 Level 01,也不授权执行 Level 01。
本轮修订:
VAL-B07旧计数残留主体契约、
lock_01_derivation.coverage与predicted_result_shape都已改为0 eligible + 41 hold + 0 killed,但VAL-B07仍写着32 eligible + 9 hold + 0 killed。同一份机器可读契约同时规定两套互斥结果,未来验证无法确定权威值。这是第二轮修订的漏改。VAL-B07散文改为0 eligible + 41 hold + 0 killed,并注明41 hold = 32 (L1-02,仅有跨膜拓扑) + 9 (L1-03,无任何注释)。expected_target_eligibility,使计数可机械比对而不必解析散文;并加validates: evidence_insufficient_binding_state与authorises_result_generation: false,明确该规则用于验证「证据不足导致无法执行」的绑定状态,不代表授权生成 Level 01 结果。test_target_eligibility_counts_agree_in_all_three_places:断言VAL-B07、lock_01_derivation.coverage、predicted_result_shape.target_eligibility三处逐键相等,且散文与结构化计数不矛盾。test_no_validation_rule_asserts_eligible_targets_while_blocked:未授权执行时任何验证规则都不得要求存在 eligible target,且每条非空 LOCK-01 规则必须 DEFER。按验收标准的全文扫描
仓库内已不存在任何把跨膜段对应的 32 个靶点写成 eligible 的执行或验证要求。剩余提到「32」的位置全部是描述性的:绑定 YAML 说明
41 hold = 32 + 9;契约第 81 行把旧映射描述为错误;handoff 第十一节是第一轮历史记录,已加取代标记指向第十二节;第十二节的对照表与变异清单本身就是修订记录。本轮变异检验
5 个全部被捕获后精确回滚,与备份
diff -q一致、恢复OK:把VAL-B07结构化计数改回 32、只把散文改回 32 使其与结构化字段矛盾、把coverage改成与VAL-B07不一致、把predicted_result_shape改成与VAL-B07不一致、声明该规则可授权生成结果。累计四轮的结论
REQUEST_CHANGES×2REQUEST_CHANGES×2REQUEST_CHANGES×1VAL-B07计数对齐,三处一致性上锁三个自查发现的执行者错误也一并记录:LOCK-03 只绑一个 dimension 会保证空池;把「statement 含 CRC」当判据是免责句造成的假阳性;
BLOCK-02把「被隔离」过度推广为「无可用输入」。关键数字(全部实测)
plasma membrane/extracellular/localization/signal peptide/GPI在已批准证据中命中数均为 0;crc_prevalence41 条全为not_available;adc_precedent33 条 supporting 中源级 CRC 单元为 0。两个阻断执行的证据缺口
EVGAP-01LOCK-01EVGAP-02LOCK-03各需独立的 contract-only PR 与
APPROVE,都不在本 PR 授权范围内。授权与不授权
授权: raw 轴绑定 + SHA-256 固定;冻结 LOCK-01 推导、LOCK-02 状态上限、LOCK-03 linkage 依据、clinical context 投影;登记两个证据缺口。
不授权: 执行 Level 01;任何证据抽取或检索运行;任何新枚举;任何筛选排序、Tier、资产推荐、实验建议;任何 Gate 执行或评分;endpoint 锁定;Level 02/03;把被隔离运行的产物重新引入。
前三轮已通过、本轮未改动
context identity 只由
indication_id决定;endpoint 不进 identity 也未锁定;顺序变化与重复行不改投影;冲突与残缺 context 进 DEFER;每个 source row 有 provenance;#53/#54 来源明确禁止;输入 SHA-256 固定;no_known_linkage_after_complete_search不可用;machine-extracted 证据可满足存在性但不得直接晋级 Level 02;跨膜段不再升级为 eligible;泛癌 precedent 不再升级为 CRC linkage;降级本身;EVGAP-01/EVGAP-02的范围定义;不运行 Gate、不评分、不排序;仓库未写入候选、证据或运行结果。架构影响
未触碰
src/、src/contracts/、genmodules/、extensions/、docs/architecture/、AGENTS.md、prompts/;未新增 Gate,未改 45-Gate 拓扑、生命周期、核心对象、envelope、Model 或 Profile。5 个文件全在docs/、tests/、logs/。不消耗 8 月月度架构修复额度。本 PR 不适用
AGENTS.md「审核豁免」,须经 ChatGPTAPPROVE。批准只代表接受 raw 轴绑定与已冻结的判据规则;不授权执行 Level 01,也不批准任何靶点、情境、筛选结果、排序、实验建议或科学结论。