Skip to content

CM0 Sub-5: execute the compiler over the candidate corpus (terminal validation) #82

Description

@usurobor

Parent: #77. Filing role: κ. Dispatch: .cdd/DISPATCH §5.2 single-session δ-as-γ — γ-axis grade capped at A− per release/SKILL.md §3.8 (state in closeout).
Mode: MCA once Sub-3 (#80) + Sub-4 (#81) are committed. Depends on: #80 and #81. Consumes #76 if available.

Problem

What exists: after Sub-3/Sub-4 there is a compiler + a report, but no evidence they discriminate — that they admit good CMs and reject a bad one.
What is expected: run coh cm-compile + the report over a candidate corpus and prove the compiler rejects the deliberately-bad CM and admits the good ones, and that a real predicate-underspecification surfaces as a reported finding (the #76 line) rather than a silent scalar wobble.
Where they diverge: the wave's central claim ("TSC refuses inconsistent instruments") is unproven until the compiler is exercised against known-good and known-bad inputs.

Impact

The wave's terminal validation. Turns "a CM must compile before it may measure" from an assertion into a demonstrated boundary move (the L7 payoff of #77).

Candidate corpus

  1. CM0 itself (cm-of-cms target) → expect admissible/provisional.
  2. The self-measure CM (methodology target) → expect admissible/provisional.
  3. The factorized-β prereg (as the CM0 Sub-2: prereg as a first-class typed compile object — schemas/prereg.cue + fixtures #79 fixture) → expect its predicate-underspecification to surface as a reported finding, classification not meter-tuning.
  4. One deliberately-bad CM (missing an α organ, OR a prereg whose gate passes without touching its declared axis, OR an LLM that owns enumerate+judge+aggregate) → expect rejected with a named refusal_reason.
  5. One simple good toy CM → expect admissible.

PINNED CONSTRAINT — self-pass is hygiene, never authority

Running CM0 over itself (candidate 1) is a hygiene gate only: a self-pass qualifies CM0 to compete and wins nothing. The validation must NOT read a CM0 self-pass as standing or authority; standing still comes from external anchors, rival CMs, and outcome correlation (existing skills/cm-of-cms/SKILL.md §5/§6). State this explicitly in the result.

#76 dependency (precise)

Source of truth

Claim / surface Canonical source Status
Compiler #80 output (coh cm-compile) Dep
Report #81 output Dep
Candidate CMs targets/{cm-of-cms,methodology}.tsc, skills/cm-of-cms, skills/self-measure Shipped
Factorized-β prereg fixture #79 output Dep
Disagreement input #76 / A3-DISAGREEMENT-REQUEST.md Input if available, else deferred

Acceptance criteria

AC1: good CMs admitted

CM0, the self-measure CM, and the good toy CM compile to admissible/provisional.

  • Oracle: coh cm-compile over each → non-rejected status with an execution_plan.

AC2: bad CM rejected with a named reason

The deliberately-bad CM is rejected with a refusal_reason naming the specific failing stage (missing organ / gameable gate / LLM-owns-all-three).

  • Negative: this is the discrimination proof — a compiler that admits the bad CM fails this sub.

AC3: factorized-β predicate-underspecification surfaces as a finding

Candidate 3 produces a reported finding about semantic predicate under-specification (the A3 line), classified — not a meter re-tune, not a consistency score.

AC4: self-pass recorded as hygiene, not standing (PINNED)

The CM0-over-CM0 result is explicitly labeled hygiene; no standing/authority claim is emitted from the self-pass.

AC5: #76-absent handled honestly

If #76 has not extracted the table, candidate 3's disagreement detail is marked deferred with a pointer to A3-DISAGREEMENT-REQUEST.md; nothing fabricated.

Proof plan

  • Invariant: the compiler discriminates — admits good CMs, rejects a bad one with a named reason, and reports (not scores) semantic ambiguity.
  • Oracle: coh cm-compile + report over the 5 candidates.
  • Positive: candidates 1, 2, 5 admitted. Negative: candidate 4 rejected at the correct stage.
  • Known gap: candidate 3's full disagreement content depends on Implement meter-found semantic ambiguity queue from factorized-β runs #76; absent it, honest-deferred per AC5.

Non-goals

No protocol change; no meter-consistency reopen; no v3.2.5; no standing promotion from any self-pass; no re-run of the terminal factorized-β experiment (used only as a fixture, FAIL unchanged).

Closure — closes the wave

Closeable (and closes #77's validation) when AC1–AC5 met: good CMs admitted, the bad CM rejected with a named reason, the factorized-β predicate-underspecification reported as a classified finding, the self-pass recorded as hygiene, the #76 path handled honestly, non-goals unviolated, γ cap A− stated in closeout.

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1High prioritycddApplied to every issue that runs a CDD cycle

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions