Skip to content

reviews: refutation battery — twelve external passes from six lineages, with fact-checked verdicts - #79

Merged
Nicholas-Keystate merged 1 commit into
Nicholas-Keystate:mainfrom
dhh1128:battery/refutation-passes
Aug 18, 2026
Merged

reviews: refutation battery — twelve external passes from six lineages, with fact-checked verdicts#79
Nicholas-Keystate merged 1 commit into
Nicholas-Keystate:mainfrom
dhh1128:battery/refutation-passes

Conversation

@dhh1128

@dhh1128 dhh1128 commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator

The external adversarial review the battery's README asks for: independent reviewer instances, blind to one another, attacking the two dockets. Six seats, one flagship per lab, none Claude-lineage (the dockets and their source of record are Claude-co-authored, so same-lineage reviewers would share the author's blind spots): GPT 5.6, DeepSeek v4 Pro, GLM 5.2, Qwen 3.8 Max, Kimi K3, Mistral Large. Twelve blind passes, verbatim and sealed under MANIFEST.txt, plus the exact stdin packets for replay and a harness report whose fact-checks are each reproducible without trusting the harness.

Docket 1 — §V.6 SURVIVES. Four lineages NOT REFUTED after prosecuting the rule-1 upgrade route across the V.1 matrix and sweeping well beyond it (Microsoft CCF and Hyperledger Fabric are the strongest new near-misses; witness-cosigned transparency, Cardano Conway, seL4 proof-CI, and the vLEI ecosystem itself were all graded and rejected). The two REFUTED verdicts do not survive fact-checking: Mistral's DNSSEC exhibit quotes three DPS sections that do not exist in the live document (checked 2026-08-14 against the 8th-edition DPS — actual titles reported in the harness report), and DeepSeek's CT upgrade route fails Part 0's own "certifying its cure" clause, with Kimi's inversion closing the door: strict Part 0 grading moves CT's P2 down (the matrix graded issuance acts, not verification runs).

One seam is ruling-grade and owed to 4.3: Part 0's P4 definition does not state the organ requirement that §V.6's "P4 textual but organ-less" dismissal of CT leans on — four seats surfaced it independently. Also convergent, on residue item 7's two named attack surfaces: "died unfunded" over-asserts (drafts' expiration verifiable, funding cause not), and the QVI exhibit conflates issuer with watcher — GLEIF's witness pool proposed as the better existence proof.

Docket 2 — split verdict, sub-claim by sub-claim.

  • Margin (i) falls, five genuine seats of five: KERI's own superseding-recovery doctrine is the prior art, verified against the spec of record (reconciliation rules A0/A1; disputed-branch retention).
  • Margin (ii) survives on novelty — five seats, convergent near-miss lists (Sheng et al. is threshold-counting, not spectral; SybilGuard/SybilLimit is the nearest genuine conductance use; Chuat et al. quantifies gossip detection without expansion pricing). The one refutation attempt had to invent a "Theorem 4.1" to land, and was struck.
  • Margin (iii) falls as stated: the corollary's own example class contains its counterexamples (outpoint double-spends, same-parent blocks, Schneier–Kelsey hash-chained logs), unless "coordinate" absorbs any consumed resource reference — in which case the lemma trends tautological. Four seats, both horns, independently.
  • The detection theorem's "iff" falls twice: GPT's temporal-ordering counterexample (connected comparison graph, no time-respecting path carrying both occupants) and Kimi's supersession-laundering (equivocate briefly, then publish a lawful superseding rotation to everyone; nothing in the six clauses preserves or transmits the superseded evidence — full connectivity, permanent evasion). Both check by hand. The convergently prescribed repair imports KERI's own "first seen, always seen, never unseen" as an explicit evidence-retention law, plus evidence-complete comparison messages and a time-indexed detection-by-t.
  • Completeness falls to the unstated predecessor-chaining premise (one-clause repair); clause 4 carries a one-line internal inconsistency ("distributed across observers" belongs to undetectedness, not the definition); soundness holds — the one attack against it misstated KERI's rule A1 and was struck.

Both struck verdicts are Mistral's; its fabrications were caught because five other lineages graded the same sources honestly — the battery's design working as intended.

Relates to #77 (the §V.6 graduation gate this battery discharges is queued there via the workspace's audit discipline).

🤖 Generated with Claude Code

…lve verdicts, sealed

Twelve blind passes from six non-Claude seats (GPT, DeepSeek, GLM,
Qwen, Kimi, Mistral via OpenRouter), verbatim and manifest-sealed,
with a harness report whose fact-checks are reproducible: docket 1's
§V.6 conjunction SURVIVES (4 NOT REFUTED; one refutation struck for
fabricated DPS quotes, checked against the live document; one graded
failing Part 0 as written) with one seam owed to 4.3 — Part 0's P4
does not state the organ requirement §V.6's CT dismissal leans on.
Docket 2: margin (ii) survives on novelty; margin (i) falls to KERI's
own superseding-recovery doctrine (five seats, spec-verified); margin
(iii) falls as stated (the corollary's example class contains its
counterexamples); the detection iff falls twice (temporal-ordering
counterexample; supersession-laundering) with the repair convergently
prescribed — evidence retention is KERI's first-seen property,
imported explicitly; completeness falls to the unstated chaining
premise; soundness holds.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Daniel Hardman <daniel.hardman@gmail.com>
@Nicholas-Keystate
Nicholas-Keystate merged commit b4ead42 into Nicholas-Keystate:main Aug 18, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants