Skip to content

v3 readiness program: consultant audit findings and execution plan #93

Description

@nedcut

Status reconciliation — 2026-08-01

Current main: a0fdec5493eaf5f702e71c519910e03f39e727f5 (through PR #104).

Completed engineering work

Still intentionally blocked

  • Model selection is provisional-blocked; exact route/privacy acceptance records remain unresolved.
  • The 15-seed private panel is pending authorized generation and has no committed identity.
  • The smoke manifest is not-started; no real SOTA-v3 artifact exists.
  • Route preflight, spend, smoke, panel, and publication authorizations are all false.
  • V3 site-switch strategy and independent external reproduction remain unresolved.

Next gate: inspect and, only with separate owner approval, run the authenticated no-spend route preflight. Private-seed generation and every provider-spend gate remain separate decisions.

This issue remains open as the active SOTA-v3 execution/readiness tracker. The original consultant checklist below is retained as historical context; its early unchecked merge/setup items are superseded by this reconciliation.


Summary

Independent consulting review of GM-Bench (2026-07-25) on feat/contract-economics / HEAD around PR #92. Local gates at review time: 541 tests passed, validate-contract green on the 24-seed canary panel, contract fingerprint sota-v3 / 0a5f0434dca31ac5.

Headline verdict: GM-Bench is an A− evaluation-engineering project and a weak model-ranking instrument. The frozen sota-v2 study is publishable under a narrow claim. sota-v3 is ready to freeze as a contract, not ready to launch as a public panel or headline refresh.

Related: #84 (framing), #89 (simulator depth), #91 (unreachable release — largely addressed in #92), PR #92 (contract economics).


Scientific assessment

Dimension Grade Why
Reproducibility & integrity A− Contract fingerprints, pre-registration, machine gates, v1 retraction, #85 P0 fixes
Statistical hygiene B+ Seed-paired analysis, Holm, tiers — n=8 structurally underpowered
Model discrimination D One overlapping tier; MDD (~62) ≫ partial-oracle gap (~19.5 on v2)
Construct validity (“GM skill”) C+ Real multi-season decisions; hand-scored bar; white-box baselines; scaffold asymmetry
Portfolio / process story A Rare honesty: withdraw bad rankings, freeze contracts, refuse to launder confounds

What the bench measures well

  • Whether a model + frozen scaffold beats a transparent scripted reference under a disclosed lane
  • Protocol competence vs strategy (strategy_score / protocol_penalty split)
  • Anti-gaming: exploit / pick-hoard / cap-hoard / accept-everything with paired-t gates

What it does not measure

  • Ordinal model ranking (predeclared and observed failure on sota-v2)
  • Pure model capability vs scaffold / view / token confounds
  • Real front-office competence (synthetic league, scripted opponents, decorative morale)
  • “Near-optimal” headroom (OracleAgent is partial, not a bound)

Durable claim (keep this): Under the frozen phase-one protocol, all eight eligible systems trailed pick-trader on every seed; no model ordering is justified.


Current state map

Layer Status
Public site / blog / release Locked to sota-v2 — correct
main sota-v3 validator (#85/#88): strict fallback, score_components, scaffold-view registered
PR #92 feat/contract-economics New simulator semantics + honest canaries + site baseline freeze
Paid v3 model panel Does not exist
External reproduction Missing
flowchart LR
  v2["sota-v2 phase-one\nfrozen, published"] --> integrity["#85 P0 integrity\nsota-v3 validator"]
  integrity --> econ["PR #92 contract economics"]
  econ --> freeze["Merge + freeze contract"]
  freeze --> measure["scaffold-view + framing"]
  measure --> lane["Pre-register v3 lane"]
  lane --> panel["Paid v3 panel\nONLY if needed"]
Loading

Contract economics (PR #92) — assessment

Merge recommendation: yes (after confirming no in-flight model checkpoints). This closes a real validity hole: bad contracts used to be free to erase.

Strengths:

  • Bounded, published dead cap
  • Incumbent extensions with structural sign-and-extend bar
  • Salaries + cap inflate together (lock-in, not league squeeze)
  • Term premium + loyalty discount with INC5 > FA1 dominance guard
  • Canaries that can fail (paired t ≥ 2.0 on 24-seed panel)
  • Site builder freezes v2 baselines from artifacts (prevents v3 sim scores under v2 labels)

Calibration honesty (24-seed panel): only pick-trader > value and shrewd > value asserted. Top four (shrewd / strategic / scaffold-view / pick-trader) sit within ~10 points — order not established. Do not invent a new ranked reference ladder. Do not compare means to the v2 411.619 bar (different fingerprint).

Residual defects from review

Sev Finding
P2 Reference agents understate multi-year dead-cap when projecting cap room (agents.py)
P2 Opponent extensions skip reservation/walkaway that bind the user
P2 Persistent-session multi-round path may not refresh full observation after queries
P3 extension_quotes includes 1-year term while extend_contract rejects years < 2
P3 Morale is published/updated but never affects strength/development/score
P3 Compact view uses top_roster (economics fields still pass through for kept players)

Priority backlog

P0 — This week / before any paid v3 spend

  • Merge PR Contract economics: dead cap, incumbent extensions, and a canary panel that can fail #92 after confirming no live checkpoints under data/model_checkpoints/
  • Run scaffold-view on seeds 11–18 at fingerprint 0a5f0434dca31ac5 (free, deterministic; unblocks observation-asymmetry measurement from Close the gap between sota-v2 evidence and framing: scoring, observation asymmetry, scaffolds, power, and site claims #84)
  • Site claim integrity while the public page still sells v2:
    • Fix Analysis.tsx ranked ladder: add CIs or unranked/tiered layout
    • Render tokens_per_decision (and input/output split where available) on score surfaces
    • Relabel “oracle ceiling” → “partial oracle reference”
    • Add public-panel adaptation / contamination caveat on the site (mirror blog)
  • Pre-register v3 publication lane before spend (config/sota_v3_lane.json, registry, smoke manifest)
  • CI: add validate-result --policy sota-v3 when first v3 artifacts land
  • Tests: smoke require_strict_fallback mutation coverage; CapHoard seed-level assertion; extend_contract simulator acceptance in prompt conformance

P1 — Credibility and interpretability

  • Independent clean-clone reproduction (external validation still Missing in PUBLISH_READINESS.md)
  • Weight sensitivity on future model rows via score_components + scripts/weight_sensitivity.py
  • Outcome-only secondary view (titles / playoff rounds / wins) alongside composite
  • Decide v3 site strategy: historical v2 page vs current v3 page
  • Propagate Holm / tier caveats to Analysis + sortable table
  • Fix multi-year dead-cap projection in shrewd/strategic release accounting
  • Drop 1-year from extension_quotes (or accept 1-year extensions consistently)
  • Soften opponent marketing language (“scripted opponent offices”)

P2 — Environment / construct validity

P3 — Polish

  • Remove residual README “MVP” wording
  • CITATION.cff + third-party result issue template
  • Architecture / evaluation-flow diagram
  • Package Grok diagnostic raw in release assets
  • Render-test: score surfaces must show ci95 + tokens_per_decision

Explicit non-goals (for now)

  • Do not authorize another paid model panel until scaffold-view + v3 lane pre-registration + site framing fixes land
  • Do not claim ordinal baseline order among shrewd / strategic / pick-trader / scaffold-view under the new contract
  • Do not market “v3 leaderboard launch” as the scientific story — the story is measurement hardening + the honest v2 case study

Recommended product pivot

Lean the public narrative toward: reproducible agent-evaluation toolkit + case study (frontier scaffolds lose to a transparent heuristic under a frozen protocol). The moat is process quality (retract, freeze, gate), not “MMLU for GMs.”


Verification checklist (carry forward)

python3 -m pytest -q
python3 -m ruff format --check gm_bench examples tests
python3 -m ruff check gm_bench examples tests
python3 -m gm_bench validate-contract
# frozen v2 evidence
find results/leaderboard -maxdepth 1 -name '*.json' -print0 | \
  xargs -0 -I{} python3 -m gm_bench validate-result {} --policy sota-v2
python3 web/scripts/build_leaderboard.py
git diff --exit-code -- web/src/data/leaderboard.json

After first v3 artifact: python3 -m gm_bench validate-result <artifact>.json --policy sota-v3

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions