You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
PR Harden SOTA-v3 pre-spend readiness #104 reported 735 passing tests; exact-head and post-merge CI, Ruff, contract, wheel, web, CodeQL, Semgrep, and Pages runs were green.
Current SOTA-v3 contract fingerprint: a523bdfcebe47bbd. Frozen SOTA-v2 evidence remains 558e8f35ea1d66b9.
Still intentionally blocked
Model selection is provisional-blocked; exact route/privacy acceptance records remain unresolved.
The 15-seed private panel is pending authorized generation and has no committed identity.
The smoke manifest is not-started; no real SOTA-v3 artifact exists.
Route preflight, spend, smoke, panel, and publication authorizations are all false.
V3 site-switch strategy and independent external reproduction remain unresolved.
Next gate: inspect and, only with separate owner approval, run the authenticated no-spend route preflight. Private-seed generation and every provider-spend gate remain separate decisions.
This issue remains open as the active SOTA-v3 execution/readiness tracker. The original consultant checklist below is retained as historical context; its early unchecked merge/setup items are superseded by this reconciliation.
Summary
Independent consulting review of GM-Bench (2026-07-25) on feat/contract-economics / HEAD around PR #92. Local gates at review time: 541 tests passed, validate-contract green on the 24-seed canary panel, contract fingerprint sota-v3 / 0a5f0434dca31ac5.
Headline verdict: GM-Bench is an A− evaluation-engineering project and a weak model-ranking instrument. The frozen sota-v2 study is publishable under a narrow claim. sota-v3 is ready to freeze as a contract, not ready to launch as a public panel or headline refresh.
Ordinal model ranking (predeclared and observed failure on sota-v2)
Pure model capability vs scaffold / view / token confounds
Real front-office competence (synthetic league, scripted opponents, decorative morale)
“Near-optimal” headroom (OracleAgent is partial, not a bound)
Durable claim (keep this): Under the frozen phase-one protocol, all eight eligible systems trailed pick-trader on every seed; no model ordering is justified.
Merge recommendation: yes (after confirming no in-flight model checkpoints). This closes a real validity hole: bad contracts used to be free to erase.
Strengths:
Bounded, published dead cap
Incumbent extensions with structural sign-and-extend bar
Salaries + cap inflate together (lock-in, not league squeeze)
Term premium + loyalty discount with INC5 > FA1 dominance guard
Canaries that can fail (paired t ≥ 2.0 on 24-seed panel)
Site builder freezes v2 baselines from artifacts (prevents v3 sim scores under v2 labels)
Calibration honesty (24-seed panel): only pick-trader > value and shrewd > value asserted. Top four (shrewd / strategic / scaffold-view / pick-trader) sit within ~10 points — order not established. Do not invent a new ranked reference ladder. Do not compare means to the v2 411.619 bar (different fingerprint).
Residual defects from review
Sev
Finding
P2
Reference agents understate multi-year dead-cap when projecting cap room (agents.py)
P2
Opponent extensions skip reservation/walkaway that bind the user
P2
Persistent-session multi-round path may not refresh full observation after queries
P3
extension_quotes includes 1-year term while extend_contract rejects years < 2
P3
Morale is published/updated but never affects strength/development/score
P3
Compact view uses top_roster (economics fields still pass through for kept players)
Render-test: score surfaces must show ci95 + tokens_per_decision
Explicit non-goals (for now)
Do not authorize another paid model panel until scaffold-view + v3 lane pre-registration + site framing fixes land
Do not claim ordinal baseline order among shrewd / strategic / pick-trader / scaffold-view under the new contract
Do not market “v3 leaderboard launch” as the scientific story — the story is measurement hardening + the honest v2 case study
Recommended product pivot
Lean the public narrative toward: reproducible agent-evaluation toolkit + case study (frontier scaffolds lose to a transparent heuristic under a frozen protocol). The moat is process quality (retract, freeze, gate), not “MMLU for GMs.”
Status reconciliation — 2026-08-01
Current
main:a0fdec5493eaf5f702e71c519910e03f39e727f5(through PR #104).Completed engineering work
a523bdfcebe47bbd. Frozen SOTA-v2 evidence remains558e8f35ea1d66b9.Still intentionally blocked
provisional-blocked; exact route/privacy acceptance records remain unresolved.not-started; no real SOTA-v3 artifact exists.Next gate: inspect and, only with separate owner approval, run the authenticated no-spend route preflight. Private-seed generation and every provider-spend gate remain separate decisions.
This issue remains open as the active SOTA-v3 execution/readiness tracker. The original consultant checklist below is retained as historical context; its early unchecked merge/setup items are superseded by this reconciliation.
Summary
Independent consulting review of GM-Bench (2026-07-25) on
feat/contract-economics/ HEAD around PR #92. Local gates at review time: 541 tests passed,validate-contractgreen on the 24-seed canary panel, contract fingerprintsota-v3/0a5f0434dca31ac5.Headline verdict: GM-Bench is an A− evaluation-engineering project and a weak model-ranking instrument. The frozen
sota-v2study is publishable under a narrow claim.sota-v3is ready to freeze as a contract, not ready to launch as a public panel or headline refresh.Related: #84 (framing), #89 (simulator depth), #91 (unreachable release — largely addressed in #92), PR #92 (contract economics).
Scientific assessment
What the bench measures well
strategy_score/protocol_penaltysplit)What it does not measure
OracleAgentis partial, not a bound)Durable claim (keep this): Under the frozen phase-one protocol, all eight eligible systems trailed
pick-traderon every seed; no model ordering is justified.Current state map
mainscore_components,scaffold-viewregisteredfeat/contract-economicsContract economics (PR #92) — assessment
Merge recommendation: yes (after confirming no in-flight model checkpoints). This closes a real validity hole: bad contracts used to be free to erase.
Strengths:
Calibration honesty (24-seed panel): only
pick-trader > valueandshrewd > valueasserted. Top four (shrewd / strategic / scaffold-view / pick-trader) sit within ~10 points — order not established. Do not invent a new ranked reference ladder. Do not compare means to the v2411.619bar (different fingerprint).Residual defects from review
agents.py)extension_quotesincludes 1-year term whileextend_contractrejects years < 2top_roster(economics fields still pass through for kept players)Priority backlog
P0 — This week / before any paid v3 spend
data/model_checkpoints/scaffold-viewon seeds 11–18 at fingerprint0a5f0434dca31ac5(free, deterministic; unblocks observation-asymmetry measurement from Close the gap between sota-v2 evidence and framing: scoring, observation asymmetry, scaffolds, power, and site claims #84)Analysis.tsxranked ladder: add CIs or unranked/tiered layouttokens_per_decision(and input/output split where available) on score surfacesconfig/sota_v3_lane.json, registry, smoke manifest)validate-result --policy sota-v3when first v3 artifacts landrequire_strict_fallbackmutation coverage; CapHoard seed-level assertion;extend_contractsimulator acceptance in prompt conformanceP1 — Credibility and interpretability
PUBLISH_READINESS.md)score_components+scripts/weight_sensitivity.pyextension_quotes(or accept 1-year extensions consistently)P2 — Environment / construct validity
valueonlyP3 — Polish
CITATION.cff+ third-party result issue templateci95+tokens_per_decisionExplicit non-goals (for now)
Recommended product pivot
Lean the public narrative toward: reproducible agent-evaluation toolkit + case study (frontier scaffolds lose to a transparent heuristic under a frozen protocol). The moat is process quality (retract, freeze, gate), not “MMLU for GMs.”
Verification checklist (carry forward)
After first v3 artifact:
python3 -m gm_bench validate-result <artifact>.json --policy sota-v3