Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
64 changes: 39 additions & 25 deletions config/sota_v3_lane.json
Original file line number Diff line number Diff line change
Expand Up @@ -6,23 +6,25 @@
"mechanics_change_policy": "Any score-, action-, observation-, or simulator-semantic change after this preregistration requires a new contract fingerprint, invalidates all sota-v3 smoke evidence, and re-locks provider spend.",
"preregistration_status": "provisional-pre-smoke",
"preregistered_at_utc": "2026-07-27T16:58:41Z",
"design_amended_at_utc": "2026-07-28T00:00:00Z",
"design_amended_at_utc": "2026-08-03T00:00:00Z",
"design_amendment": {
"amendment_id": "sota-v3-design-amendment-1",
"amendment_id": "sota-v3-design-amendment-2",
"status": "pre-data",
"supersedes": "docs/run_logs/sota-v3-statistical-design-audit-2026-07-28.md",
"record": "docs/run_logs/sota-v3-design-amendment-2026-07-28.md",
"rationale": "The superseded design powered a superiority claim (+40 above pick-trader) that the frozen sota-v2 evidence contradicts: all eight eligible models trailed the reference by 180-282 score points with a 0.0 seed win rate on every model. No allocation in the 9-20 seed by 1-3 repeat grid reached 0.80 sensitivity power for that claim, and the 20-seed ceiling caps sensitivity power near 0.36 even at unbounded repeats, so the block was structural rather than a budget shortfall.",
"supersedes": "docs/run_logs/sota-v3-design-amendment-2026-07-28.md",
"record": "docs/run_logs/sota-v3-design-amendment-2026-08-03.md",
"rationale": "Owner-directed pre-data cohort update before any smoke or panel evidence exists: the OpenAI anchor moves from GPT-5.6 Luna Pro to plain GPT-5.6 Luna (the originally intended route, now healthy, without the Pro variant's reasoning.mode inconsistency), and DeepSeek V4 Flash 0731 plus Tencent Hy3 join as open-weight anchors on first-party FP8 routes; Thinking Machines Inkling Small was evaluated and found ineligible because no healthy route advertises response_format under the lane's frozen JSON-mode and require-parameters options. A family of ten tightens the Holm first step from 0.00625 to 0.005, so the allocation was reselected with the identical frozen power machinery; 15x1 no longer holds the 0.80 sensitivity floor at family ten and the smallest qualifying allocation is 16 seeds x 1 repeat.",
"changes": [
"Primary claim restated directionally: each registered model trails pick-trader, replicating the frozen sota-v2 finding under the sota-v3 contract.",
"Planning effect changed from +40 to -100 score points, conservative against the -180 to -282 observed on sota-v2.",
"Panel repeats changed from 3 to 1: at a fixed episode budget the candidate-minus-reference lift keeps its full seed component, so repeats shrink only the within-seed noise term and seeds dominate for discrimination.",
"Evaluated seed grid narrowed from 9-20 to 9-16, since the qualifying allocation falls inside it."
"Registered model family grows from eight to ten; Holm family size is now 10 and the first-step threshold is 0.05/10 = 0.005.",
"GPT-5.6 Luna Pro is replaced by plain GPT-5.6 Luna on the first-party OpenAI route, removing the unresolved Pro reasoning.mode inconsistency.",
"DeepSeek V4 Flash 0731 and Tencent Hy3 are added as open-weight anchors, both on first-party FP8 routes.",
"Selected allocation moves from 15 seeds x 1 repeat to 16 seeds x 1 repeat (16 episodes/model): sensitivity power 0.8488, Wilson 95% CI [0.841645, 0.855688]; base power 0.9527, Wilson 95% CI [0.948363, 0.95669].",
"The pending private seed panel is now 16 seeds; its identity remains unfrozen pending separately authorized generation."
],
"unchanged": [
"alpha 0.05, Holm-Bonferroni across a fixed family of eight",
"alpha 0.05, Holm-Bonferroni across the fixed registered family",
"exact two-sided enumeration sign-flip test with the seed as the unit of inference",
"0.80 conservative-sensitivity familywise all-reject power target",
"planning effect -100 score points and the 0.80 conservative-sensitivity familywise all-reject power target",
"historical variance components, sensitivity multipliers, 10000 trials, simulation seed 2026072800",
"no model-to-model tiers or ordinal ranking",
"every execution, spend, and publication authorization remains false"
]
Expand All @@ -40,9 +42,15 @@
"status": "frozen",
"claim_direction": "trails-reference",
"primary_claim": "Every registered model-plus-compact-scaffold system trails deterministic pick-trader on seed-paired mean lift under the frozen sota-v3 mechanics.",
"evaluated_seed_range": [9, 16],
"evaluated_repeat_range": [1, 3],
"holm_family_size": 8,
"evaluated_seed_range": [
9,
16
],
"evaluated_repeat_range": [
1,
3
],
"holm_family_size": 10,
"target_effect_score_points": -100,
"target_effect_basis": "Conservative against the frozen sota-v2 observed lifts of -180 to -282 score points at a 0.0 seed win rate for all eight eligible models.",
"target_familywise_all_reject_power": 0.8,
Expand All @@ -60,24 +68,30 @@
"interaction_variance_floor_as_historical_noise_fraction": 0.1
},
"selected_allocation": {
"seed_count": 15,
"seed_count": 16,
"repeats": 1,
"episodes_per_model": 15,
"minimum_exact_two_sided_sign_flip_p_value": 6.103515625e-05,
"holm_first_step_threshold_at_alpha_0_05": 0.00625,
"episodes_per_model": 16,
"minimum_exact_two_sided_sign_flip_p_value": 3.0517578125e-05,
"holm_first_step_threshold_at_alpha_0_05": 0.005,
"exact_sign_flip_holm_feasible": true,
"base_power_estimate": 0.9461,
"base_power_wilson_ci95": [0.9415, 0.950357],
"sensitivity_power_estimate": 0.8357,
"sensitivity_power_wilson_ci95": [0.828309, 0.842834]
"base_power_estimate": 0.9527,
"base_power_wilson_ci95": [
0.948363,
0.95669
],
"sensitivity_power_estimate": 0.8488,
"sensitivity_power_wilson_ci95": [
0.841645,
0.855688
]
},
"selection_rule": "Smallest allocation with 9-16 seeds and 1-3 repeats whose conservative-sensitivity familywise-power Wilson 95% lower bound is at least 0.80.",
"decision_required": null
},
"seed_panel": {
"status": "pending-authorized-generation",
"name": null,
"count": 15,
"count": 16,
Comment thread
coderabbitai[bot] marked this conversation as resolved.
"sha256": null
},
"reference_agent": "pick-trader",
Expand All @@ -87,7 +101,7 @@
"smoke_manifest": "config/sota_v3_smoke_manifest.json",
"publication_protocol": "config/sota_v3_publication_protocol.json",
"pricing_snapshot": "config/sota_v3_pricing_snapshot.json",
"minimum_headline_models": 8,
"minimum_headline_models": 10,
"reasoning_policy": "catalog-pinned-pending-live-route-verification",
"output_token_cap": 4096,
"output_budget_status": "provisional-pre-smoke-validation",
Expand All @@ -101,7 +115,7 @@
"panel_execution_authorized": false,
"publication_authorized": false,
"blockers": [
"Generate and commit the 15-seed private panel under separate owner authorization; seed identity is unfrozen until its salted commitment hash is recorded.",
"Generate and commit the 16-seed private panel under separate owner authorization; seed identity is unfrozen until its salted commitment hash is recorded.",
"Complete authenticated exact-route and privacy verification for the catalog-selected model cohort before any provider call.",
"Live-verify and pin each exact provider route, endpoint name, supported parameters, privacy policy, and reasoning requirements.",
"Validate the provisional common 4,096-token safety ceiling on one strict smoke per registered route; any predeclared cap-pressure trigger invalidates all v3 smokes and requires one symmetric pre-panel amendment.",
Expand Down
Loading