Skip to content
Merged
112 changes: 65 additions & 47 deletions config/sota_v3_models.json
Original file line number Diff line number Diff line change
Expand Up @@ -11,14 +11,14 @@
"repeats": 1,
"selection_status": "route-preflight-ready",
"selection_frozen_at_utc": null,
"selection_revision": "2026-08-03-public-catalog-cohort-v2",
"selection_revision": "2026-08-04-public-catalog-lineup-refresh-v3",
"catalog_snapshot_status": "frozen-public-metadata-only",
"catalog_checked_at_utc": "2026-08-03T15:53:59Z",
"catalog_checked_at_utc": "2026-08-04T16:00:40Z",
"catalog_sources": [
"https://openrouter.ai/api/v1/models",
"https://openrouter.ai/api/v1/models/{model_id}/endpoints"
],
"selection_policy": "Ten-model pre-data cohort selected from the 2026-08-03 public OpenRouter catalog, superseding the eight-model 2026-07-28 cohort before any smoke or panel evidence exists. GPT-5.6 Luna replaces GPT-5.6 Luna Pro: the plain Luna route that was unhealthy at the prior snapshot has recovered, and the Pro variant carried an unresolved reasoning.mode inconsistency between its public description and the structured catalog. DeepSeek V4 Flash 0731 and Tencent Hy3 (both on first-party FP8 routes) are added as open-weight anchors. Thinking Machines Inkling Small was evaluated for the tenth slot and found ineligible at this snapshot: no healthy route advertises the response_format parameter the lane's frozen JSON-mode and require-parameters options demand. Exact public endpoint metadata and prices are pinned below, but the registry remains provisional-blocked because public metadata does not prove authenticated exact-route access or provider privacy and retention behavior.",
"selection_policy": "Ten-model pre-data cohort originally selected from the 2026-08-03 public OpenRouter catalog and refreshed on 2026-08-04 before any smoke or panel evidence existed. The refresh replaces Qwen 3.7 Plus with Qwen 3.8 Max and substitutes the same-model MiniMax M3 and DeepSeek V4 Flash 0731 slots from minimax/fp8 and deepseek/fp8 onto deepinfra/fp8 and cloudflare/fp8 under docs/ROUTE_SUBSTITUTION_POLICY.md. Cohort size, Holm family, and the 16x1 allocation are unchanged. Exact public endpoint metadata and undiscounted list prices are pinned below, but the registry remains route-preflight-ready rather than frozen because authenticated exact-route access, parameter behavior, and provider privacy and retention acceptance remain unresolved.",
"models": [
{
"id": "openrouter-gpt-5.6-luna-openai",
Expand All @@ -32,7 +32,7 @@
"endpoint_tag": "openai",
"endpoint_name": "OpenAI | openai/gpt-5.6-luna-20260709",
"catalog_route_status": 0,
"catalog_uptime_last_30m": 99.31771054990628,
"catalog_uptime_last_30m": 99.62251080512623,
"catalog_supported_parameters": [
"reasoning",
"include_reasoning",
Expand Down Expand Up @@ -79,7 +79,7 @@
"endpoint_tag": "amazon-bedrock/global",
"endpoint_name": "Amazon Bedrock | anthropic/claude-sonnet-5-20260630",
"catalog_route_status": 0,
"catalog_uptime_last_30m": 99.56639566395664,
"catalog_uptime_last_30m": 99.62314342717801,
"catalog_supported_parameters": [
"reasoning",
"include_reasoning",
Expand Down Expand Up @@ -126,7 +126,7 @@
"endpoint_tag": "google-ai-studio",
"endpoint_name": "Google AI Studio | google/gemini-3.6-flash-20260721",
"catalog_route_status": 0,
"catalog_uptime_last_30m": 99.78925184404636,
"catalog_uptime_last_30m": 97.20506031185643,
"catalog_supported_parameters": [
"reasoning",
"include_reasoning",
Expand Down Expand Up @@ -172,7 +172,7 @@
"endpoint_tag": "xai/zdr",
"endpoint_name": "xAI | x-ai/grok-4.5-20260708",
"catalog_route_status": 0,
"catalog_uptime_last_30m": 100.0,
"catalog_uptime_last_30m": 99.97348886532343,
"catalog_supported_parameters": [
"reasoning",
"include_reasoning",
Expand Down Expand Up @@ -222,7 +222,7 @@
"endpoint_tag": "novita/fp8",
"endpoint_name": "Novita | z-ai/glm-5.2-20260616",
"catalog_route_status": 0,
"catalog_uptime_last_30m": 99.70362506358791,
"catalog_uptime_last_30m": 98.38854073410921,
"catalog_supported_parameters": [
"reasoning",
"include_reasoning",
Expand Down Expand Up @@ -260,27 +260,35 @@
"role": "Z.ai open-weight anchor on the previously smoke-validated Novita FP8 route"
},
{
"id": "openrouter-minimax-m3-minimax",
"id": "openrouter-minimax-m3-deepinfra",
"provider": "openrouter",
"model": "minimax/minimax-m3",
"canonical_slug": "minimax/minimax-m3-20260531",
"transport": "gateway-api",
"cohort": "open-weight",
"upstream_provider": "Minimax",
"upstream_provider_slug": "minimax/fp8",
"endpoint_tag": "minimax/fp8",
"endpoint_name": "Minimax | minimax/minimax-m3-20260531",
"upstream_provider": "DeepInfra",
"upstream_provider_slug": "deepinfra/fp8",
"endpoint_tag": "deepinfra/fp8",
"endpoint_name": "DeepInfra | minimax/minimax-m3-20260531",
"catalog_route_status": 0,
"catalog_uptime_last_30m": 99.45594322885867,
"catalog_uptime_last_30m": 99.82795698924731,
"catalog_supported_parameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"stop",
"frequency_penalty",
"presence_penalty",
"repetition_penalty",
"top_k",
"seed",
"min_p",
"response_format",
"tool_choice",
"tools"
"logit_bias",
"tools",
"tool_choice"
],
"reasoning_policy": "disabled",
"reasoning_effort": null,
Expand All @@ -293,21 +301,21 @@
"absent_options": [
"OPENROUTER_REASONING_EFFORT"
],
"role": "MiniMax open-weight anchor on the first-party FP8 route"
"role": "MiniMax open-weight anchor, substituted from the first-party `minimax/fp8` route onto `deepinfra/fp8` on 2026-08-04 after `minimax/fp8` was deranked to status -2 (30m uptime 94.59%). Same FP8 quantization and identical published rates; highest 24h uptime (99.63%) among eligible FP8 routes. Per docs/ROUTE_SUBSTITUTION_POLICY.md."
},
{
"id": "openrouter-qwen3.7-plus-alibaba",
"id": "openrouter-qwen3.8-max-alibaba",
"provider": "openrouter",
"model": "qwen/qwen3.7-plus",
"canonical_slug": "qwen/qwen3.7-plus-20260602",
"model": "qwen/qwen3.8-max",
"canonical_slug": "qwen/qwen3.8-max-20260803",
"transport": "gateway-api",
"cohort": "open-weight",
"upstream_provider": "Alibaba",
"upstream_provider_slug": "alibaba/fp8",
"endpoint_tag": "alibaba/fp8",
"endpoint_name": "Alibaba | qwen/qwen3.7-plus-20260602",
"upstream_provider_slug": "alibaba",
"endpoint_tag": "alibaba",
"endpoint_name": "Alibaba | qwen/qwen3.8-max-20260803",
"catalog_route_status": 0,
"catalog_uptime_last_30m": 99.99335742374322,
"catalog_uptime_last_30m": 100,
"catalog_supported_parameters": [
"reasoning",
"include_reasoning",
Expand All @@ -317,11 +325,15 @@
"seed",
"presence_penalty",
"response_format",
"logprobs",
"top_logprobs",
"tools",
"tool_choice",
"structured_outputs"
"structured_outputs",
"logprobs",
"top_logprobs",
"top_k",
"frequency_penalty",
"stop",
"reasoning_effort"
],
"reasoning_policy": "disabled",
"reasoning_effort": null,
Expand All @@ -335,7 +347,7 @@
"absent_options": [
"OPENROUTER_REASONING_EFFORT"
],
"role": "Qwen open-weight frontier anchor"
"role": "Qwen frontier anchor on the first-party Alibaba route"
},
{
"id": "openrouter-mistral-medium-3.5-mistral",
Expand All @@ -349,7 +361,7 @@
"endpoint_tag": "mistral",
"endpoint_name": "Mistral | mistralai/mistral-medium-3.5-20260430",
"catalog_route_status": 0,
"catalog_uptime_last_30m": 100.0,
"catalog_uptime_last_30m": 100,
"catalog_supported_parameters": [
"reasoning",
"include_reasoning",
Expand Down Expand Up @@ -385,32 +397,38 @@
"role": "Mistral European open-weight anchor"
},
{
"id": "openrouter-deepseek-v4-flash-0731-deepseek",
"id": "openrouter-deepseek-v4-flash-0731-cloudflare",
"provider": "openrouter",
"model": "deepseek/deepseek-v4-flash-0731",
"canonical_slug": "deepseek/deepseek-v4-flash-20260731",
"transport": "gateway-api",
"cohort": "open-weight",
"upstream_provider": "DeepSeek",
"upstream_provider_slug": "deepseek/fp8",
"endpoint_tag": "deepseek/fp8",
"endpoint_name": "DeepSeek | deepseek/deepseek-v4-flash-20260731",
"upstream_provider": "Cloudflare",
"upstream_provider_slug": "cloudflare/fp8",
"endpoint_tag": "cloudflare/fp8",
"endpoint_name": "Cloudflare | deepseek/deepseek-v4-flash-20260731",
"catalog_route_status": 0,
"catalog_uptime_last_30m": 99.99164403593065,
"catalog_uptime_last_30m": 99.28610653487095,
"catalog_supported_parameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"stop",
"top_k",
"seed",
"repetition_penalty",
"frequency_penalty",
"presence_penalty",
"logprobs",
"top_logprobs",
"min_p",
"stop",
"logit_bias",
"response_format",
"structured_outputs",
"tools",
"tool_choice",
"response_format",
"logprobs",
"top_logprobs",
"reasoning_effort"
],
"reasoning_policy": "disabled",
Expand All @@ -431,7 +449,7 @@
"absent_options": [
"OPENROUTER_REASONING_EFFORT"
],
"role": "DeepSeek open-weight anchor on the first-party FP8 route"
"role": "DeepSeek open-weight anchor, substituted from the first-party `deepseek/fp8` route onto `cloudflare/fp8` on 2026-08-04 after `deepseek/fp8` was deranked to status -5 (30m uptime 78.93%). Same FP8 quantization and identical published rates; selected on 24h uptime (99.75%, best of sixteen endpoints), not spot health. Per docs/ROUTE_SUBSTITUTION_POLICY.md."
},
{
"id": "openrouter-hy3-tencent",
Expand All @@ -445,7 +463,7 @@
"endpoint_tag": "tencent/fp8",
"endpoint_name": "Tencent | tencent/hy3-20260706",
"catalog_route_status": 0,
"catalog_uptime_last_30m": 99.88329118848473,
"catalog_uptime_last_30m": 99.80510276399717,
"catalog_supported_parameters": [
"reasoning",
"include_reasoning",
Expand Down Expand Up @@ -563,7 +581,7 @@
"evidence_sha256": null
}
},
"openrouter-minimax-m3-minimax": {
"openrouter-minimax-m3-deepinfra": {
"route_identity_sha256": null,
"authenticated": false,
"verified_at_utc": null,
Expand All @@ -579,7 +597,7 @@
"evidence_sha256": null
}
},
"openrouter-qwen3.7-plus-alibaba": {
"openrouter-qwen3.8-max-alibaba": {
"route_identity_sha256": null,
"authenticated": false,
"verified_at_utc": null,
Expand Down Expand Up @@ -611,7 +629,7 @@
"evidence_sha256": null
}
},
"openrouter-deepseek-v4-flash-0731-deepseek": {
"openrouter-deepseek-v4-flash-0731-cloudflare": {
"route_identity_sha256": null,
"authenticated": false,
"verified_at_utc": null,
Expand Down Expand Up @@ -651,10 +669,10 @@
"openrouter-gemini-3.6-flash-google-ai-studio",
"openrouter-grok-4.5-xai",
"openrouter-glm-5.2-novita",
"openrouter-minimax-m3-minimax",
"openrouter-qwen3.7-plus-alibaba",
"openrouter-minimax-m3-deepinfra",
"openrouter-qwen3.8-max-alibaba",
"openrouter-mistral-medium-3.5-mistral",
"openrouter-deepseek-v4-flash-0731-deepseek",
"openrouter-deepseek-v4-flash-0731-cloudflare",
"openrouter-hy3-tencent"
],
"shared_fixed_options": {
Expand Down
36 changes: 16 additions & 20 deletions config/sota_v3_pricing_snapshot.json
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
"contract": "sota-v3",
"contract_fingerprint": "a523bdfcebe47bbd",
"status": "catalog-frozen-public-metadata-only",
"checked_at_utc": "2026-08-03T15:53:59Z",
"checked_at_utc": "2026-08-04T16:00:40Z",
"source": "Unauthenticated HTTP GET of https://openrouter.ai/api/v1/models and each selected model's https://openrouter.ai/api/v1/models/{model_id}/endpoints response; no completion or chat endpoint was called.",
"currency": "USD",
"rates_are_per_token": true,
Expand All @@ -29,8 +29,8 @@
"completion": 1e-05
},
"deepseek/deepseek-v4-flash-0731": {
"provider_slug": "deepseek/fp8",
"endpoint_name": "DeepSeek | deepseek/deepseek-v4-flash-20260731",
"provider_slug": "cloudflare/fp8",
"endpoint_name": "Cloudflare | deepseek/deepseek-v4-flash-20260731",
"prompt": 1.4e-07,
"completion": 2.8e-07
},
Expand All @@ -42,8 +42,8 @@
"internal_reasoning": 7.5e-06
},
"minimax/minimax-m3": {
"provider_slug": "minimax/fp8",
"endpoint_name": "Minimax | minimax/minimax-m3-20260531",
"provider_slug": "deepinfra/fp8",
"endpoint_name": "DeepInfra | minimax/minimax-m3-20260531",
"prompt": 3e-07,
"completion": 1.2e-06
},
Expand All @@ -56,24 +56,19 @@
"openai/gpt-5.6-luna": {
"provider_slug": "openai",
"endpoint_name": "OpenAI | openai/gpt-5.6-luna-20260709",
"prompt": 1e-07,
"completion": 6e-07,
"prompt": 2e-07,
"completion": 1.2e-06,
"long_context_override": {
"min_prompt_tokens": 272000,
"prompt": 2e-07,
"completion": 9e-07
}
},
"qwen/qwen3.7-plus": {
"provider_slug": "alibaba/fp8",
"endpoint_name": "Alibaba | qwen/qwen3.7-plus-20260602",
"prompt": 3.2e-07,
"completion": 1.28e-06,
"long_context_override": {
"min_prompt_tokens": 256000,
"prompt": 9.6e-07,
"completion": 3.84e-06
}
"qwen/qwen3.8-max": {
"provider_slug": "alibaba",
"endpoint_name": "Alibaba | qwen/qwen3.8-max-20260803",
"prompt": 2e-06,
"completion": 6e-06
},
"tencent/hy3": {
"provider_slug": "tencent/fp8",
Expand All @@ -95,8 +90,8 @@
"z-ai/glm-5.2": {
"provider_slug": "novita/fp8",
"endpoint_name": "Novita | z-ai/glm-5.2-20260616",
"prompt": 7.266e-07,
"completion": 2.2836e-06
"prompt": 1.4e-06,
"completion": 4.4e-06
}
},
"public_metadata_limitations": [
Expand All @@ -108,5 +103,6 @@
"route_preflight_authorized": false,
"smoke_execution_authorized": false,
"panel_execution_authorized": false,
"publication_authorized": false
"publication_authorized": false,
"pricing_basis": "Undiscounted list rates for the exact pinned route. Promotional discounts are deliberately not reserved against: the GLM 5.2 Novita discount moved from 55.1% to 50% within hours on 2026-08-04, and a reservation computed from a promo is wrong the moment the promo ends. A live discount only ever brings the run in under reserve."
}
3 changes: 2 additions & 1 deletion config/sota_v3_publication_protocol.json
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,8 @@
"cost_estimate_artifact": "results/analysis/sota-v3-pre-smoke-cost-estimate.json",
"operator_must_pass_max_spend_usd": true,
"spend_authorized": false,
"operator_ceiling_usd": null
"operator_ceiling_usd": 150.0,
"operator_ceiling_basis": "Owner-set hard cap, raised 120.00 -> 150.00 on 2026-08-04. The $120 figure was chosen against a $119.76 reservation that turned out to depend on 50%-off promotional rates on openai/gpt-5.6-luna and z-ai/glm-5.2; pinning undiscounted list rates moved the reservation to $127.29 and put the committed plan over its own ceiling. $150.00 clears the current reservation with headroom for a further route substitution or list-price move without another ceiling decision. Enforced by the runner ahead of the cell loop: --max-spend-usd above this value is rejected before any endpoint probe or child process. Projected actual spend is ~$35-45 from July smoke telemetry reprojected at current rates (models emit 48-640 output tokens per decision against a 4,096-token reservation), so this is a backstop, not a forecast. Raising it again is a deliberate edit here."
},
"publication_authorized": false
}
3 changes: 3 additions & 0 deletions docs/PUBLISH_READINESS.md
Original file line number Diff line number Diff line change
Expand Up @@ -860,6 +860,9 @@ decision and why.

| Date | Decision | Evidence / rationale | Effect |
| --- | --- | --- | --- |
| 2026-08-04 | Adopt `qwen/qwen3.8-max` in place of `qwen/qwen3.7-plus`, and adopt a written [route substitution policy](ROUTE_SUBSTITUTION_POLICY.md). | Qwen 3.8 Max shipped 2026-08-03 and the benchmark is intended to track new releases rather than freeze once. Three substitutions in two days had each been decided ad hoc, which is how a benchmark quietly starts measuring whichever host is cheapest this week. | Cohort stays at ten, so the Holm family, the 16x1 allocation, and the power selection are untouched. The Max tier bills 6.25x/4.69x the Plus tier per token. The policy fixes eligibility, forbids price/throughput/first-party status as substitution criteria, and requires re-establishing route and privacy acceptance for any new counterparty. |
| 2026-08-04 | Substitute `deepseek/fp8` -> `cloudflare/fp8` and `minimax/fp8` -> `deepinfra/fp8`. | Both first-party routes were deranked the same day (status `-5` at 78% availability, and `-2`). Both replacements are the same model at the same FP8 quantization, chosen on highest 24h availability, at identical published rates. | No cost or cohort-size effect. The DeepSeek route recovered on its own within the day, so in hindsight that substitution was not strictly necessary; it is kept because Cloudflare holds the better 24h record. Route and privacy acceptance do **not** carry over — Cloudflare and DeepInfra are counterparties this project has never reviewed. |
| 2026-08-04 | Pin undiscounted list rates instead of promotional rates, and enforce `operator_ceiling_usd` in the runner. | `openai/gpt-5.6-luna` and `z-ai/glm-5.2` were both pinned at 50%-off promos; the GLM discount moved 55.1% -> 50% within hours of being recorded. A reservation computed from a promo is wrong the moment it ends. Separately, `operator_ceiling_usd` had sat in config unread, so nothing enforced the committed cap. | Reservation $119.76 -> **$127.29**, which exceeded the then-committed $120.00 ceiling — the plan only ever appeared to fit because of the discounts. **Ceiling raised to $150.00 on 2026-08-04** to clear the reservation with headroom for a further substitution or list-price move. The reservation and the ceiling are now compared by `test_the_committed_plan_fits_under_the_committed_ceiling`, and the committed cost artifact is checked against the configs it claims to describe, so neither can drift silently again. Projected actual spend is ~$35-45 from July smoke telemetry, so the ceiling is a backstop rather than the expected bill. Logged in [`docs/run_logs/sota-v3-lineup-refresh-2026-08-04.md`](run_logs/sota-v3-lineup-refresh-2026-08-04.md). |
| 2026-07-13 | Treat v1 model rows as archived historical evidence, not a current ranking. | Scout-key mismatch affected models unevenly; failed queries were invisible. | Current claims require `sota-v2`; v1 remains auditable under `sota-v1`. |
| 2026-07-13 | Separate API and coding-harness lanes. | Archived rows mixed provider API behavior with uncontrolled CLI harness context and very different output usage. | API becomes the headline lane; CLI harnesses remain diagnostic. |
| 2026-07-13 | Withhold the v2 ranking pending an output-budget sweep. | Archived scores tracked output allowance strongly enough to confound model comparison. | Run the planned cap matrix and freeze a compute policy before the full panel. |
Expand Down
Loading