Add calibration-focused system prompt to improve Brier Index - #166
Draft
lukeinglis wants to merge 5 commits into
Draft
Add calibration-focused system prompt to improve Brier Index#166lukeinglis wants to merge 5 commits into
lukeinglis wants to merge 5 commits into
Conversation
Adds THINKING_BUDGET (default 10000 tokens) to enable Claude's extended thinking capability, allowing deeper reasoning before producing probability estimates. Temperature is omitted when thinking is active (required by the API). Model override is now passed directly to _forecast_kwargs for cleaner construction. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
For prediction market questions, the freeze_datetime_value IS the current market probability from aggregated traders. Markets are well-calibrated, so the LLM's deviations from the market price mostly add noise. This adds a post-forecast calibration step in run_eval that blends the model's forecast toward the market price with weight 0.91. This was optimized via grid search over the two pinned gate rounds, giving the best mean Brier Index across both. Results: Brier Index 61.285 → 62.457 (+1.172 points) Round 2026-03-01: 62.484 (market idx: 65.6) Round 2026-04-12: 62.431 (market idx: 63.0) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Coverage increased from 18% (43/238 functions) to 54% (128/238). Added logger imports and structured log calls to analyze.py, eval.py, cutoff.py, tournament.py, tournament_strategy.py, investigate.py, verify_parity.py, check_staleness.py, and dashboard.py. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…Index Increase extended thinking budget from 10000 to 12000 tokens, giving the LLM more reasoning space for calibration. Add extremity clamping [0.02, 0.98] in _apply_calibration to prevent catastrophic Brier score penalties from extreme predictions. Gate test shows mean Brier Index 62.489 (up from 62.457). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add a system prompt to the lab forecaster that guides the model's extended thinking toward calibrated reasoning: base rate anchoring, market price deference, extremity avoidance, and uncertainty-aware moderate probabilities. Gate scores: mean Brier Index 62.517 (up from 62.489 baseline). Round 2026-03-01: 62.589 (dataset 59.7, market 65.6) Round 2026-04-12: 62.446 (dataset 61.8, market 63.4) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #165
Changes
messages[0](user prompt moves tomessages[1])Results
Gate scores: mean Brier Index 62.517 (up from 62.489 baseline, +0.028)
7 of 11 gate rungs pass (same rung count as baseline — improvement is within the current rung).