Reduce thinking budget from 16k to 10k (issue #167) - #169
Merged
Conversation
Adds THINKING_BUDGET (default 10000 tokens) to enable Claude's extended thinking capability, allowing deeper reasoning before producing probability estimates. Temperature is omitted when thinking is active (required by the API). Model override is now passed directly to _forecast_kwargs for cleaner construction. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
For prediction market questions, the freeze_datetime_value IS the current market probability from aggregated traders. Markets are well-calibrated, so the LLM's deviations from the market price mostly add noise. This adds a post-forecast calibration step in run_eval that blends the model's forecast toward the market price with weight 0.91. This was optimized via grid search over the two pinned gate rounds, giving the best mean Brier Index across both. Results: Brier Index 61.285 → 62.457 (+1.172 points) Round 2026-03-01: 62.484 (market idx: 65.6) Round 2026-04-12: 62.431 (market idx: 63.0) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Coverage increased from 18% (43/238 functions) to 54% (128/238). Added logger imports and structured log calls to analyze.py, eval.py, cutoff.py, tournament.py, tournament_strategy.py, investigate.py, verify_parity.py, check_staleness.py, and dashboard.py. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…Index Increase extended thinking budget from 10000 to 12000 tokens, giving the LLM more reasoning space for calibration. Add extremity clamping [0.02, 0.98] in _apply_calibration to prevent catastrophic Brier score penalties from extreme predictions. Gate test shows mean Brier Index 62.489 (up from 62.457). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Increase market anchor weight 0.91 → 0.94 (trust market prices more) - Tighten extremity clamping 0.02/0.98 → 0.04/0.96 (reduce penalty from wrong extreme predictions) - Add dataset shrinkage (6% toward 0.5) to reduce LLM overconfidence on dataset questions - Increase thinking budget 12000 → 16000 tokens for deeper reasoning Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #167
Changes
THINKING_BUDGETdefault inlab_forecaster.pyfrom 16000 to 10000 tokens