/proof RR-001 scoreboard H2

Hypothesis H2

WordingWe meaningfully beat vanilla LLM + web search on calibration on the same prompt set
Falsification thresholdVanilla LLM median Brier <= ours minus 0.02 OR equal within +/-0.01
Current statuspending
Evidence summaryAwaiting evaluator run
Pre-disclosed expected outcome (honesty disclosure):
Likely inconclusive or falsified — frontier LLM + browse is competitive on raw calibration. We lead on structure + tradecraft, not Brier.

How this hypothesis resolves

The threshold above is the pre-registered falsification rule. Once the evaluator runs (per the schedule in the protocol document), the `currentStatus` column on `reference_run_hypotheses` updates and this page reflects the verdict. All evaluator runs are cryptographically sealed via the daily Merkle anchor.

Source of truth