/proof › RR-001 scoreboard › H2
Hypothesis H2
WordingWe meaningfully beat vanilla LLM + web search on calibration on the same prompt set
Falsification threshold
Vanilla LLM median Brier <= ours minus 0.02 OR equal within +/-0.01Current statuspending
Evidence summaryAwaiting evaluator run
Pre-disclosed expected outcome (honesty disclosure):
Likely inconclusive or falsified — frontier LLM + browse is competitive on raw calibration. We lead on structure + tradecraft, not Brier.
Likely inconclusive or falsified — frontier LLM + browse is competitive on raw calibration. We lead on structure + tradecraft, not Brier.
How this hypothesis resolves
The threshold above is the pre-registered falsification rule. Once the evaluator runs (per the schedule in the protocol document), the `currentStatus` column on `reference_run_hypotheses` updates and this page reflects the verdict. All evaluator runs are cryptographically sealed via the daily Merkle anchor.
Source of truth
- Schema:
reference_run_hypotheses(migration 0104) - Evaluator:
lib/referenceRun/hypothesisEvaluator.ts - Writer:
DrizzleHypothesisDbWriterinlib/referenceRun/dbWriters.ts - Locked wording:
LOCKED_HYPOTHESESinlib/referenceRun/hypothesisEvaluator.ts