/proof RR-001 scoreboard H10

Hypothesis H10

WordingSCO artifacts pass the structural fidelity test
Falsification thresholdFresh LLM cannot reconstruct strategic landscape from artifacts alone (downstream summary scores <60% on rubric)
Current statuspending
Evidence summaryAwaiting evaluator run
Pre-disclosed expected outcome (honesty disclosure):
Expected supported — downstreamFidelity test design holds in unit-test scaffold; production validation pending.

How this hypothesis resolves

The threshold above is the pre-registered falsification rule. Once the evaluator runs (per the schedule in the protocol document), the `currentStatus` column on `reference_run_hypotheses` updates and this page reflects the verdict. All evaluator runs are cryptographically sealed via the daily Merkle anchor.

Source of truth