/proof › RR-001 scoreboard › H10
Hypothesis H10
WordingSCO artifacts pass the structural fidelity test
Falsification threshold
Fresh LLM cannot reconstruct strategic landscape from artifacts alone (downstream summary scores <60% on rubric)Current statuspending
Evidence summaryAwaiting evaluator run
Pre-disclosed expected outcome (honesty disclosure):
Expected supported — downstreamFidelity test design holds in unit-test scaffold; production validation pending.
Expected supported — downstreamFidelity test design holds in unit-test scaffold; production validation pending.
How this hypothesis resolves
The threshold above is the pre-registered falsification rule. Once the evaluator runs (per the schedule in the protocol document), the `currentStatus` column on `reference_run_hypotheses` updates and this page reflects the verdict. All evaluator runs are cryptographically sealed via the daily Merkle anchor.
Source of truth
- Schema:
reference_run_hypotheses(migration 0104) - Evaluator:
lib/referenceRun/hypothesisEvaluator.ts - Writer:
DrizzleHypothesisDbWriterinlib/referenceRun/dbWriters.ts - Locked wording:
LOCKED_HYPOTHESESinlib/referenceRun/hypothesisEvaluator.ts