RESEARCH DEMONSTRATION

SCB/SCO Reference Run #001 — AI Compute (30 days)

Public longitudinal demonstration of structured strategic-context artifacts (entity snapshots, analytic deltas, evidence packs, risk timeseries) produced over the AI compute domain. Pre-registered protocol cryptographically anchored on day 0; all 12 hypotheses with falsification thresholds locked before run start; publication date committed before any results exist.

What this page is. A research-demonstration corpus produced under Leadership Under Uncertainty.
What it is not. Not investment advice. Not a commercial product launch. Not a claim of complete domain coverage. Both successes and failures are published.

Run timeline

Pre-registration sealing: 2026-05-07 23:59 UTC (cryptographic anchor of locked protocol)
Run period: 2026-05-07 → 2026-06-05 (30 days)
Validation week: 2026-06-06 → 2026-06-12 (conformance, 5 case studies, capability card, ablations, downstream-fidelity test)
Publication week (committed): 2026-06-13 → 2026-06-19 (corpus, technical note, all 12 hypothesis results — including failures)

The 12 falsifiable hypotheses

Each hypothesis has an explicit falsification threshold locked before run start. Status is updated through the run; final results published in the validation week. Whether each hypothesis is supported, falsified, or inconclusive, the result is published.

IDHypothesisFalsification thresholdStatus
H1Our 30-day calibration is at or better than Manifold/Metaculus median on shared questions
Expected: Likely inconclusive or falsified — public prediction markets are hard to beat on raw Brier; this is not where we lead.
Median Brier on resolved markets exceeds market median Brier by >0.05pending
H2We meaningfully beat vanilla LLM + web search on calibration on the same prompt set
Expected: Likely inconclusive or falsified — frontier LLM + browse is competitive on raw calibration. We lead on structure + tradecraft, not Brier.
Vanilla LLM median Brier <= ours - 0.02 OR equal within +/-0.01pending
H3We dominate public-research-org analyst content (CSET, RAND, Brookings, Epoch AI, EIA, LBL) on the four core tradecraft dimensions (provenance, alternatives, gaps, update discipline)
Expected: Expected supported — NorthBrief enforces all 4 dimensions by construction (renderers + judge gate); analyst content is unevenly structured.
Aggregated public-research-org SCB-rubric score >= ours on any of the four dimensionspending
H4Anti-thrash gates improve calibration
Expected: Expected supported — small effect (5-15 basis points). Tech-debt rounds R20-R50 already show stability primitives prevent oscillation.
No-gates ablation Brier improves by >=0.01 vs gates-onpending
H5Cross-domain edge modeling improves downstream summary quality
Expected: Inconclusive over 30 days — cross-domain effects take longer than a month to manifest; will revisit at 90-day mark.
No-edges ablation downstream-LLM fidelity score >= with-edges within +/-2%pending
H6Our information gaps are the right unknowns
Expected: Operator action — expert panel recruitment deferred per founder decision. Marked as TBD; if no panel, hypothesis cannot resolve.
Informal expert review (n>=5) rates our gaps wrong/missing on >40% of itemspending
H7Our binding-constraint identification (BMI) matches expert consensus
Expected: Probably supported — BMI substrate aligns with current Epoch AI + CSET public analysis on chip-control + power-bottleneck framings.
Epoch AI, CSET, RAND, Brookings, or LBL public analysis explicitly disagrees with our BMI top-3 binding constraints during the runpending
H8Pipeline runs >=27 of 30 days without intervention
Expected: Expected supported — production cron jobs (daily/weekly/monthly publish) are now wired with per-user lock + idempotency; outages should be rare.
>=4 outage days requiring manual fixpending
H9Per-domain marginal cost stays under $10/day
Expected: Expected supported — per-tenant cost cap enforced architecturally at lib/llm/costCap.ts.
Mean daily spend exceeds $10/day over the 30 dayspending
H10SCO artifacts pass the structural fidelity test
Expected: Expected supported — downstreamFidelity test design holds in unit-test scaffold; production validation pending.
Fresh LLM cannot reconstruct strategic landscape from artifacts alone (downstream summary scores <60%)pending
H11Group B prospects recognize the value proposition
Expected: Operator action — recruitment + calls deferred. Marked as TBD.
0 of 8 prospects in Group B express interest in seeing more after the callpending
H12AI compute is the right reference domain
Expected: Operator action — recruitment + calls deferred. Marked as TBD.
>=3 of 8 Group B prospects say 'interesting but you should be doing this in [other domain] for me'pending

Comparator suite

Five comparator categories, locked at protocol seal:

  • Calibration spine — 12-15 questions on Manifold, Metaculus, Polymarket, INFER. Weekly probability snapshots from us + each market. Brier on resolved.
  • Vanilla LLM baselines — 4 configurations (Claude alone, Claude+search, Gemini+grounding, Perplexity Pro) run weekly on identical prompts.
  • Public-research-org tradecraft scoring — Epoch AI, CSET, RAND, Brookings, LBL, EIA scored on the SCB rubric (provenance, alternatives, gaps, update discipline). Subscription-gated content (SemiAnalysis, Stratechery, equity research) is deliberately excluded — see Open-Sources-Only Commitment below.
  • Internal ablations — no-anti-thrash, no-cross-domain, cheap-models-only.
  • Downstream LLM structural fidelity test — fresh LLM reconstructs strategic landscape from SCO artifacts only; scored against ground-truth.

Live data surface

The AI compute domain artifacts are exposed at the existing public-data endpoints. These remain available before, during, and after the run.

Pre-registration documents

The locked protocol, locked questions JSON, and locked comparators JSON are hashed (SHA-256), bundled, and the bundle hash is anchored cryptographically on day 0. Verification instructions are published alongside the corpus in the publication week.

  • Protocol document: /docs/reference-runs/RR_001_AI_COMPUTE_PROTOCOL.md
  • Locked questions: /data/reference-runs/rr-001-questions.json
  • Locked comparators: /data/reference-runs/rr-001-comparators.json

Limitations (will be expanded post-run)

  • 30 days is enough to demonstrate the artifact pipeline; not enough for long-term track-record claims.
  • Public-source bias — no private allocation data, no expert call corpus, no proprietary equity research.
  • Sample size on calibration spine: ~12-15 questions. Brier confidence intervals will be reported.
  • The reference run is demonstration corpus, not held-out benchmark — must not be incorporated into SCB-Core evaluation tasks.
  • Vanilla-LLM ecosystem changes monthly; baseline reproducibility requires pinned model snapshots (locked at seal).
  • Open-Sources-Only Commitment. This run consumes only publicly accessible sources. Subscription-gated analyst content (SemiAnalysis, Stratechery, equity research, expert call libraries, paid market intelligence) is not consumed by the pipeline and not used as research input by the operator. This preserves the independent-derivation claim by construction — external researchers can replicate the corpus from publicly accessible inputs only. Such products are acknowledged as separately-existing offerings (chip-level supply-chain depth, strategic-frame analysis) that we do not consume, score, or compete with.