SCB/SCO Reference Run #001 — AI Compute (30 days)
Public longitudinal demonstration of structured strategic-context artifacts (entity snapshots, analytic deltas, evidence packs, risk timeseries) produced over the AI compute domain. Pre-registered protocol cryptographically anchored on day 0; all 12 hypotheses with falsification thresholds locked before run start; publication date committed before any results exist.
What it is not. Not investment advice. Not a commercial product launch. Not a claim of complete domain coverage. Both successes and failures are published.
Run timeline
Pre-registration sealing: 2026-05-07 23:59 UTC (cryptographic anchor of locked protocol)
Run period: 2026-05-07 → 2026-06-05 (30 days)
Validation week: 2026-06-06 → 2026-06-12 (conformance, 5 case studies, capability card, ablations, downstream-fidelity test)
Publication week (committed): 2026-06-13 → 2026-06-19 (corpus, technical note, all 12 hypothesis results — including failures)
The 12 falsifiable hypotheses
Each hypothesis has an explicit falsification threshold locked before run start. Status is updated through the run; final results published in the validation week. Whether each hypothesis is supported, falsified, or inconclusive, the result is published.
| ID | Hypothesis | Falsification threshold | Status |
|---|---|---|---|
| H1 | Our 30-day calibration is at or better than Manifold/Metaculus median on shared questions Expected: Likely inconclusive or falsified — public prediction markets are hard to beat on raw Brier; this is not where we lead. | Median Brier on resolved markets exceeds market median Brier by >0.05 | pending |
| H2 | We meaningfully beat vanilla LLM + web search on calibration on the same prompt set Expected: Likely inconclusive or falsified — frontier LLM + browse is competitive on raw calibration. We lead on structure + tradecraft, not Brier. | Vanilla LLM median Brier <= ours - 0.02 OR equal within +/-0.01 | pending |
| H3 | We dominate public-research-org analyst content (CSET, RAND, Brookings, Epoch AI, EIA, LBL) on the four core tradecraft dimensions (provenance, alternatives, gaps, update discipline) Expected: Expected supported — NorthBrief enforces all 4 dimensions by construction (renderers + judge gate); analyst content is unevenly structured. | Aggregated public-research-org SCB-rubric score >= ours on any of the four dimensions | pending |
| H4 | Anti-thrash gates improve calibration Expected: Expected supported — small effect (5-15 basis points). Tech-debt rounds R20-R50 already show stability primitives prevent oscillation. | No-gates ablation Brier improves by >=0.01 vs gates-on | pending |
| H5 | Cross-domain edge modeling improves downstream summary quality Expected: Inconclusive over 30 days — cross-domain effects take longer than a month to manifest; will revisit at 90-day mark. | No-edges ablation downstream-LLM fidelity score >= with-edges within +/-2% | pending |
| H6 | Our information gaps are the right unknowns Expected: Operator action — expert panel recruitment deferred per founder decision. Marked as TBD; if no panel, hypothesis cannot resolve. | Informal expert review (n>=5) rates our gaps wrong/missing on >40% of items | pending |
| H7 | Our binding-constraint identification (BMI) matches expert consensus Expected: Probably supported — BMI substrate aligns with current Epoch AI + CSET public analysis on chip-control + power-bottleneck framings. | Epoch AI, CSET, RAND, Brookings, or LBL public analysis explicitly disagrees with our BMI top-3 binding constraints during the run | pending |
| H8 | Pipeline runs >=27 of 30 days without intervention Expected: Expected supported — production cron jobs (daily/weekly/monthly publish) are now wired with per-user lock + idempotency; outages should be rare. | >=4 outage days requiring manual fix | pending |
| H9 | Per-domain marginal cost stays under $10/day Expected: Expected supported — per-tenant cost cap enforced architecturally at lib/llm/costCap.ts. | Mean daily spend exceeds $10/day over the 30 days | pending |
| H10 | SCO artifacts pass the structural fidelity test Expected: Expected supported — downstreamFidelity test design holds in unit-test scaffold; production validation pending. | Fresh LLM cannot reconstruct strategic landscape from artifacts alone (downstream summary scores <60%) | pending |
| H11 | Group B prospects recognize the value proposition Expected: Operator action — recruitment + calls deferred. Marked as TBD. | 0 of 8 prospects in Group B express interest in seeing more after the call | pending |
| H12 | AI compute is the right reference domain Expected: Operator action — recruitment + calls deferred. Marked as TBD. | >=3 of 8 Group B prospects say 'interesting but you should be doing this in [other domain] for me' | pending |
Comparator suite
Five comparator categories, locked at protocol seal:
- Calibration spine — 12-15 questions on Manifold, Metaculus, Polymarket, INFER. Weekly probability snapshots from us + each market. Brier on resolved.
- Vanilla LLM baselines — 4 configurations (Claude alone, Claude+search, Gemini+grounding, Perplexity Pro) run weekly on identical prompts.
- Public-research-org tradecraft scoring — Epoch AI, CSET, RAND, Brookings, LBL, EIA scored on the SCB rubric (provenance, alternatives, gaps, update discipline). Subscription-gated content (SemiAnalysis, Stratechery, equity research) is deliberately excluded — see Open-Sources-Only Commitment below.
- Internal ablations — no-anti-thrash, no-cross-domain, cheap-models-only.
- Downstream LLM structural fidelity test — fresh LLM reconstructs strategic landscape from SCO artifacts only; scored against ground-truth.
Live data surface
The AI compute domain artifacts are exposed at the existing public-data endpoints. These remain available before, during, and after the run.
- /proof/ai-compute — daily-updated charter, BMI, capex credibility, assumption fragility, contagion, watchlist, spillover map
- /api/v1/public/ai-compute — JSON API index
Pre-registration documents
The locked protocol, locked questions JSON, and locked comparators JSON are hashed (SHA-256), bundled, and the bundle hash is anchored cryptographically on day 0. Verification instructions are published alongside the corpus in the publication week.
- Protocol document:
/docs/reference-runs/RR_001_AI_COMPUTE_PROTOCOL.md - Locked questions:
/data/reference-runs/rr-001-questions.json - Locked comparators:
/data/reference-runs/rr-001-comparators.json
Limitations (will be expanded post-run)
- 30 days is enough to demonstrate the artifact pipeline; not enough for long-term track-record claims.
- Public-source bias — no private allocation data, no expert call corpus, no proprietary equity research.
- Sample size on calibration spine: ~12-15 questions. Brier confidence intervals will be reported.
- The reference run is demonstration corpus, not held-out benchmark — must not be incorporated into SCB-Core evaluation tasks.
- Vanilla-LLM ecosystem changes monthly; baseline reproducibility requires pinned model snapshots (locked at seal).
- Open-Sources-Only Commitment. This run consumes only publicly accessible sources. Subscription-gated analyst content (SemiAnalysis, Stratechery, equity research, expert call libraries, paid market intelligence) is not consumed by the pipeline and not used as research input by the operator. This preserves the independent-derivation claim by construction — external researchers can replicate the corpus from publicly accessible inputs only. Such products are acknowledged as separately-existing offerings (chip-level supply-chain depth, strategic-frame analysis) that we do not consume, score, or compete with.