Research program observatory

A single represent → evaluate → govern → learn loop on strategic judgment. Every study below is pre-registered: its hypothesis, metric, and falsification threshold are hash-sealed into a public integrity chain before any data is collected, and results are published regardless of outcome. 36 studies across four papers.

JCO

AI cheapens generation faster than evaluationsemi
jco-p1 · JCO-P1
AI lowers the marginal cost of generating analysis faster than it lowers the cost of evaluating judgment quality.
Returns to frozen vs live judgment by environmentmanual
jco-p2 · JCO-P2
Returns to frozen judgment rise with task frequency, stability, and low causal ambiguity; returns to live judgment rise with regime change, ambiguity, delayed feedback, and irreversibility.
Meta-judgment premium rises with frozen-judgment stocksemi
jco-p3 · JCO-P3
As the stock of frozen judgment rises, the marginal value of meta-judgment infrastructure rises rather than falls.
Structured representation beats prose-only judgmentmanual
jco-p4 · JCO-P4
Structured representation of judgment artifacts improves calibration, post-hoc learning, and accountability relative to prose-only workflows, holding analysts and evidence fixed.
Decorrelated evaluators detect weak judgments betterautomated
jco-p5 · JCO-P5
Decorrelated evaluators (different model families) detect weak or misleading judgments more reliably than same-family evaluators.
Human+AI with explicit rules beats either alonemanual
jco-p6 · JCO-P6
Human-plus-AI processes using explicit confidence, abstention, and escalation rules outperform humans-alone or models-alone on delayed-ground-truth tasks.
Live-judgment value concentrates in illiquid/regime-shift settingssemi
jco-p7 · JCO-P7
In capital allocation, the excess value of live judgment concentrates in illiquid, weakly price-discovered, and regime-shifting settings; scale flows to frozen judgment.
Judgment-centric employment share is large and risingautomated
jco-p8 · JCO-P8
The judgment-centric share of employment and value added is large (low tens of percent) and rising, fastest in its meta-judgment component.
Forecaster value migrates from production to question-selectionsemi
jco-p9 · JCO-P9
As machine forecasters reach expert parity, the value of a forecasting function migrates from producing forecasts to selecting questions, auditing reasoning, and governing reliance.
Prediction-market amenability predicts where judgment can freezesemi
jco-p10 · JCO-P10
The amenability of a question to a liquid prediction market predicts where judgment can be efficiently frozen; judgment-centric work concentrates in the complement.
Inferred external-actor decisions, as a paired-arm ablationsemi
jco-p14 · JCO-P14
A record capturing inferred external-actor decisions under explicit alternatives and calibrated confidence supports earlier and better-calibrated anticipation than one recording only stated positions, provided the inferences are scored against realized outcomes.
Only feedback-closing systems instrument judgmentsemi
jco-p11 · JCO-P11
Only systems-level AI (persistent memory, outcome feedback, recursive adaptation) closes the represent-evaluate-govern-learn loop; chatbots and episodic agents do not.
Oversight variety must match judgment varietyautomated
jco-p12 · JCO-P12
An oversight architecture detects weak judgments only when its variety is at least that of the judgments it governs; reducing evaluator variety degrades detection in proportion.
Governance value is capability-invariantsemi
jco-p13 · JCO-P13
The value of judgment-governance infrastructure does not fall as model capability rises, and may rise with it.

SCO

Substrate primitive load-bearing ablation (SCO Phase 4)semi
sco-s1 · SCO-S1
Each of the nine kernel primitives is load-bearing: its controlled removal collapses calibration, auditability, or adaptability beyond a pre-registered threshold.
Attribution cross-domain ablation (is the forecast-family null domain-conditional?)semi
sco-s1b · SCO-S1b
Attribution is load-bearing in an attribution-salient (multi-actor, accountability) judgment domain; its null on the weakly-agentic forecast family is domain-conditional, not a kernel-membership error.
Cross-domain substrate sufficiency (SCO Phase 5: legal + scientific)semi
sco-s2 · SCO-S2
The substrate alone is sufficient to encode coherent judgment in the legal and scientific domains (as already shown for medical and joint-doctrine).
Multi-engine SHACL validation agreementautomated
sco-s3 · SCO-S3
Two independent W3C SHACL engines agree on the conformance verdict for all 60 trial instances.
Independent producer/consumer bake-offmanual
sco-s4 · SCO-S4
Independent parties can implement an SCO parser + generator and round-trip a HandoffEnvelope without loss above threshold.
Domain-portfolio sustained-operation validationsemi
sco-s15 · SCO-S15
Spillover routing and strategic-contagion scoring remain stable and meaningful under sustained multi-domain operation.
Causal-mechanism strength updating stabilityautomated
sco-s16 · SCO-S16
Outcome-fed Beta-Binomial mechanism-strength posteriors converge (anti-thrash holds) over ≥50 resolved mechanisms.
Ontology — living evidence statesemi
sco-evidence-state · SCO-STATE
The ontology's served namespace, conformance registry and accruing substrate evidence are reportable as a sealed, reproducible snapshot at any time.

SCB

First SCB-Live calibration runsemi
scb-s5 · SCB-S5
The scaffolded system's Brier on a fresh, future-resolving market question set is no worse than the vanilla baseline, with leakage controlled by construction.
90-day per-grammar / per-operator insight Brierautomated
scb-s6 · SCB-S6
Per-grammar insight calibration (Brier-equivalent) is reportable and discriminates operators once ≥30 insights resolve.
Per-grammar performance card + retirement cycleautomated
scb-s8 · SCB-S8
A composite keepScore (critic-survival, feedback-balance, signpost-trip-rate, Brier) discriminates operators well enough to drive quarterly retirement.
Advisor-honesty proxy (human-vs-AI read)semi
scb-s9 · SCB-S9
An LLM judge reliably distinguishes whether an insight reads as a human strategic advisor vs an AI, and the signal is actionable for framer tuning.
Cross-user grammar pattern release (k≥5 anonymity)automated
scb-s10 · SCB-S10
Grammar-effectiveness patterns generalise across similar user archetypes once ≥5 users have histories.
Insight lifecycle governance auditautomated
scb-s17 · SCB-S17
Every surfaced insight carries a cryptographic seal, kinetic-guard attestation, and change-reason record (lifecycle integrity is complete).
Obligation Repair Track readinessautomated
scb-s18 · SCB-S18
Structured obligation feedback (the Verifier Feedback Contract) is sufficient for a competent agent to converge a broken judgment artifact to conformance — and the track's scorer discriminates full from partial repair.
Do retrieved resolved-question exemplars reduce forecast Brier?semi
icl-exemplar-brier-v1 · SCB-ICL1
Injecting k=4 similarity-retrieved resolved questions (with outcomes) into the vanilla forecaster's context reduces mean Brier vs a no-exemplar control.
Does injecting an empirical reference-class base rate beat self-recall?semi
icl-empirical-baserate-brier-v1 · SCB-ICL2
Rendering the persisted empirical reference-class base rate into the forecast context reduces Brier vs both no-prime and the model's self-recalled base-rate prime.

SIP

Insight vs hired-analyst blind comparisonmanual
sip-s11 · SIP-S11
System-surfaced insights are rated at least as novel/useful as a hired analyst's on the same user-days by a blind panel.
Operator retirement / addition cycle (SIP)automated
sip-s12 · SIP-S12
The quarterly operator-lifecycle cycle (retire below keepScore, add candidates) improves aggregate insight quality over time.
Insight paper — living evidence statesemi
sip-evidence-state · SIP-STATE
The insight architecture's per-grammar track record, portfolio state and withdrawal ledger are reportable as a sealed, reproducible snapshot at any time.

RR

Reference Run #001 — AI-Compute Delta Feed hypothesessemi
rr-001 · RR-001
The 12 pre-registered AI-compute hypotheses (H1–H12) resolve as recorded, published regardless of outcome.

BOARD

Board-governance human-AI uplift (N=16 crossover)manual
board-s14 · BOARD-S14
The board-governance extension improves oversight-decision quality in a within-subject crossover.
Part of the Leadership Under Uncertainty research programme. The launch paper, the companion manuscripts and how to take part are at the programme front page.