Leadership Under Uncertainty
Launch paper

The Judgment Layer: Making AI Judgment Explicit, Testable and Cumulative

Christopher Berzins, Visiting Professor, Graduate School of Public and International Affairs, University of Ottawa.
Draft 11, September 2026. Correspondence: cberzins@uottawa.ca.
Download the PDF
A working draft, circulated for comment. Comments and corrections are welcome at the address above.

Abstract

In September 1962 the United States Intelligence Board judged it unlikely that the Soviet Union would place strategic missiles in Cuba. The estimate was carefully reasoned, and it was wrong. It was corrected by an official who set the probability aside and asked what the new air-defence sites in Cuba were there to protect. AI systems that carry out open-ended work now make judgments of this kind all the time, about what deserves attention, which explanations to pursue and when the evidence justifies action, and usually nobody has examined how they make them.

This paper treats judgment as a capability of AI systems that can be specified, disciplined and tested, and proposes that it can be made cumulative, so that experience scored against outcomes improves the questions a system asks. Intelligence Community Directive 203 supplies the institutional precedent, and a judgment layer carries it into AI design. When the analyst is a system whose procedures can be switched on and off, the effect of a single tradecraft standard can be isolated, which has not been possible with human analysts. Every claim about the programme behind the paper is labelled as specified, implemented, operating or measured.

Three measured results are reported. Six widely used models, left unprompted, omitted the sources, alternatives or gaps the standards require. The programme's own architecture produced citations that resolved to the document without supporting the claims attached to them. And retrieval, structured prompting and deliberation added nothing measurable to forecast accuracy. A decisive study is specified, the objection that better models will make the layer redundant is addressed, and ways to take part are set out.

Keywords. judgment; artificial intelligence; strategic cognition; insight generation; analytic standards; intelligence analysis; decision-making under uncertainty; AI safety; human–AI collaboration

How to cite

Berzins, C. (2026). The Judgment Layer: Making AI Judgment Explicit, Testable and Cumulative. Draft 11. Leadership Under Uncertainty research programme, Graduate School of Public and International Affairs, University of Ottawa. https://leadershipunderuncertainty.org/papers/the-judgment-layer-draft-11.pdf
Companion manuscripts

Six manuscripts the launch paper draws on

Each is at manuscript stage and not yet public. They are available to collaborators on request, and they will be listed here with identifiers as they are released.

The philosophical account
From Information to Judgment: A Judgment-First Account of Cognition in Bounded Agents
Manuscript, 2026. Not yet public.
Why the minimal unit of consequential cognition for an agent acting under uncertainty is a situated, revisable commitment: a claim held with a stated confidence, resting on attributed evidence, standing against represented alternatives and answerable to an outcome.
The shared representation
The Strategic Cognition Ontology: A Validation-Oriented Application Ontology for AI-Assisted Strategic Analysis
Manuscript, 2026. Not yet public.
The nine-element vocabulary for judgments under uncertainty, its strategic extensions, the machine-readable schema and the public validator against which artifacts are checked.
The evaluation instrument
SCB: A Two-Layer Benchmark for Strategic Cognition in AI Systems
Manuscript, 2026. Not yet public.
Outcome-scored forecasting alongside tradecraft-scored analysis, with the judge-reliability audit, the capped composite and the fidelity ladder that grades what a judgment record can show.
The insight architecture
SIP: A Substrate-Grounded Architecture and Evaluation Methodology for Strategic Insight Generation in AI Systems
Manuscript, 2026. Not yet public.
The closed vocabulary of search patterns, the gates between generation, criticism and audit, and the evaluation design under which the architecture is compared with a plain model call.
The organizational theory
AI for Judgment-Centric Organizations: A Research Program for Representing, Evaluating, and Governing Strategic Judgment
Manuscript, 2026. Not yet public.
The three conditions that define judgment-centric work, live and frozen judgment, and twenty-one pre-specified propositions, including the predictions that answer the objection that better models will make a judgment layer redundant.
The board application
SCO-Board: An Evidence-Gated Application Profile and Pre-Registered Evaluation Protocol for AI-Assisted Board Oversight
Manuscript, 2026. Not yet public.
The application profile for board and committee documents, the comparison against guard-matched prompting, the citation-support audit and the pre-registered human study that has not yet run.