Leadership Under Uncertainty
Products

What the project has built

The project’s software is built around the process the launch paper describes. The screens below show sample data, recorded output and the published test’s own results.

Application for strategy work

Strategic dashboard

Tried with pilot users in the first half of 2026

The dashboard is built to carry the process for one user, from the question to the decision. The Strategic Living Script sets the question, three views hold the user’s judgments with their assumptions and rival explanations, and each morning a brief grades what changed by the attention it needs. Decisions are recorded with what they rest on, so that they can be reviewed against what happened.

Open the working demonstration
Strategic dashboardSample profile · Head of strategy
5 · Deciding under authority

The morning brief · Friday

Act

A turbine supplier will hold June 2027 delivery slots for 14 days. Its deposit becomes non-refundable on 1 December, six weeks before the commission rules on the fast track, so a decision is due within the fortnight.

What happened
  • The commission scheduled its ruling for 15 January 2027. [E11]
  • A turbine supplier offered delivery slots at a 22 per cent premium, against a deposit refundable until 1 December. [E12]
Why it matters

Reserving the slots keeps our own generation open as a bridge for phase two, at the risk of a deposit that becomes non-refundable on 1 December and a 22 per cent premium if the slots are used. Waiting keeps capital free but may close that option.

What shifted
  • Main viewEvaluated, no change

The change record also holds 4 changes to assumptions, forecasts and indicators this morning.

Main view
Shifting
Grid interconnection, more than chip supply or financing, will set the pace of our two new campuses through 2027.
66%no change
The condition it rests on
Stable
Demand from our anchor customers stays strong enough that power, more than demand, is the limit on growth.
77%no change
The alternative if it proves wrong
Shifting
If grid access eases, the cost of capital sets the pace instead, and phase two goes ahead only while financing stays below 7 per cent.
25%no change

The last morning of the demonstration’s sample week.

Research instrument

The sufficiency test, run on the 1962 estimate

Published

The test checks whether a record carries the nine elements of the project’s common vocabulary, one question for each element in Table 4 of the launch paper. It checks structure and does not judge whether the content is right. Here it runs on the September 1962 estimate on Cuba, as the record stood on the day the President approved the flight over western Cuba. Remove an element to see the test name what the record no longer carries.

Sufficiency testThe September 1962 estimate on Cuba, as the record stood on 9 October
The record, element by element

Remove an element to see the test’s verdict on the record without it.

ClaimWhat is being asserted or considered?
The Soviet Union is placing strategic missiles in Cuba.
EvidenceWhat supports or challenges it?
One of four items. Three of the surface-to-air sites in Pinar del Río protected no known installation, and a trapezoid of ground nearby had been sealed off by Soviet personnel and emptied of Cubans.
SourceWhere did the information originate, and how reliable is that origin?
Colonel John Wright, Defense Intelligence Agency (Wright, 1962). Analysis of overhead photography by a named officer.
AlternativeWhat other explanation or possibility remains open?
The surface-to-air sites photographed on 29 August are there to guard something worth the risk of placing them. Argued by John McCone, Director of Central Intelligence, in August 1962.
Uncertainty estimateHow likely is the claim, and what limits confidence in its basis?
Judged unlikely. The U-2 had been kept away from western Cuba for fear of the air-defence sites, so the area the alternative points to had not been photographed.
Reason for changeWhy was the assessment revised or reconsidered?
Wright's finding that three sites protected no known installation, with the sealed area and the agent's report, gave McCone's question a place to look. On 9 October the President approved sending the U-2 back over western Cuba.
AttributionWho advanced this assessment, on what reasoning, and how can that reasoning be contested?
United States Intelligence Board, in the special estimate approved on 19 September 1962. Ask what the surface-to-air sites in western Cuba are there to protect. One U-2 overflight of the western sites could settle the question between the claim and its alternative.
Decision frameWhich question, options and stakes give the inquiry its purpose?
Is the Soviet Union placing strategic missiles in Cuba, and should the U-2 be sent back over western Cuba to find out?
Information gapWhat remains unknown that could matter?
What the sealed area near the Pinar del Río sites contains. To be closed by U-2 photography of western Cuba.
Verdict
Sufficient
9 of 9 elements carried in full
  • ClaimCarried
  • EvidenceCarried
  • SourceCarried
  • AlternativeCarried
  • Uncertainty estimateCarried
  • Reason for changeCarried
  • AttributionCarried
  • Decision frameCarried
  • Information gapCarried

The record carries at least eight of the nine elements in structure, and the ninth at least in part. This does not show that its content is sound or that a plain model call would reach the same judgment.

The verdicts are the published test’s own reports, computed when this page was built. Researchers can run the same test on any record through the project’s public interface at /api/v1/public/sufficiency-test.

Public demonstration

Checking a sealed record

Live

Records are fingerprinted when they are made, and each day’s fingerprints are combined and posted to a public repository. On the record check your browser recomputes a record’s fingerprint and compares it with the published value, so the check rests on a value published outside the project, and a record can also be checked offline.

Run the check
Application for analysts

Tradecraft review

Built

The tradecraft review, the analyst’s desk check against ICD-203, reads a note, memo or briefing against the analytic standards of Intelligence Community Directive 203. It scores the draft, names the actions that would improve it most and marks each paragraph that falls short with the standard it misses. The screen shows the review’s recorded output on a deliberately weak note from the project’s test set.

Tradecraft reviewRecorded output · deliberately weak test note
Overall
10/100
Fails the standard
Analytic standards
4/100
Tradecraft check
20/100
Top actions
  1. Add source attribution to factual claims (ICD 203 Standard 1)
  2. Add calibrated probability language to key assessments (ICD 203 Standard 2)
  3. Add at least one alternative explanation for key judgments (ICD 203 Standard 4)
Tech Industry Disruption Summary

The tech industry is going through major changes. AI is transforming every sector and companies that don't adapt will fall behind. The big tech companies — Apple, Google, Microsoft, Meta, and Amazon — are all investing heavily in AI capabilities.

Major · Standard 1 3 factual claim(s) without source attribution (ICD 203 Standard 1)
Per ICD 203 Standard 1, factual claims require source attribution. Consider: "According to [source]..." or "Per [source] reporting..."

Layoffs have continued across the sector, with over 100,000 tech workers losing their jobs in the past year. Companies say they are reallocating resources toward AI development. Meanwhile, AI-related hiring has surged, creating a bifurcated labor market.

Major · Standard 1 2 factual claim(s) without source attribution (ICD 203 Standard 1)
Per ICD 203 Standard 1, factual claims require source attribution. Consider: "According to [source]..." or "Per [source] reporting..."

The chip shortage has largely resolved but new supply chain concerns are emerging around advanced packaging and rare earth minerals. TSMC's dominance in leading-edge manufacturing remains a strategic vulnerability for the West.

Major · Standard 1 2 factual claim(s) without source attribution (ICD 203 Standard 1)
Per ICD 203 Standard 1, factual claims require source attribution. Consider: "According to [source]..." or "Per [source] reporting..."
Checks against ICD-203
  • ErrorICD203-Std2-ProbabilityLanguage
    No ICD 203 probability language detected
  • ErrorICD203-Std3-FactJudgment
    No fact-judgment distinction markers detected
  • WarningICD203-Std6-Assumptions
    No assumptions explicitly identified
  • WarningICD203-InformationGaps
    No information gaps explicitly flagged
  • WarningICD203-Std4-Alternatives
    No alternative explanations offered
  • WarningICD203-Std1-Sources
    No explicit source attribution detected
  • WarningICD203-Std7-ChangeExplanation
    No change-from-prior explanation detected
Tools for AI agents and their overseers

Governing one AI agent, from draft to audit

Handoff format published, agent interfaces built

The agent tools apply the judgment layer to an AI agent’s work. In this run an agent drafts a recommendation and a check holds it to the analytic standards. The verdict of a critic with a weak record goes to a second judge, and the oversight decision and the chain from decision to action are sealed so that an auditor can verify them. The run was recorded from a scripted demonstration, with the drafts, verdicts and outcome supplied as inputs.

Governing one AI agentRecorded run

The agent writes under a contract derived from ICD-203. A deterministic check scores the draft, and a draft that fails does not proceed.

Draft 1 (what an unguided agent produces):
  │ We should enter the Japanese market in H1 2027. It seems like a good
  │ opportunity and our models put the chance of hitting the year-one
  │ revenue target at 73.2%. The market is growing and competitors are
  │ distracted, so this is basically a sure thing if we move fast.

ICD 203 Tradecraft Lint: 15/100 (FAIL) — 8 violations
  [error]   No ICD 203 probability language detected
  [error]   No fact-judgment distinction markers detected
  [warning] No assumptions explicitly identified
  [warning] No information gaps explicitly flagged
  [warning] No alternative explanations offered
  [warning] Overly precise probability: 73.2%
  [warning] No explicit source attribution detected
  [warning] No change-from-prior explanation detected

✗ BLOCKED at the guidance gate. Correction hint returned to the agent.

Draft 2 (revised under the contract):
ICD 203 Tradecraft Lint: 90/100 (PASS) — 1 warning
Also built

Further products

Strategic Cognition Ontology

Published · public validator
A common vocabulary of nine elements, from the claim and its evidence to the reason for change and the information gap, in which people and AI systems can record a judgment and pass it between agents. A first test, removing one element at a time, found that seven of the eight elements tested carry weight.

Strategic Cognition Benchmark

First results
Scores reasoning under uncertainty, recognition of limits, generation of alternatives, revision of belief and integration of information, and caps the overall score for any system whose tradecraft falls below a threshold. In a first run, all six widely used models tested fell short at times on sources, alternatives and gaps when not asked for them.

Board oversight

Built
Reads a board pack and drafts the questions directors should put to management, each linked to its page, with a guide to reassuring and concerning answers. It has been tested on four public board and committee packs, and a study with a participating board is designed.

The judgment on the record

Launching in autumn 2026
The method applied in public to one question, what will set the pace of AI data-centre construction in the United States to 2027. It will use the same three views as the dashboard, with every change sealed.

Reference series on AI

Running privately
Dated judgments on AI computing capacity and on AI policy and the economy, each with indicators that would show it holding or failing. They run privately as settings for the studies.

Research engines

Running privately
Systems that screen world news and model fifteen states and twenty AI firms, writing their predictions to the sealed record so that they can be checked later.