What the project has built
The project’s software is built around the process the launch paper describes. The screens below show sample data, recorded output and the published test’s own results.
Strategic dashboard
The dashboard is built to carry the process for one user, from the question to the decision. The Strategic Living Script sets the question, three views hold the user’s judgments with their assumptions and rival explanations, and each morning a brief grades what changed by the attention it needs. Decisions are recorded with what they rest on, so that they can be reviewed against what happened.
Open the working demonstrationThe last morning of the demonstration’s sample week.
The sufficiency test, run on the 1962 estimate
The test checks whether a record carries the nine elements of the project’s common vocabulary, one question for each element in Table 4 of the launch paper. It checks structure and does not judge whether the content is right. Here it runs on the September 1962 estimate on Cuba, as the record stood on the day the President approved the flight over western Cuba. Remove an element to see the test name what the record no longer carries.
The verdicts are the published test’s own reports, computed when this page was built. Researchers can run the same test on any record through the project’s public interface at /api/v1/public/sufficiency-test.
Checking a sealed record
Records are fingerprinted when they are made, and each day’s fingerprints are combined and posted to a public repository. On the record check your browser recomputes a record’s fingerprint and compares it with the published value, so the check rests on a value published outside the project, and a record can also be checked offline.
Run the checkTradecraft review
The tradecraft review, the analyst’s desk check against ICD-203, reads a note, memo or briefing against the analytic standards of Intelligence Community Directive 203. It scores the draft, names the actions that would improve it most and marks each paragraph that falls short with the standard it misses. The screen shows the review’s recorded output on a deliberately weak note from the project’s test set.
Governing one AI agent, from draft to audit
The agent tools apply the judgment layer to an AI agent’s work. In this run an agent drafts a recommendation and a check holds it to the analytic standards. The verdict of a critic with a weak record goes to a second judge, and the oversight decision and the chain from decision to action are sealed so that an auditor can verify them. The run was recorded from a scripted demonstration, with the drafts, verdicts and outcome supplied as inputs.