Leadership Under Uncertainty

Provable, not pitchable.

Every artifact below links to something you can audit, run, or verify without trusting us. Pick any one — there's no slide deck behind it.

NEW
Strategic overlay for Anthropic's Claude Tag
Watch the judgment layer catch a decision that contradicts a logged bet, check groupthink, and seal a call — live in a simulated Slack channel.
See the demo →

Live track record

These artifacts update over time and accumulate evidence. Hard to fake without leaving traces.

Forecast integrity chain

Every forecast is hashed, sealed daily into a Merkle tree, and committed to a public GitHub repo. Verify any forecast against a published root.

Updates: Updates daily at 23:59 UTC
How to verify: SHA-256 + Merkle proof + GitHub commit timestamp. Offline verifier available.

Live anchored forecasting tournament

System forecasts on real Manifold and Polymarket questions, each cryptographically anchored at creation. System Brier vs. market Brier shown for every resolved question.

Updates: New forecasts every 6 hours; resolutions polled hourly
How to verify: Each forecast links to its anchor commit. Outcomes pulled directly from the prediction market.

Calibration leaderboard

Public rankings by Brier score and calibration error. Filterable by tier, timeframe, and domain.

Updates: Refits weekly on Sunday 05:30 UTC
How to verify: Brier scores derived from resolved predictions in the integrity chain.

Source credibility leaderboard

Which news outlets, think tanks, and analysts actually predict accurately? Rankings derived from the predictive value of sources cited in resolved forecasts.

Updates: Recomputes nightly; decays 0.95/month without new data
How to verify: Predictive value = baseline Brier − Brier when source cited. 95% CI bounds, minimum 10 resolved predictions per source.

Fed-Watch

Live FOMC forecasts vs CME Fed Funds futures-implied probabilities. System Brier vs market Brier per meeting. Each forecast cryptographically anchored at creation.

Updates: Updated as Fed Funds futures move; resolved at each FOMC meeting
How to verify: Multi-class Brier vs CME Fed Funds futures on the same outcome space. Forecast hash + GitHub anchor commit URL per meeting.

Earnings-Watch

Pre-earnings strategic briefs on top US public companies. EPS direction + price reaction forecasts. ICD 203-graded prior-quarter forward-looking commentary.

Updates: Refresh 24-72h before each report; resolves at earnings + 1 day
How to verify: Multi-class EPS Brier (beat / miss / in-line) and 5-class price-reaction Brier. Each forecast hashed before earnings release.

Pre-registered predictions

Every public forecast hashed (SHA-256), sealed daily into a Merkle tree, and committed to GitHub BEFORE resolution. Cherry-picking is mechanically impossible.

Updates: Refresh daily after 23:59 UTC sealing job
How to verify: Each entry shows content hash, seal date, Merkle root, and anchor commit URL. Brier shown for resolved entries.

Research

Open standards and benchmarks. Submitted to or in submission for top venues.

Strategic Cognition Benchmark leaderboard

Frontier-model evaluations on the SCB benchmark — 205 tasks across 5 dimensions of strategic cognition, grounded in ICD 203.

Updates: Refreshed when new models are evaluated
How to verify: Each entry is signed by the SCB scoring service. Track filter (core / live / hidden) shows which evaluation tier was used.

Strategic Cognition Benchmark — paper

v15 candidate specification (~35,000 words). 205 tasks, JCS non-compensatory gating, ICD 203 grounding, BetterBench self-assessment 91/100.

Updates: v15 — submission to AIES 2027 in progress
How to verify: Open spec, CC-BY-SA. SCB-Core 25-task starter pack ships with the paper.

Strategic Cognition Ontology — paper

v21 candidate specification (~28,000 words). 35 OWL classes, 18 SKOS schemes, 8 Exchange Profiles, A2A protocol extension.

Updates: v21 — submission to ISWC 2027 in progress
How to verify: JSON-LD + OWL 2 RL + SHACL machine-readable artifacts published.

Capability Cards

Per-AI-system capability summaries (SCI / SCI-H, archetype, dimension breakdown, strengths, limitations). Citable at /capability-cards/<system>/<date>.

Updates: Refreshed as SCB evaluations land
How to verify: Each card derived from a signed SCB result entry. Citation string stable; URL slug includes evaluation date.

Interactive demos

Try the platform on your own input. No account needed for a sample run.

Strategic overlay for Claude Tag

A simulated Slack channel showing the judgment layer respond on top of Anthropic's Claude Tag: calibrated forecasts, ambient contradiction catches over logged assumptions, groupthink checks, and Merkle-sealed decisions captured in-thread.

Updates: Static simulation; the capabilities shown are production primitives
How to verify: Each moment maps to a live primitive — calibrated forecasting, the declared-vs-revealed self-mirror, ICD 203 scoring, and the Merkle attestation chain. Default deployment holds no Slack token.

SCO conformance demo

Enter a strategic topic. Watch the system generate validated SCO artifacts (Landscape, Drivers, Predictions, Scenarios) in real-time, with conformance violations surfaced.

Updates: On-demand
How to verify: JSON-LD + SHACL validation runs client-visible. Conformance violations linked to specific shape rules.

Brief Critic — ICD 203 public lab

Paste any analytical document. Get an ICD 203 tradecraft scorecard with paragraph-level findings. Share results via permanent link.

Updates: On-demand
How to verify: Same scoring engine as the GA Word add-in. Each result hash is shareable; tampering changes the hash.

OKR strategic twin

Paste a Google Sheets URL or upload a CSV of your OKRs. See them as a strategic landscape with forecast probabilities and external-signal binding.

Updates: Refreshes on every paste; Cascade / ClearPoint OAuth available for live sync
How to verify: Output cryptographically anchored. CSV upload deleted after demo unless explicit opt-in.

WhatsApp bot demo

Daily intelligence briefs over WhatsApp. Voice messages transcribed, audio briefs delivered, war-game scenarios you can play through interactively.

Updates: Live conversation
How to verify: Twilio Business API, 24-hour template window, audio via ElevenLabs multilingual.

Multi-persona memo critic

Paste a strategic memo. Run it through 4 personas in parallel: McKinsey partner, Sequoia GP, Pentagon J5 strategist, Chief Risk Officer.

Updates: On-demand
How to verify: Each persona surfaces different failure modes (MECE / TAM / branches & sequels / tail risk). Stub heuristic deterministic; LLM mode ICD 203-graded.

Storyline arc tracker

8 industry-shaping narratives (Fed pivot, China-Taiwan, AI Act, energy transition, etc.) tracked from emergence → rising → climax → resolution. Posture, signposts, and assumption decay.

Updates: Refreshes when an arc state changes or a signpost hits
How to verify: Each storyline shows assumption initial-vs-current confidence and recent-update trajectory impact.

Watch the AI think

ADW Delphi decomposition trace step-by-step. Semantic parse → decomposition → base-rate anchor → panel vote → adversarial round → ensemble synth → ICD 203 check → signposts.

Updates: Sample trace evergreen; live ADW streaming on operator wiring
How to verify: Each step has an output hash for chain-of-custody. 10-step pipeline; voter votes shown per sub-question.

Counterfactual sandbox

Pick a forecast. Toggle off individual sources / signals / assumptions. See the probability shift live, with attribution math per element.

Updates: On-demand interactive
How to verify: Pure-function recompute over removal deltas. Shows which evidence is actually load-bearing.

Scenario room

Pick a strategic question. Get 3-5 distinct futures with conditional probabilities, decision implications, pre-mortem signals, and key drivers.

Updates: Sample questions evergreen; live decompositions on operator wiring
How to verify: Probabilities sum to 1.0 (within 0.02 tolerance, validated). Each scenario carries a falsifiable signpost set.

Source vs source

Pick any two news sources. Head-to-head predictive value (Brier improvement vs uninformed base rate) with 95% CIs and per-category breakdown.

Updates: Refreshes weekly via source_quality_update job
How to verify: Statistical significance flagged when CIs do not overlap. Sample size (≥30 resolved) gating before publishing.

Posture matrix

For each strategic situation, system recommends a posture (Act / Probe / Watch / Hedge / Exit) with explicit rationale, the regret accepted, and 3 leading indicators that would shift the call.

Updates: Refreshes as situations update; stable matrix view
How to verify: Postures derived from strategic_views.posture; signposts from storyline arc primitives. Each call is a falsifiable artifact tied to specific triggers.

Information gaps

Paste a thesis. System ranks the load-bearing unknowns by relevance × knowability ÷ cost. The research priority list — 'if you only have time to find out 3 things, here they are.'

Updates: On-demand
How to verify: Heuristic synthesizer is deterministic; same thesis → same gaps. Each gap carries a concrete research path.

Coherence check

Paste a strategy. System identifies internal contradictions, tensions, unstated assumptions, reasoning gaps, circular reasoning — with paragraph-level citations and suggested fixes.

Updates: On-demand
How to verify: Six classes of failures (direct-contradiction → scope-mismatch). Coherence score 0-100 with explicit weighting (-25 critical / -10 material / -3 minor).

Assumption audit

Paste a memo. System extracts every load-bearing belief, scores initial vs. current confidence, names the kill-criterion, and tracks decay across accumulated evidence.

Updates: On-demand
How to verify: Each assumption carries a kill-criterion (specific evidence that would tip you over). Weighted-drift stat heavily weights load-bearing assumptions.

Decision frame

Pose a strategic question. System emits the operational decision frame: who decides, options with cost/yield/reversibility, recommended option with rationale, change-of-mind criteria, deadline.

Updates: On-demand
How to verify: Recommended option picked by yield ÷ cost with reversibility bonus. Cross-links to /information-gaps for research priorities.

Pre-mortem

Paste a thesis. System runs adversarial passes producing failure modes — each with leading-indicator signposts and the load-bearing assumption that breaks. 'If I see X, I'm wrong.'

Updates: On-demand
How to verify: Inverts the prediction lens. Expected-loss score = sum(P × severity). Top-3 risks identified for monitoring focus.

Perspective pack

Two persona-trained reasoners attack the same question. System aligns claims, identifies agreement zones, and surfaces divergences with reconciliation paths. SCO §6 Exchange Profile in action.

Updates: On-demand
How to verify: Consensus declared only when 2+ agreement zones reach ≥75% mean confidence — guard against false consensus. Each divergence kind classified.

Landscape composer

Pose a strategic question. Compose the landscape graph: actors, drivers, dependencies, decision-points, actions, scenarios, signals, constraints. Live SCO SHACL validation as you toggle nodes.

Updates: On-demand interactive
How to verify: Validation surfaces structural issues (missing actors / drivers / decision-points, dangling edges, scenarios un-linked to drivers). Output exportable as JSON-LD per SCO ontology.

Governing one AI agent, end to end

One agent's strategic call passes through five governance surfaces: an ICD 203 tradecraft contract that blocks the sloppy draft, a critic whose own reliability is measured and whose verdict escalates when it is low, a hash-chained decision ledger, outcome scoring against what actually happened, and an offline auditor check.

Updates: Verbatim output of a real run; script re-runnable on demand
How to verify: Every hash, chain link and seal is real; the transcript includes a tamper test where a single edited field breaks the hash. The judgment inputs are authored — the page states exactly what is real and what is staged.

Developer surfaces

Install, integrate, and run yourself. Code is on GitHub or npm.

Connect to Claude Tag

Add the LUU strategic tools to your Claude Tag in Slack (Mode ①). Generate a key, add an MCP connector, and @Claude a verb. LUU holds no Slack token — it receives only task-scoped inputs.

Updates: Live MCP server; connector config evergreen
How to verify: Default deployment requests zero Slack scopes. Same 35-tool MCP server documented with a public JSON-RPC contract.

AI agent framework showcase

Three runnable agents — daily-briefing, enterprise-risk-monitor, geopolitical-advisor — built on CrewAI, LangChain, Semantic Kernel, and Google ADK.

Updates: Templates evergreen; live run latency depends on backend load
How to verify: Each template clones, installs, and runs without an LUU account beyond an MCP API key.

@luu/attestation-verifier

Offline verifier for LUU evidence packs. Zero network I/O, native crypto, timing-safe comparisons. Verifies content hashes, Merkle proofs, and Ed25519 signatures.

Updates: v0.1.0 awaiting @luu npm scope provisioning
How to verify: Run against any tampered evidence pack and watch verification fail. 18 unit tests pass.

MCP server (live)

Model Context Protocol server exposing strategic-intelligence tools to AI agents over JSON-RPC 2.0. Live endpoint; Tier 0 anonymous access available (5 queries/day per IP).

Updates: Live; canonical luu_* tools, legacy northbrief_* aliases sunset 2026-10-28
How to verify: Public OpenAPI / JSON-RPC 2.0. Cost estimates per tool published. Capability Cards exposed at scb://capability-cards/{system}/latest.

Infrastructure

The plumbing behind everything else. Verifiable on its own merits.

CloudEvents subscriber

Subscribe to platform events (decision-created, action-executed, outcome-observed, seal-anchored). HMAC-SHA256 signed, RFC-compliant CloudEvents 1.0.

Updates: Live event stream
How to verify: Each event signed with X-LUU-Signature. Test endpoint provided for synthetic events.

Public anchors repo

GitHub repository of daily Merkle-root commits. Append-only, signed commits, branch-protected. Walk the chain offline with the bundled examples/walk-chain.mjs.

Updates: Daily commits at 23:59 UTC
How to verify: Every commit signed with the LUU attestation machine identity. CC0 1.0 — hashes are facts, no attribution required.

Missing an artifact you'd expect to see here? Tell us. If we can produce it, we will.

Want this for your team?

Drop your email and we'll be in touch when access opens for your vertical.