Leadership Under Uncertainty

The record has not started. It begins on the day the system's starting confidences are fixed in it, and nothing is published before then. The question and the rule are fixed before the record begins, and each change shows its evidence and its reasons. The rule and what is not claimed are on the method page.

How the judgment is kept on the record

The question and its three assessments were fixed before the record began, and the system set a starting confidence for each from the previous three weeks of reports. Since then each confidence has moved only as the rule below allows. The rule, and the rules that derive each daily state, are published here word for word, and their fingerprints are sealed in the record’s first day.

Update rule, version v1

The rule, verbatim

Update rule, version 1 Each assessment carries a stated confidence. It starts at the confidence the system set for it when the record began, and it moves only as this rule allows. The rule was fixed before the record began. Changing any part of it creates a new version, fingerprinted and announced in the change log, and thresholds are never lowered to make an assessment move. Every change after the first day is made by this rule, and the daily check refuses any change the rule does not allow. 1. Evidence. Only a dated report with a web address from an outside source, or a reading of a public data series admitted for this purpose, can move an assessment. The project's own analysis never can. 2. Events. Reports of the same development are folded into one event, which counts once however many outlets repeat it. 3. Weight. An event bearing on a causal factor has the weight e = stance × sign × grade × materiality × diagnosticity. Stance is +1 when the event strengthens the factor and -1 when it weakens it. Sign is +1 when a stronger factor supports the assessment and -1 when it undercuts it. Grade is 1.0 for a source graded high, 0.6 for medium and 0.3 for unknown, taken from the best-graded outlet reporting the event. Materiality is 1.0 for a material development and 0.5 for a routine one. Diagnosticity is 1 when the event bears differently on the assessment than on at least one of its alternatives, and 0 when it fits them equally. An event with no diagnosticity carries no weight and is listed as considered and set aside. 4. Size. An assessment moves by 0.10 × Σe in log-odds, summed over the events applied to it. One high-grade, material, diagnostic event therefore moves it by 0.10 in log-odds, about 2.5 percentage points at even odds, and no single event can move it further. 5. Corroboration. An event moves an assessment only once at least 2 distinct domains have reported it within 7 days of its first report. Until then it is listed as noted and awaiting corroboration. After 7 days without a second domain it is listed as uncorroborated and moves nothing by itself. 6. Cooldown. A causal factor moves its assessment at most once in 48 hours. Several events on the same factor in one update move it together, once. Evidence on a factor inside its cooldown waits until the cooldown ends. 7. Daily limit. The net move of one assessment within one UTC day stays within 0.20 in log-odds in either direction. When the events ready to apply would take it further, they are taken in the order they were first reported, each applied only if it keeps the net move within the limit, and the rest wait for the next update. 8. Challenge. Before a material move is applied, a second model, from a different family than the one that read the event, looks for the strongest reason the reports do not support it. A move it disputes is not applied. It is recorded as set aside, with the challenge. A move whose challenge could not run waits for the next update. 9. Attention level. The attention level shown changes to the computed one once the computed level has held for 3 consecutive daily states. The activation of the contingent assessment changes the same way. 10. No drift. There is no decay and no pull toward even odds. A stated confidence stays where the evidence left it.

SHA-256 of the text above, as UTF-8: . The record has not started, so this text has not been sealed yet.

State rules, version v1

How each daily state is derived

How the daily state is derived, version 1 1. Starting point. Before the record begins, the system takes up to 10 of the developments it has linked to each assessment's causal factors over the previous 21 days, those reported by at least 2 distinct domains first, then the most widely reported, then the newest. Developments outside the record's scope are left out. A model reads them and chooses, for each assessment, one of five terms from the ICD 203 scale of likelihood, from very unlikely to very likely. When an assessment has no linked development, the model chooses from its general knowledge. Each assessment starts at the middle of its term's range: 12.5 per cent for very unlikely, 32.5 for unlikely, 50 for even chance, 67.5 for likely and 87.5 for very likely. The reading, the developments it was shown and the model that made it are kept in the first day's state. 2. Probability. A stated confidence held as log-odds L is shown as the probability 1/(1+e^-L), rounded to four decimal places, with its term from the ICD 203 scale of likelihood. 3. Thesis status. The thesis holds when the main assessment is at or above the level set for it, the assessment it rests on is at or above its level, and the contingent assessment is dormant. It is inverted when the main assessment is below the lower level set for it or the contingent assessment is active. Otherwise it is strained. The levels were fixed with the question before the record began. 4. Trend. Up or down when today's stated confidence differs by at least 5 percentage points from the mean of the previous 6 daily states, and stable otherwise, or when fewer than 3 daily states exist. 5. Stability and attention level. Each day, moves of at least 0.10 in log-odds in either direction in the previous 48 hours raise an assessment's stability from stable to watch. The contingent assessment is also at watch while it is computed as showing early signals, and stressed while it is computed as active. The attention level is then read from the published ladder, using stability, stated confidence and trend. The attention level shown changes to the computed one once the computed level has held for 3 consecutive daily states. 6. Contingent activation. Each day the contingent assessment is computed as dormant, showing early signals or active, from moves of at least 0.10 in log-odds toward it in the previous 48 hours. It steps down at most one level a day. The activation shown changes to the computed one once it has held for 3 consecutive daily states. 7. Record. One state is written for each UTC day, fingerprinted and linked to the one before it. A day the job missed is written later, marked as recorded late, and computed from the changes that had been recorded by that day's close. 8. Reading. Each state after the first also records what the update loop read in its window: how many reports it compared with the causal factors for the first time, how many it linked to a factor, how many repeated a development already reported, how many developments wait for a second outlet, and the developments it set aside because they fitted an assessment and each of its alternatives equally, listed up to 20 and counted beyond that. Reports excluded for scope are counted and never listed. This is a record of what was read. Nothing in it moves an assessment.

SHA-256 of the text above, as UTF-8: . The record has not started, so this text has not been sealed yet.

What is learned and what is fixed

Each step of the update loop, and how it is measured

Models find, fold and read the evidence. Only the rule moves a confidence, and the rule is not learned.

StepWhat it doesHow it is measuredDoes it learn?
Starting pointBefore the record begins, a model reads up to 10 linked developments for each assessment and chooses a term from the ICD 203 scale. The assessment starts at the middle of that term's range.The developments it was shown and the ones its reason cites, kept in the first day's record.No. It is read once, before the record begins.
AttentionFinds new outside reports that bear on a causal factor: a shortlist by text similarity, then a relevance check by a small model.The reports read and linked each day, in each daily state. Recall is not measured.No. The threshold is set in a parameter file, and every change records a fingerprint of that file.
ConsolidationFolds reports of the same development into one event, so it counts once however many outlets repeat it.The repeated reports folded each day, in each daily state.No. The threshold is set in a parameter file, and every change records a fingerprint of that file.
InterpretationReads what an event says about the causal factor, the assessment and each alternative. Evidence that fits the assessment and its rival equally carries no weight.The developments set aside as fitting an assessment and its alternatives equally.No.
Source gradingGrades each outlet high, medium or unknown.The share of evidence with a grade.No. The grades are fixed for the record.
RevisionApplies the published rule below to move a stated confidence.How often a move is reversed within seven days.No. A change to the rule is a new, fingerprinted version.
ChallengeA model from a different family argues against each material move. A disputed move is set aside, with the challenge on the record.The share of material moves it disputes.No.
Quality monitors

How the record's own behaviour is watched

Each monitor is measured from the sealed record alone and published whatever its value, with its n and the dates its data span.

Not yet measured: the record has not started.

Scope

What the judgment covers, and what it leaves out

The scope is published with the record’s first day.

Reports on topics the project keeps out of this record are excluded before they can move anything, and the public pages refuse to show a record that carries one. Only the count is published: no excluded report is listed, linked or quoted anywhere on these pages.

Tests we're watching

When a data series may count as a test

No data series is admitted as a test in this version of the rule. A later version may admit one only if, over all the readings it will get, its chance of firing on noise alone is 5 per cent or less. Hourly regional electricity demand, for example, fails that bar, because weather and the daily cycle move it more than any buildout would. The arithmetic and a replay are on when a resolution criterion is a real test.

Fingerprints

How records are fingerprinted

One daily state is recorded for each day in UTC, and one record for each change. Each names the one before it by its fingerprint, so a record cannot be rewritten without breaking the chain. At 23:59 UTC every such record not yet fingerprinted joins that day’s fingerprint of all of them, and the day’s fingerprint is pushed to a public repository on GitHub. When a day is missed, its records join the next day’s, and the pages give the date of the fingerprint that contains each record.

The state and the changes are published as data: the state and the changes, as JSON Lines, each record in full with its fingerprint, so anyone can recompute it. The project’s sealed protocols and pre-registrations are listed in the methods ledger.

What we do not claim

  • No accuracy, skill, calibration, Brier score, hit rate or track record of any kind is claimed. Nothing here predicts events, and nothing here is said to beat experts or markets.
  • The numbers are not measured probabilities. They are stated judgments, set by the system from the three weeks of reports before the record began and revised since then only by a published rule.
  • The system did not choose the question. The thesis, its assessments and the rule were fixed before the record began, and the system works within them.
  • Nothing is learned from outcomes. There is no learned model of source quality and no reliability score for the challenge step.
  • The record is not real-time: it is updated once a day. It is not full-text: the update loop reads headlines and summaries. It is not comprehensive: it covers one slice of the United States, drawn from a web-search service and public datasets. It does not establish causation.
  • A timestamped fingerprint does not show that a judgment was right. It shows that the record existed no later than the day its fingerprint was published. When a day's fingerprint is missed, that day's records join the next day's, and the pages give the date of the fingerprint that contains each record.
  • Nothing is reconstructed after the fact. The only material from before the first day is the evidence the starting point was read from, kept in the first day's record.