When is a resolution criterion a real test?
A criterion that settles whether a claim has come true should rarely be met when nothing has happened. This page works out how often a criterion labelled “three standard deviations” is met on noise alone. It then replays one of the project’s own criteria, since withdrawn, on the public electricity-demand series it was written against, using only the readings that had arrived by each date.
The page calls a series in which nothing is happening a null. A normal and independent null is one whose readings scatter at random around a fixed level, each unrelated to the last.
Three standard deviations, checked many times
The 532 is the number of readings in the withdrawn supporting criterion’s thirty-day window on 25 July 2026, including revised figures counted a second time, as given for this criterion in the gate’s own code that night (commit 75b7ae9d9, 22:29 UTC). They cover 413 of the window’s 720 hours. The replay below reproduces both counts. The curve shows the same arithmetic for any number of checks. Drag the handle or use the arrow keys.
The curve as a table (3.0σ)
| Checks | Chance of at least one firing |
|---|---|
| 1 | 0.1% |
| 10 | 1.3% |
| 30 | 4.0% |
| 100 | 12.6% |
| 168 | 20.3% |
| 532 | 51.3% |
| 720 | 62.2% |
| 1,000 | 74.1% |
| 2,160 | 94.6% |
| 3,000 | 98.3% |
- Checks in the window
- 532
- Chance of firing on one check
- 0.135%
- Chance of at least one firing on noise
- 51.3%
- Times the label’s rate
- 380
- Expected false firings
- 0.72
- Against the project’s 5 per cent ceiling
- Refused
- Smallest threshold the gate would admit at 532 checks
- 3.73σ
532 counts rows, including revised figures the gate counted a second time. Readings arrived for 413 of the window’s 720 hours. Moving any control recounts in hours.
A pair of criteria
On noise alone the falsifying side is 117.5 times easier to meet. The project’s ceiling on a pair’s imbalance is 3 to 1.
The same criterion on two dates
Between 2 July and 20 July 2026 the project sealed 37 assessments whose resolution criteria were written against hourly electricity demand on the grid run by MISO, the operator for much of the U.S. Midwest, as published by the U.S. Energy Information Administration. The supporting criterion fired if demand rose 3σ above its baseline within thirty days. The falsifying criterion fired if demand fell 1σ below within twenty-one. The assessments belong to a reference feed that is not public. Only the criteria’s form and the public series are shown here. The file below is the project’s copy of the series as it was ingested, with the time each value arrived, which is what the gate saw.
window 2026-06-25 13:00 to 2026-07-25 13:00 UTC
The window as a table, by local day
| Day (CDT) | Readings | Lowest MW | Highest MW | Supporting fired | Falsifying fired |
|---|---|---|---|---|---|
| 2026-06-28 | 3 | 90,370 | 97,976 | 0 | outside window |
| 2026-06-29 | 24 | 75,960 | 119,185 | 9 | outside window |
| 2026-06-30 | 24 | 82,912 | 121,514 | 9 | outside window |
| 2026-07-01 | 9 | 86,228 | 98,007 | 0 | outside window |
| 2026-07-02 | 2 | 98,057 | 103,393 | 0 | outside window |
| 2026-07-03 | 24 | 77,844 | 109,432 | 0 | outside window |
| 2026-07-04 | 22 | 71,321 | 101,987 | 0 | 0 |
| 2026-07-10 | 19 | 75,683 | 109,498 | 0 | 0 |
| 2026-07-11 | 47 | 71,905 | 103,185 | 0 | 0 |
| 2026-07-12 | 47 | 68,432 | 108,099 | 0 | 6 |
| 2026-07-13 | 22 | 73,934 | 116,702 | 6 | 0 |
| 2026-07-15 | 2 | 99,702 | 105,647 | 0 | 0 |
| 2026-07-16 | 24 | 79,934 | 115,311 | 6 | 0 |
| 2026-07-17 | 45 | 79,103 | 120,031 | 15 | 0 |
| 2026-07-18 | 45 | 79,164 | 114,473 | 7 | 0 |
| 2026-07-19 | 24 | 72,150 | 109,891 | 0 | 0 |
| 2026-07-20 | 24 | 74,809 | 114,803 | 5 | 0 |
| 2026-07-21 | 45 | 81,165 | 113,853 | 5 | 0 |
| 2026-07-22 | 33 | 73,665 | 98,675 | 0 | 0 |
| 2026-07-23 | 24 | 71,024 | 97,280 | 0 | 0 |
| 2026-07-24 | 23 | 71,338 | 100,282 | 0 | 0 |
As the data stood at 13:00 UTC on 25 July 2026, fifteen days after 10 July and with no change to the claim, the supporting criterion had fired on 62 of 532 readings (11.65 per cent) and the falsifying criterion on 6 of 439 (1.37 per cent). The pair now leaned 8.5 to 1 towards confirmation. July demand rose above a baseline drawn from June readings.
| Supporting: 3σ above, 30-day window | Falsifying: 1σ below, 21-day window | |
|---|---|---|
| Baseline readings | 431 | 525 |
| Baseline mean | 80,332.2 MW | 83,159.5 MW |
| Baseline standard deviation | 10,372.3 MW | 12,717.1 MW |
| Threshold | 111,449 MW | 70,442 MW |
| Readings in the window | 532 | 439 |
| Distinct hours | 413 | 320 |
| Firings | 62 | 6 |
| Firing rate | 11.65% | 1.37% |
| Verdict under the gate’s thresholds | Refused: the two sides differ by more than 3 to 1, fires on more than 5 per cent of readings, chance of a firing on noise above 5 per cent | Refused: the two sides differ by more than 3 to 1, chance of a firing on noise above 5 per cent |
The two sides differ by 8.5 to 1, towards confirmation. Ceiling: 3 to 1.
Recorded, and recomputed here
| Firing rate, per cent | Recorded | Recomputed |
|---|---|---|
| Supporting, as of 10 July | 0.44 | 0.44 |
| Falsifying, as of 10 July | 7.04 | 7.04 |
| Supporting, as of 25 July | 11.65 | 11.65 |
| Falsifying, as of 25 July | 1.37 | 1.37 |
Recorded in the project’s review of 25 July 2026 (repository commit 2315470e3, 19:30 UTC). Recomputed in your browser from the published file with the gate’s own function.
The review recorded dates only. 13:00 UTC is the hour at which the replay reproduces all four rates, and on 10 July it is the only such hour.
Where the replay qualifies the record
- The gate counts every row, including revisions. When EIA revised an hour’s figure the revision arrived as a new row, and both were counted: the 532 readings of 25 July cover 413 distinct hours. Counting each hour once, with its latest figure, the supporting criterion fired in 30 of 413 hours (7.26 per cent). The gate’s verdict is the same either way.
- The number of readings depends on the moment the gate runs, because revisions arrive and old hours leave the window. Checked hourly from 10 July to 26 July, the supporting window held between 413 and 577 readings. At the replay instant on 25 July it held 532, the count the gate’s own code gave for this criterion that night (commit 75b7ae9d9, 22:29 UTC), and the arithmetic above uses that count. Draft 11 of the paper put the count at 526, and Drafts 12 and 13 give 532.
- Over each criterion’s own window (30 days supporting, 21 falsifying), the falsifying criterion fired only in the early morning, around the daily low in demand, and on few days: 3 of the 12 days in its window as of 10 July and 1 of 15 as of 25 July, always between 02:00 and 08:59 Central Daylight Time. The gate itself evaluates both sides of a pair over one window, the longer of the two, here 30 days. Over that window the falsifying criterion fired on 113 of 457 readings (24.73 per cent) as of 10 July, on 10 of 19 days, between 00:00 and 10:59 Central Daylight Time, and on 4 of 532 (0.75 per cent) as of 25 July, on 1 of 21 days. Counted that way the pair leaned 56.5 to 1 towards falsification and then 15.5 to 1 towards confirmation. Draft 11 of the paper said criteria of this kind were met every night by the ordinary overnight trough. The replay bears out the hour of the firings, and in neither window did the criterion fire every night. Draft 12 said only that the criterion fired in the early-morning low, and Draft 13 does not give the hour.
When consecutive readings are correlated
The 51.3 per cent assumes independent checks. Hourly demand is not independent: between 29 June and 26 July 2026 the correlation between consecutive hours was 0.971. In a simulated series with no change in level but each hour correlated with the last, the chance falls, to 17.8 per cent at 0.95 and 9.1 per cent at 0.98, and at 0.97 it is 12.3 per cent, above the 5 per cent ceiling. A common shortcut, treating n(1−ρ)/(1+ρ) checks as independent, would give 3.7 per cent at ρ = 0.9 where simulation gives 27.3 per cent. This page does not use it. Neither model has the daily cycle that real demand has, which is why the replay above carries the argument.
| Correlation between consecutive readings (ρ) | Chance of at least one firing |
|---|---|
| 51.2% | |
| 48.8% | |
| 37.8% | |
| 27.3% | |
| 17.8% | |
| 12.3% | |
| 9.1% | |
| 5.2% |
Five thresholds, each with its reason
- A pair’s two sides may not differ in realized firing rate by more than 3 to 1.Reason given in the code: “Observed: the MISO pair sat at ~16:1 at mint [when the criterion was issued] and had inverted to ~8.5:1 fifteen days later.”Measured later on this page: the replay puts the pair at 16.1 to 1 in one direction as the data stood on 10 July and 8.5 to 1 in the other on 25 July, each side counted over its own window. Over the gate’s single 30-day window the figures are 56.5 and 15.5.
- A side may not fire on more than 5 per cent of readings.Reason given in the code: “A side firing on more than this share of observations is saturating, not detecting.”
- The chance of at least one firing on noise, over the criterion’s actual number of checks, may not exceed 5 per cent.Reason given in the code: “The single largest error in the original design: a ‘3σ’ criterion scanned hourly for 30 days sits at ~51% here against the 0.135% its label implies.”Measured later on this page: the window held 532 hourly readings on 25 July, not a full 720, and at 532 checks a 3σ criterion fires on noise about half the time.
- A baseline needs at least 8 distinct values.Reason given in the code: “Below this, a sigma is describing a near-constant series, not rarity.”
- A baseline needs at least 30 readings.Reason given in the code: “Below this, the baseline statistics are not worth acting on.”
The thresholds were committed, together with the result of the gate’s first run, at 22:29 UTC on 25 July 2026 (75b7ae9d9). The repository cannot show that they were written before that run. The reasons above are quoted from that commit, with the editor’s gloss in square brackets. The commit message says seven thresholds were declared. The code it committed declares five, the ones listed here, and nothing else in that commit accounts for the other two. The gate first ran in a mode that recorded its verdicts without refusing anything. A change committed at 06:13 UTC on 28 July 2026 (e8411eb8a) made refusal the default for any criterion that breaks one of them. The dates come from the project’s own repository and are not externally anchored.
What counting past crossings missed
The five thresholds count how often a criterion would have fired on the readings already in its window. A review in September 2026 found that this count is blind to drift. Thirty criteria behind the project’s most decisive scored beliefs had passed the gate with an estimated chance of firing on nothing of 1.2 to 4.0 per cent, and each then fired on 52 to 100 per cent of the historical dates at which it could have been issued, because their series were trending. On 26 September 2026 the gate gained a second test (commit bd6432831). It replays each proposed criterion over its own series’ history, issued at every earlier date it could have been, and refuses it when the simulated chance of firing on nothing exceeds 5 per cent, or when the history is too short to tell. It also refuses any series nobody has reviewed as a fit measure for the claim. The five thresholds are unchanged. The figures here come from that commit’s own record and have not been recomputed on this page.
Sealed withdrawals
The project’s log records 81 criteria across 49 open assessments withdrawn on the evening of 25 July 2026, and 2 more that night, 83 across 51 in all. The withdrawal records that survive from that evening were written between 20:37 and 20:38 UTC and carry 78 criteria across 46 open assessments. They were folded into a later daily seal. The records for the other assessments do not survive. A defect disclosed on 28 July 2026 deleted integrity records of assessments that had already been sealed, and it is the likely cause. Further withdrawal records were written on 28 July, 15 August, 29 August and 30 August 2026. As of 26 September 2026, 192 criteria across 127 assessments carry the sealed outcome “withdrawn as non-discriminating”. Draft 11 of the paper gave the log’s figure of 83 criteria across 51 assessments. Draft 12 gave the sealed figure of 78 across 46, and Draft 13 gives the total withdrawn by the end of August, 192 across 127.
The first withdrawal came before the five thresholds existed. It applied a rule fixed earlier that day: a criterion was withdrawn if its query failed structural validation or named a series that the project’s registry marks as unsuitable for settling claims. Each withdrawal is a new sealed record. The original criteria are never edited, and windows are not extended to make room for replacements.
| Date | Time | Criteria | Assessments |
|---|---|---|---|
| 25 July 2026 | 20:37–20:38 | 78 | 46 |
| 28 July 2026 | 06:21–09:40 | 9 | 9 |
| 15 August 2026 | 14:22–14:23 | 10 | 10 |
| 29 August 2026 | 11:33–11:34 | 54 | 45 |
| 30 August 2026 | 01:14–14:17 | 41 | 41 |
An assessment can appear on more than one date. Total: 192 criteria across 127 assessments.
| Source | Form | Criteria |
|---|---|---|
| Public code repositories (GitHub) | a set number of standard deviations | 77 |
| U.S. electricity demand (EIA) | a set number of standard deviations | 77 |
| Public code repositories (GitHub) | fixed level | 28 |
| Company filings (SEC) | not machine-checkable | 4 |
| U.S. electricity demand (EIA) | fixed level | 3 |
| Company filings (SEC) | a set number of standard deviations | 2 |
| Federal Register | fixed level | 1 |
73 of the 192 have the form of the MISO pair replayed above: 37 supporting criteria (3σ above, 30 days) and 36 falsifying criteria (1σ below, 21 days).
Counts exported from the sealed records at 14:52 UTC on 26 September 2026.
In 1962 John McCone, the Director of Central Intelligence, reportedly read the new air-defence sites in western Cuba as meant to blind American reconnaissance. On 1 October Colonel John Wright, a Defense Intelligence Agency officer, noted that three of those sites could not be connected with any significant military installation. The reading pointed to an observation that could settle it, a single overflight of western Cuba, which photographed the launchers on 14 October (see Section 2 of the launch paper). The criterion gate asks something similar of every machine-checkable criterion: that it should rarely be met when nothing has happened.
Code and data
The gate imports these same functions. A test fails if the page and the gate diverge. Data files: the MISO series as ingested (SHA-256, source EIA-930, exported 26 September 2026) and the simulation output.
The file below is lib/insights/criterionPowerMath.ts as it stands today. The thresholds block in it, comments included, is byte for byte as committed in 75b7ae9d9, and a test pins it. Comments elsewhere in the file were revised on 26 September 2026, and a dated note after the block corrects two of its figures. The threshold values are unchanged since 75b7ae9d9.
/**
* The pure arithmetic of the criterion gate.
*
* A resolution criterion of the form "the metric moves k standard deviations
* away from its baseline" is only informative if that move is rare when nothing
* has happened. Everything in this file answers that question with closed-form
* arithmetic over numbers the caller supplies: a normal tail probability, the
* chance of at least one firing across many checks, the realized firing rates of
* a pair of criteria, and the pre-registered thresholds the gate judges them
* against.
*
* WHY THIS FILE HAS NO IMPORTS
* The gate (`criterionPower.ts`) reads the observations table and re-exports
* everything here. The public methods page at /methods/resolution-criteria
* imports this file directly and runs it in the reader's browser. One
* implementation serves both, so the page computes with the same function that
* decides admissibility, and a test fails if the two diverge
* (tests/insights/criterionPowerMath.test.ts). Keep it free of imports: adding
* one would pull server code into a static page.
*
* @module lib/insights/criterionPowerMath
*/
/** One side of a supports/falsifies pair. */
export interface NullPowerSide {
sigma: number;
direction: "above" | "below";
/** Observations in the scan window: the number of chances to fire. */
scanN: number;
/** How many of them actually crossed the threshold. */
crossings: number;
/** crossings / scanN, the REALIZED rate, which is what matters. */
fireRate: number;
/** Per-observation p under a normal null. What the "k sigma" label implies. */
nominalP: number;
/** 1 - (1 - nominalP)^scanN: the honest probability of >=1 fire on noise. */
familywiseP: number;
}
export interface NullPowerReport {
metricKey: string;
dimensions: Record<string, unknown> | null;
/** The instant the report describes; both issue date and today are useful. */
asOf: string;
baseline: {
n: number;
mean: number | null;
stddev: number | null;
/** Few distinct values => a sigma is not a rarity measure. */
distinctValues: number;
};
sides: NullPowerSide[];
/**
* max(fireRate) / max(min(fireRate), floor). Null when a side is missing.
* >3 means the pair is structurally biased toward one verdict.
*/
asymmetryRatio: number | null;
/** Human-readable observations, safe to publish. */
notes: string[];
}
/** Standard normal CDF via Abramowitz & Stegun 7.1.26 (|error| < 1.5e-7). */
export function normalCdf(z: number): number {
const sign = z < 0 ? -1 : 1;
const x = Math.abs(z) / Math.SQRT2;
const t = 1 / (1 + 0.3275911 * x);
const y =
1 -
((((1.061405429 * t - 1.453152027) * t + 1.421413741) * t - 0.284496736) *
t +
0.254829592) *
t *
Math.exp(-x * x);
return 0.5 * (1 + sign * y);
}
/** One-sided p under a normal null for a k-sigma threshold. */
export function nominalTailP(sigma: number): number {
return 1 - normalCdf(Math.abs(sigma));
}
/**
* P(at least one crossing | null) across `scanN` chances.
*
* This is the number that should appear next to any "k sigma" claim. Treating
* the scans as independent is generous to us. Hourly grid demand is strongly
* autocorrelated, which REDUCES the effective number of independent trials, so
* this is an upper bound on the familywise rate and a lower bound on how bad the
* overstatement is. Stated that way on purpose.
*/
export function familywiseP(sigma: number, scanN: number): number {
if (scanN <= 0) return 0;
return 1 - Math.pow(1 - nominalTailP(sigma), scanN);
}
/** Pure core: given values + baseline stats, what would each side have done? */
export function assessCriterionPair(
scanValues: number[],
baseline: { mean: number | null; stddev: number | null },
sides: Array<{ sigma: number; direction: "above" | "below" }>,
): { sides: NullPowerSide[]; asymmetryRatio: number | null } {
const { mean, stddev } = baseline;
const out: NullPowerSide[] = sides.map((s) => {
let crossings = 0;
if (mean !== null && stddev !== null && stddev > 0) {
for (const v of scanValues) {
const z = (v - mean) / stddev;
if (s.direction === "above" ? z >= s.sigma : z <= -s.sigma) crossings++;
}
}
const scanN = scanValues.length;
return {
sigma: s.sigma,
direction: s.direction,
scanN,
crossings,
fireRate: scanN > 0 ? crossings / scanN : 0,
nominalP: nominalTailP(s.sigma),
familywiseP: familywiseP(s.sigma, scanN),
};
});
let asymmetryRatio: number | null = null;
if (out.length >= 2) {
const rates = out.map((s) => s.fireRate);
const hi = Math.max(...rates);
const lo = Math.min(...rates);
// Floor the denominator at one observation's worth of rate so a
// never-fires side yields a large finite ratio rather than Infinity.
const floor = out[0].scanN > 0 ? 1 / out[0].scanN : 1;
asymmetryRatio = hi === 0 ? 1 : hi / Math.max(lo, floor);
}
return { sides: out, asymmetryRatio };
}
/**
* PRE-REGISTERED rejection thresholds.
*
* Declared here, with a justification each, BEFORE being run against the live
* population — deliberately, because "tune until it rejects exactly the known-bad
* set and nothing else" is curve-fitting wearing a gate's clothes, and it is the
* first thing a sceptical reader would accuse us of. Whatever the gate rejects
* when it runs is an OUTCOME to be reported, including surprises.
*
* Each threshold maps to an observed failure, not to a preference.
*/
export const CRITERION_POWER_THRESHOLDS = {
/**
* Realized firing rates of a supports/falsifies pair may not differ by more
* than this. Observed: the MISO pair sat at ~16:1 at mint and had inverted to
* ~8.5:1 fifteen days later. NOT a nominal |k|-equality rule — on a skewed,
* capacity-bounded series equal thresholds give unequal rates by construction
* (43.5% at or above +1σ against 1.37% at or below −1σ).
*/
maxPairAsymmetryRatio: 3,
/**
* A side firing on more than this share of observations is saturating, not
* detecting. 5% is one observation in twenty — already far from "anomaly".
*/
maxSideFireRate: 0.05,
/**
* P(at least one fire | null) across the criterion's actual scan count. The
* single largest error in the original design: a "3σ" criterion scanned hourly
* for 30 days sits at ~51% here against the 0.135% its label implies.
*/
maxFamilywiseP: 0.05,
/** Below this, a sigma is describing a near-constant series, not rarity. */
minBaselineDistinctValues: 8,
/** Below this, the baseline statistics are not worth acting on. */
minBaselineN: 30,
} as const;
/*
* Note dated 26 September 2026, kept outside the declared block above so that
* block stays as committed in 75b7ae9d9 (25 July 2026, 22:29 UTC). Two of its
* justifications are loose: the MISO pair's ~16:1 describes the data as it
* stood on 10 July 2026, not the moment of issue, and the "3σ" criterion's
* 30-day window held 532 hourly readings on 25 July 2026 rather than a full
* 720. The threshold values are unchanged. See /methods/resolution-criteria.
*/
interface CriterionPowerVerdict {
/** False ⇒ the criterion should not be issued (in `enforce`). */
admissible: boolean;
/** Machine-readable rejection codes, empty when admissible. */
violations: string[];
report: NullPowerReport;
}
/**
* Judge a criterion (or a pair) against the pre-registered thresholds.
*
* Fail-OPEN on missing data: a report with no scan observations cannot show a
* criterion is bad, and refusing to issue on absent data would silently empty the
* portfolio whenever an adapter stalls. Absence of evidence is reported in
* `notes`, not converted into a rejection.
*/
export function judgeCriterionPower(
report: NullPowerReport,
): CriterionPowerVerdict {
const t = CRITERION_POWER_THRESHOLDS;
const violations: string[] = [];
const measurable = report.sides.some((s) => s.scanN > 0);
if (measurable) {
if (
report.asymmetryRatio !== null &&
report.asymmetryRatio > t.maxPairAsymmetryRatio
) {
violations.push(
`pair_asymmetry:${report.asymmetryRatio.toFixed(1)}x>${t.maxPairAsymmetryRatio}x`,
);
}
for (const s of report.sides) {
if (s.scanN === 0) continue;
if (s.fireRate > t.maxSideFireRate) {
violations.push(
`saturating:${s.direction}${s.sigma}sigma@${(s.fireRate * 100).toFixed(1)}%`,
);
}
if (s.familywiseP > t.maxFamilywiseP) {
violations.push(
`familywise:${s.direction}${s.sigma}sigma@${(s.familywiseP * 100).toFixed(0)}%over${s.scanN}scans`,
);
}
}
if (report.baseline.n > 0 && report.baseline.n < t.minBaselineN) {
violations.push(`thin_baseline:n=${report.baseline.n}`);
}
if (
report.baseline.distinctValues > 0 &&
report.baseline.distinctValues < t.minBaselineDistinctValues
) {
violations.push(
`low_cardinality:${report.baseline.distinctValues}values`,
);
}
}
return { admissible: violations.length === 0, violations, report };
}
/**
* Per-SIDE admissibility within a group report.
*
* WHY THIS EXISTS. {@link judgeCriterionPower} answers "is this SERIES' criterion
* set admissible", which is the right unit for the pair-level checks and the
* wrong unit for a decision about one row. Several criteria written by
* DIFFERENT generators can land on the same series. Live example, 2026-07-27:
* `github.NVIDIA.Megatron-LM.commits_7d` carried a legacy 3σ criterion
* (familywise 3.5%, admissible) and a new 2.5σ one (familywise 15.0%, not).
* Judging the group and acting on the group verdict either withdraws the sound
* criterion along with the unsound one, or keeps both. Neither is correct.
*
* Side-scoped checks (`saturating`, `familywise`) are attributed to the side
* they describe. Baseline checks (`thin_baseline`, `low_cardinality`) and the
* PAIR check (`pair_asymmetry`) are properties of the series or of the pair as a
* unit, so they disqualify every side in the group. An asymmetric pair has no
* innocent half.
*/
export function judgeCriterionSide(
report: NullPowerReport,
side: { sigma: number; direction: "above" | "below" },
): { admissible: boolean; violations: string[] } {
const t = CRITERION_POWER_THRESHOLDS;
const violations: string[] = [];
// Group-scoped: shared by every side on this series.
if (
report.asymmetryRatio !== null &&
report.asymmetryRatio > t.maxPairAsymmetryRatio
) {
violations.push(
`pair_asymmetry:${report.asymmetryRatio.toFixed(1)}x>${t.maxPairAsymmetryRatio}x`,
);
}
const measurable = report.sides.some((s) => s.scanN > 0);
if (measurable) {
if (report.baseline.n > 0 && report.baseline.n < t.minBaselineN) {
violations.push(`thin_baseline:n=${report.baseline.n}`);
}
if (
report.baseline.distinctValues > 0 &&
report.baseline.distinctValues < t.minBaselineDistinctValues
) {
violations.push(
`low_cardinality:${report.baseline.distinctValues}values`,
);
}
}
// Side-scoped: only this side's own realized behaviour.
const own = report.sides.find(
(s) => s.direction === side.direction && s.sigma === side.sigma,
);
if (own && own.scanN > 0) {
if (own.fireRate > t.maxSideFireRate) {
violations.push(
`saturating:${own.direction}${own.sigma}sigma@${(own.fireRate * 100).toFixed(1)}%`,
);
}
if (own.familywiseP > t.maxFamilywiseP) {
violations.push(
`familywise:${own.direction}${own.sigma}sigma@${(own.familywiseP * 100).toFixed(0)}%over${own.scanN}scans`,
);
}
}
return { admissible: violations.length === 0, violations };
}
/**
* The sigma BOTH sides of a supports/falsifies pair must use.
*
* THE DEFECT THIS CLOSES. All three criterion generators paired a 3σ supporter
* with a 1σ falsifier. Under a normal null those are not comparable tests: the
* one-sided tail p is 0.001350 at 3σ against 0.158655 at 1σ, so the falsifier
* was **117.5x** easier to satisfy than its partner. Against a declared
* `maxSideFireRate` of 0.05 that is not a rounding error. It is a portfolio
* biased toward one verdict by construction, which is worse for a published
* calibration record than being wrong, because the bias survives every honest
* effort to score it. An earlier fix equalised the two sides' WINDOWS
* (`validateWindowSymmetry`) and left the sigmas untouched.
*
* WHY 3 AND NOT SOMETHING SMALLER. It falls out of the pre-registered
* `maxFamilywiseP` of 0.05. A criterion scanned once a day for 30 days gets ~30
* chances, so admissibility needs a per-scan tail p below
* `1 - 0.95^(1/30)` = 0.00171, i.e. |z| >= 2.93. 3σ clears that; 2.5σ (p =
* 0.0062, familywise ~17%) does not. A pair set below this is rejected by the
* gate it was issued under, which is the coherence that was missing.
*/
export const PAIRED_CRITERION_SIGMA = 3;
/**
* Ratio of the two sides' NOMINAL tail probabilities: the asymmetry that is
* knowable from the thresholds alone, before any data.
*
* Distinct from {@link NullPowerReport.asymmetryRatio}, which is measured from
* REALIZED firing rates and needs the metric's history. Both matter: nominal
* asymmetry catches a mis-specified pair at write time (and in a unit test);
* realized asymmetry catches a nominally-fair pair on a skewed series. Returns 1
* for a single side (nothing to compare) and Infinity is impossible, since the
* normal tail is never exactly 0 at finite z.
*/
export function nominalAsymmetryRatio(sides: Array<{ sigma: number }>): number {
if (sides.length < 2) return 1;
const ps = sides.map((s) => nominalTailP(s.sigma));
return Math.max(...ps) / Math.min(...ps);
}
/**
* The smallest threshold, in standard deviations, whose chance of at least one
* firing on noise over `scanN` checks does not exceed `ceiling`.
*
* Bisection on [0, 8] for 60 iterations: familywiseP is strictly decreasing in
* sigma for scanN >= 1, so the interval halves onto the boundary. Returns 0 when
* there are no checks (nothing can fire).
*/
export function minAdmissibleSigma(
scanN: number,
ceiling: number = CRITERION_POWER_THRESHOLDS.maxFamilywiseP,
): number {
if (scanN <= 0) return 0;
let lo = 0;
let hi = 8;
for (let i = 0; i < 60; i++) {
const mid = (lo + hi) / 2;
if (familywiseP(mid, scanN) <= ceiling) hi = mid;
else lo = mid;
}
return hi;
}
/**
* Number of checks a criterion gets: checks per day, times days in the window,
* times the share of the window that actually has readings.
*/
export function checksFor(
perDay: number,
days: number,
coverage: number,
): number {
return Math.round(perDay * days * coverage);
}
If you have a criterion, an indicator set or a series on which this arithmetic should be tried, write to cberzins@uottawa.ca.