15  The preregistered record

Earlier chapters reported registered results where their content belonged. This chapter collects the complete confirmatory record in the form specified by the preregistration, together with the registration history itself. It consolidates source-backed results and timing rather than introducing a new analysis. The record establishes which results retain confirmatory status; it does not require later descriptive or exploratory synthesis to imitate the confirmatory presentation.

15.1 What was locked, and when

The analysis plan was locked on 2026-07-24, before the production wave ran, and its SHA-256 fingerprint travels in every confirmatory output row. The locked content includes the replicate-level paired log loss-ratio estimand (2,400 rows, 20 per condition), the linear mixed model with its fixed structure and form-family random intercept, measured centered reliability as the moderator, one Holm family over exactly five p-values at \(\alpha = 0.05\), the non-inferiority margin log(1.05) for H4, the opportunity region (non-normal shapes at N of 200 and 500), the tie band [0.95, 1.05] and safety boundaries 1.05 and 1.10, the inclusion policy (every completed fit enters; diagnostics gate wording, not membership), and the refusal rule under which a result that cannot be computed as registered is reported as such rather than approximated. The registration history after lock consists of the records in Table 15.1. The timing column distinguishes three states that should not be collapsed: a pilot outcome may already exist, fitting may already be under way, and production analysis may still not exist.

Table 15.1: The registration record: the locked plan and its amendments, with the timing of each relative to pilots, production fitting, and outcomes. Source: tables/T-addenda.rds.
Record Timing Content
Preregistration (locked 2026-07-24) before the production wave estimands, H1-H5, one Holm family, inclusion policy, thresholds
Addendum 001 2026-07-26: after P-2; before P-3 and production reference grid for H1: which reading of 'at mean logN' the contrast uses
Addendum 002 2026-07-27: after P-2; before P-3 and production symmetric floor on both operands of registered replicate-level log ratios; non-positive losses are floored, not dropped; condition-level ratios are outside its scope
Addendum 003 (+ amendment 001) 2026-07-29/30: after P-3; before any production fit separate H6-H7 Holm pair; amendment: H7 uses a share-only floor and form_family clustering (form_rep sensitivity)
Addendum 004 2026-07-31: 1,798 fits complete; no analyzer or outcome optional direction-blind replication extension; the locked wording ladder was not added here (trigger did not fire)
Addendum 005 2026-08-03: post-fit, before production analysis common reporting scale repair plus restriction of H6 to the realized reliability/item-count/information ladder
Addendum 006 2026-08-06: after production outcomes were known post-outcome reporting policy for H4 safety: full 7-cell anchor in the confirmatory record; concise warnings and cross-references permitted in later synthesis
Source: preregistration and Addenda 001-006 registry. Dates, the 1,798-fit state at Addendum 004, and amendment timing are historical registry facts, not analytical outcomes recomputed from data/tidy.

Read chronologically, Addenda 001 and 002 follow P-2 but precede P-3 and production. Addendum 003 follows P-3, discloses that the pilot prompted H6 and H7, and precedes every production fit; its Amendment 001 specifies H7’s share-only floor and form_family clustering. Addendum 004 was logged with 1,798 production fits complete, but before a production analyzer or outcome existed. It added an optional direction-blind replication extension, not the wording ladder already present in the lock. Addendum 005 is post-fit and pre-analysis: it repairs the reporting scale and restricts H6 to the realized reliability/item-count/information ladder. Addendum 006 alone is post-outcome.

The Addendum 006 policy is applied proportionately in this book. The full seven-cell table, ratios and intervals appear with the confirmatory H4 record here and with the main H4 result in Section 12.1. Later summaries may state the bimodal, low-reliability exception concisely and cross-reference one of those anchors. They need not reproduce seven rows every time H4 is mentioned. That policy keeps the local harm visible without turning a post-outcome reporting rule into a constraint on unrelated distributional claims or free exploratory synthesis.

15.2 The primary family

Table 15.2: The five primary hypotheses under the locked specification, with the calibration control reported alongside as a gate. Source: tables/T-confirmatory.rds.
Hypothesis Estimate SE 95% CI p (Holm) Decision
H1 flexible prior improves distribution recovery (non-normal) -0.221 0.013 [-0.246, -0.195] 1.35e-28 supported
H2 the advantage grows with sample size -0.114 0.004 [-0.123, -0.106] 1.24e-131 supported
H3 model-by-shape dissociation -0.127 0.023 [-0.175, -0.078] 0.0001 supported
H4 no individual-accuracy harm (opportunity region)1 -0.070 0.005 [-0.080, -0.060] 2.29e-23 supported
H5 focused beats broad -0.011 0.003 [-0.017, -0.005] 0.0001 supported
Calibration control (normal cells; gate, not hypothesis) 0.016 0.007 [0.002, 0.030] - pass (0/40 wins)
1 H4 is a one-sided non-inferiority test at margin 0.0488; its one-sided upper 95% bound is -0.0614.
Source: primary rows of confirmatory_model_rows.tsv. Log loss-ratio scale; negative favors the flexible pipeline. One Holm family over five p-values, alpha = 0.05. Satterthwaite df; 2,400 frozen replicate rows (H4: 800; H1: 1,600; H5: 1,600).

All five primary hypotheses are supported under the locked specification, with the smallest Holm-adjusted p-value at \(1.2 \times 10^{-131}\) (H2) and the largest at \(0.000\) (tied for H3 and H5 by Holm’s step ordering). The calibration control passed with 0 wins among 40 normal cells. Effect sizes on the ratio scale: H1 corresponds to a loss ratio of .80 at the design center; H2 to a further multiplication by .89 per doubling of N; H4’s point estimate corresponds to a 7 percent individual-accuracy improvement over the opportunity region, with the seven flagged cells of Table 15.3 qualifying its scope; H5 to a 1 percent focused-over-broad edge whose magnitude is specification-dependent (Section 14.5).

Table 15.3: The seven flagged opportunity-region cells that qualify the registered H4 mean non-inferiority result. Addendum 006 is post-outcome. Source: tables/T-safety-cells.rds.
Condition MSEL ratio 95% CI Label
2PL, N = 500, reliability 0.5, bimodal 1.138 [1.115, 1.162] block
2PL, N = 200, reliability 0.5, bimodal 1.122 [1.095, 1.149] block
Rasch, N = 200, reliability 0.5, bimodal 1.094 [1.048, 1.136] caution
2PL, N = 200, reliability 0.6, bimodal 1.092 [1.070, 1.116] caution
Rasch, N = 500, reliability 0.5, bimodal 1.076 [1.059, 1.093] caution
2PL, N = 500, reliability 0.6, bimodal 1.072 [1.059, 1.084] caution
Rasch, N = 200, reliability 0.6, bimodal 1.052 [1.030, 1.077] caution
Source: generated F$safety$fired rows. The 7 opportunity-region cells whose DP-versus-Gaussian MSEL ratio exceeds the 1.05 pass boundary; every interval excludes 1. Addendum 006 is a post-outcome reporting policy: the full block anchors the confirmatory H4 record, while later synthesis may use a concise warning and cross-reference.

15.3 The secondary family

Table 15.4: The secondary family registered by Addendum 003, with its within-normal diagnostic. Source: tables/T-secondary.rds.
Hypothesis Estimate SE 95% CI p (Holm, secondary) Decision
H6 reliability slope of the log ratio (non-normal, equal shape weight) -0.514 0.037 [-0.587, -0.441] 4.38e-41 supported
H7 summary-by-goal dissociation 0.281 0.031 [0.210, 0.351] 8.57e-06 supported
H6 diagnostic: the same slope within normal cells -0.029 0.052 [-0.131, 0.074] 0.5831 (raw) not evaluated (diagnostic)
Registered by Addendum 003 after the P-3 pilot and before any production fit; own Holm family of exactly two, no alpha shared with the primary family. Amendment 001 specifies H7's share-only floor and form_family clustering.

Both secondary hypotheses are supported in their own two-member Holm family. H6’s slope, -0.514 per unit of achieved reliability, describes moderation along the realized ladder, where reliability, item count and finite-test information move together; it is not a reliability-only effect. Its decomposition by shape in Section 10.2 is a post-registration exploratory analysis, and its normal-cell diagnostic is flat as required. H7’s dissociation, +0.281 log units, is registered; the broader lever analysis of Chapter 9 is a post-registration synthesis.

15.4 Every estimate under every registered specification

Table 15.5: All seven registered hypotheses under the locked primary specification and the three preregistered sensitivity variants. Source: tables/T-sensitivity.rds.
Hypothesis Primary (locked) Equal condition weight Strict-pass only CR2 cluster-robust
H1 -0.221 [-0.246, -0.195] -0.221 [-0.246, -0.195] -0.232 [-0.249, -0.215] -0.221 [-0.285, -0.156]
H2 -0.114 [-0.123, -0.106] -0.114 [-0.123, -0.106] -0.116 [-0.124, -0.108] -0.114 [-0.123, -0.106]
H3 -0.127 [-0.175, -0.078] -0.127 [-0.175, -0.078] -0.121 [-0.174, -0.069] -0.127 [-0.158, -0.095]
H4 -0.070 [-0.080, -0.060] -0.070 [-0.080, -0.060] -0.069 [-0.078, -0.060] -0.070 [-0.081, -0.059]
H5 -0.011 [-0.017, -0.005] -0.011 [-0.017, -0.005] -0.003 [-0.006, 0.000] -0.011 [-0.015, -0.007]
H6 -0.514 [-0.587, -0.441] -0.514 [-0.587, -0.441] -0.486 [-0.558, -0.414] -0.514 [-0.604, -0.424]
H7 0.281 [0.210, 0.351] 0.281 [0.210, 0.351] 0.286 [0.218, 0.353] 0.281 [0.242, 0.320]
Entries are estimates with 95% intervals from confirmatory_model_rows.tsv and secondary_confirmatory_rows.tsv on the registered log scale. CR2 intervals use 3.7-4.0 degrees of freedom and issue no family verdict by convention.

The grid contains 28 estimates; its reading is given in Section 14.5. For completeness of the confirmatory record, all six adequacy checks returned claim_wording_step = full, so no registered confirmatory statement required a wording downgrade. This status does not assign a wording tier to later exploratory claims.

15.5 Interpretation boundaries set by the registration

The registered record supports the following confirmatory claim classes. H1, H2, H3 and H5 support their stated distributional-recovery claims within their registered domains. H6 supports moderation along the realized reliability/item-count/information ladder, not a reliability-only causal effect. H4 supports mean non-inferiority over the opportunity region, qualified by the seven bimodal low-reliability cells displayed here and in Section 12.1. H7 supports the registered summary-by-goal dissociation.

The book is not otherwise confined to those claim classes. The evidence map, quantile and tail families, lever decompositions, shape-specific H6 slopes and other post-registration analyses may be used for descriptive explanation and exploratory synthesis. Their boundary is provenance: they do not acquire Holm protection, alter a registered verdict or license unqualified extrapolation beyond the simulated grid. Exploratory status is marked at the entry to a section or where a statement could reasonably be mistaken for a confirmatory one, rather than repeated at every use. Tail ratios remain especially fragile because of their zeros and floor incidence, and the evidence map’s harm-flag label remains a residual category rather than a harm verdict.