Appendix F — Sensitivity of the exploratory layer

The exploratory displays of Chapter 9, Chapter 10 and Chapter 7 rest on choices that the preregistration does not fix: which cell-level ratio to divide, how to summarize a lever, and which combination to call a winner when the margin is small. This appendix reports what those choices are worth. Nothing here is a confirmatory test, and none of it carries a familywise error rate; the purpose is to let a reader see how far a conclusion would move if a different defensible convention had been adopted.

F.1 Two cell-ratio estimands

Two ratios can be formed from the same replicate-paired data, and the book uses both.

Table F.1: The two cell-ratio estimands. Both are computed from the same frozen replicate losses; they differ in the order of division and averaging. Source: data/derived/ratio_estimand_metadata.rds.
Estimand Numerator Denominator Aggregation Role Uncertainty
ratio_of_condition_means equal-form-weighted condition mean loss for arm A equal-form-weighted condition mean loss for arm B divide the two equal-form-weighted condition means registered display ratio when read from cell_evidence_rows; descriptive elsewhere hierarchical bootstrap interval is stored only for registered evidence-map contrasts
paired_geometric_ratio replicate-paired arm-A loss within dataset replicate-paired arm-B loss within dataset mean replicate log ratio within form; equal-weight forms; exponentiate exploratory sensitivity used by selected figures form-level MCSE of the paired log ratio

Across the 120 cells of the primary contrast the two agree closely. The largest absolute difference on the log scale is 0.040, and 2 cells change side of parity: rasch|50|0.7|skew_pos_strong and rasch|50|0.8|normal. Both flips are Rasch cells at N = 50 whose ratios sit within half a percent of 1, so the disagreement is between two descriptions of a null, not between two findings. Every registered statement in the book uses the evidence-map ratio, which is the first row of Table F.1; the paired geometric ratio appears only where a figure says so.

The crossed-lever count of Section 9.4 is the display most exposed to the choice, and it survives it. Gaussian + GR has the lower loss in 100 of 120 cells under one convention and 100 under the other; the exception rosters differ by two cells. Among the 20 paired-convention exceptions the winning combination is DP + GR in 18 cells and DP + CB in 2.

F.2 Alternative summaries of the two levers

Section 9.1 summarizes each lever by a within-cell range over three levels. Two alternatives give the same ordering. The first replaces ranges with the loss actually obtained by moving one factor away from the operational default, averaged over the non-normal cells on the KS family:

Table F.2: Loss ratios against the Gaussian-plus-PM default on the KS family, non-normal cells, under both ratio conventions. The independent product is what the two single swaps would deliver if they did not interact. Source: data/derived/book-facts.rds (F$crossed_levers).
Quantity Paired geometric Ratio of condition means
Summary swap only (Gaussian + GR) 0.746 0.745
Prior swap only (DP + PM) 0.879 0.884
Both swaps (DP + GR) 0.593 0.595
Product of the two single swaps 0.656 0.658
Log interaction (both minus product) -0.102 -0.101

The summary swap moves the loss further than the prior swap under either convention, and the two swaps together move it further than their product, which is the compounding reported in Section 9.4. The second alternative is the descriptive factorial decomposition already shown in Section 9.1; its shares agree with the ranges in ordering for every loss family except the rank family, where both levers are negligible.

None of these summaries is a causal decomposition. They describe how far the loss moves within the studied nine-combination roster, and they inherit the goal-matching structure discussed in Section 9.2.

F.3 How stable are the winner maps?

The winner maps of Chapter 7 select, in each cell, the combination with the smallest condition mean over five forms. Because the selection uses the same five forms that estimate the margin, the counts are optimistic by an unknown amount. We bound the instability by resampling forms exhaustively: for each cell we recompute the winner over all 3,125 equally likely resamples of the five observed form positions, exhaustive 5^5 resampling with replacement and record how often the observed winner survives.

Table F.3: Exhaustive form-resampling stability of the winner maps. The target summary is GR for the distributional goal and PM for the individual goal. Counts are the number of cells (of 120) whose winner uses the target summary. Source: data/derived/winner_resampling.rds.
Goal Target summary Observed Expected Range Summary held Prior held
Distribution recovery (KS) GR 113 112.7 104-116 0.967 0.937
Individual accuracy (MSEL) PM 116 115.1 108-117 0.988 0.937

The choice of summary is far more stable than the choice of prior. Averaged over cells, the winning summary survives resampling in 0.967 (KS) and 0.988 (MSEL) of resamples, while the winning prior survives in about 0.94 and 0.94. The counts themselves move little: the observed 113 GR cells sit against an expected 112.7 with a range of 104 to 116 across resamples. Leaving one form out entirely gives the same picture (0.977 and 0.993 summary stability). The headline regularity of Chapter 7, that the goal selects the summary while the prior varies by regime, is therefore not an artifact of which five item banks were drawn; the prior tiles are the part a reader should treat as provisional, which is what the map’s washed-out cells already say.

F.4 Constrained Bayes: implemented against exact

Appendix C records that production computes the posterior-independence specialization of the constrained Bayes action, using the mean marginal posterior variance rather than the trace of the centered joint posterior covariance. Both are positive affine maps of the posterior-mean vector, so they preserve its ordering and its standardized shape, and they differ only in the rescaling constant.

We can measure the realized constant but not its exact counterpart. The rescaling factor implied by the stored estimates, \(\mathrm{sd}(\hat\theta^{\mathrm{CB}})/\mathrm{sd}(\hat\theta^{\mathrm{PM}})\), runs from about 1.04 in the high-reliability exemplar cells to about 1.49 in the reliability-0.5 bimodal cell, and differs between the two priors by less than 0.03 in every exemplar. Recomputing the exact joint constant would require the full person-by-draw posterior array for all 7,200 fits, which the reporting layer does not carry; that recomputation is recorded as a limitation in Chapter 17 rather than approximated here. What can be said is that the two estimators cannot differ in the direction of any comparison this book reports, because a positive affine map cannot reverse an ordering, and that CB’s role in the results is in any case a middle one: it is never the recommended summary for either goal.

F.5 What this appendix does not do

It does not attach error rates to exploratory quantities, re-test any registered hypothesis, or extend any claim beyond the studied grid. The resampling calculations condition on the five realized form positions and on the frozen response store; they describe stability within this study, not sampling variability of a future study.