| Estimand | Numerator | Denominator | Aggregation | Role | Uncertainty |
|---|---|---|---|---|---|
| ratio_of_condition_means | equal-form-weighted condition mean loss for arm A | equal-form-weighted condition mean loss for arm B | divide the two equal-form-weighted condition means | registered display ratio when read from cell_evidence_rows; descriptive elsewhere | hierarchical bootstrap interval is stored only for registered evidence-map contrasts |
| paired_geometric_ratio | replicate-paired arm-A loss within dataset | replicate-paired arm-B loss within dataset | mean replicate log ratio within form; equal-weight forms; exponentiate | exploratory sensitivity used by selected figures | form-level MCSE of the paired log ratio |
Appendix F — Sensitivity of the exploratory layer
The exploratory displays of Chapter 9, Chapter 10 and Chapter 7 rest on choices that the preregistration does not fix: which cell-level ratio to divide, how to summarize a lever, and which combination to call a winner when the margin is small. This appendix reports what those choices are worth. Nothing here is a confirmatory test, and none of it carries a familywise error rate; the purpose is to let a reader see how far a conclusion would move if a different defensible convention had been adopted.
F.1 Two cell-ratio estimands
Two ratios can be formed from the same replicate-paired data, and the book uses both.
Across the 120 cells of the primary contrast the two agree closely. The largest absolute difference on the log scale is 0.040, and 2 cells change side of parity: rasch|50|0.7|skew_pos_strong and rasch|50|0.8|normal. Both flips are Rasch cells at N = 50 whose ratios sit within half a percent of 1, so the disagreement is between two descriptions of a null, not between two findings. Every registered statement in the book uses the evidence-map ratio, which is the first row of Table F.1; the paired geometric ratio appears only where a figure says so.
The crossed-lever count of Section 9.4 is the display most exposed to the choice, and it survives it. Gaussian + GR has the lower loss in 100 of 120 cells under one convention and 100 under the other; the exception rosters differ by two cells. Among the 20 paired-convention exceptions the winning combination is DP + GR in 18 cells and DP + CB in 2.
F.2 Alternative summaries of the two levers
Section 9.1 summarizes each lever by a within-cell range over three levels. Two alternatives give the same ordering. The first replaces ranges with the loss actually obtained by moving one factor away from the operational default, averaged over the non-normal cells on the KS family:
| Quantity | Paired geometric | Ratio of condition means |
|---|---|---|
| Summary swap only (Gaussian + GR) | 0.746 | 0.745 |
| Prior swap only (DP + PM) | 0.879 | 0.884 |
| Both swaps (DP + GR) | 0.593 | 0.595 |
| Product of the two single swaps | 0.656 | 0.658 |
| Log interaction (both minus product) | -0.102 | -0.101 |
The summary swap moves the loss further than the prior swap under either convention, and the two swaps together move it further than their product, which is the compounding reported in Section 9.4. The second alternative is the descriptive factorial decomposition already shown in Section 9.1; its shares agree with the ranges in ordering for every loss family except the rank family, where both levers are negligible.
None of these summaries is a causal decomposition. They describe how far the loss moves within the studied nine-combination roster, and they inherit the goal-matching structure discussed in Section 9.2.
F.3 How stable are the winner maps?
The winner maps of Chapter 7 select, in each cell, the combination with the smallest condition mean over five forms. Because the selection uses the same five forms that estimate the margin, the counts are optimistic by an unknown amount. We bound the instability by resampling forms exhaustively: for each cell we recompute the winner over all 3,125 equally likely resamples of the five observed form positions, exhaustive 5^5 resampling with replacement and record how often the observed winner survives.
| Goal | Target summary | Observed | Expected | Range | Summary held | Prior held |
|---|---|---|---|---|---|---|
| Distribution recovery (KS) | GR | 113 | 112.7 | 104-116 | 0.967 | 0.937 |
| Individual accuracy (MSEL) | PM | 116 | 115.1 | 108-117 | 0.988 | 0.937 |
The choice of summary is far more stable than the choice of prior. Averaged over cells, the winning summary survives resampling in 0.967 (KS) and 0.988 (MSEL) of resamples, while the winning prior survives in about 0.94 and 0.94. The counts themselves move little: the observed 113 GR cells sit against an expected 112.7 with a range of 104 to 116 across resamples. Leaving one form out entirely gives the same picture (0.977 and 0.993 summary stability). The headline regularity of Chapter 7, that the goal selects the summary while the prior varies by regime, is therefore not an artifact of which five item banks were drawn; the prior tiles are the part a reader should treat as provisional, which is what the map’s washed-out cells already say.
F.4 Constrained Bayes: implemented against exact
Appendix C records that production computes the posterior-independence specialization of the constrained Bayes action, using the mean marginal posterior variance rather than the trace of the centered joint posterior covariance. Both are positive affine maps of the posterior-mean vector, so they preserve its ordering and its standardized shape, and they differ only in the rescaling constant.
We can measure the realized constant but not its exact counterpart. The rescaling factor implied by the stored estimates, \(\mathrm{sd}(\hat\theta^{\mathrm{CB}})/\mathrm{sd}(\hat\theta^{\mathrm{PM}})\), runs from about 1.04 in the high-reliability exemplar cells to about 1.49 in the reliability-0.5 bimodal cell, and differs between the two priors by less than 0.03 in every exemplar. Recomputing the exact joint constant would require the full person-by-draw posterior array for all 7,200 fits, which the reporting layer does not carry; that recomputation is recorded as a limitation in Chapter 17 rather than approximated here. What can be said is that the two estimators cannot differ in the direction of any comparison this book reports, because a positive affine map cannot reverse an ordering, and that CB’s role in the results is in any case a middle one: it is never the recommended summary for either goal.
F.5 What this appendix does not do
It does not attach error rates to exploratory quantities, re-test any registered hypothesis, or extend any claim beyond the studied grid. The resampling calculations condition on the five realized form positions and on the frozen response store; they describe stability within this study, not sampling variability of a future study.