9  The consequence map

Part IV reports what actually changes on the thirteen tests when the two choices change. This chapter gives the whole answer at once, in one display, before the family-by-family chapters take it apart, because the shape of the whole is itself the book’s organizing result: whether a choice “matters” is a property of the output family being reported and of the lever being pulled, before it is a property of the dataset.

All shape-conditioned rows in Part IV group the case-cells by the standardized shape class of Section 5.4. The grouping is descriptive: it says where the portfolio’s departures from normality were measured, not which report is closer to truth.

9.1 Both levers, all cases, all families

Figure 9.1 shows every case-cell (rows), every output family (columns), and both levers (panels), with each change expressed as a multiple of its family’s pre-declared materiality threshold.

Two heatmap panels, prior swap and summary swap, of 26 case-cells by four output families, colored by change relative to the materiality threshold.
Figure 9.1: Whether the analyst’s choice matters depends on what is being reported and on which lever is pulled, before it depends on the dataset. Each row is one of the 26 case-cells (labeled case and item model, grouped by standardized shape class and ordered by reliability within class, with the one case-cell reported non-normal without a class in a group of its own). Columns are the four output families of a score report; the left panel changes the prior with the summary held at PM, the right panel changes the summary with the prior held Gaussian. That contrast uses the same stored posterior draws, so it is paired and has reduced, not zero, Monte Carlo variation. The fill shows the size of the resulting change as a multiple of that family’s pre-declared materiality threshold (log scale: red above the bar, blue below; scores, median move over 0.10 SD; distribution, KS over 0.05 or spread ratio outside [0.95, 1.05]; rankings, more than 5% of ranks moving over 2 percentile points or top-decile Jaccard under 0.90; tails, any fixed cut moving over 2 points). Black dots mark the cells whose current materiality flag is set (the flags consider each family’s full rule, not only the statistic plotted). The color scale saturates at 1/8 and 8: 18 tiles lie below and 6 above those endpoints, but black-dot flags use the unsaturated ratios and each family’s full rule. Values from the frozen v1 comparison rowset; 26 case-cells, thresholds fixed before any model was fitted.

Reading the left panel down its columns gives the prior swap’s geography. On individual scores the red is confined to the skewed rows; on the reported distribution it extends through the skewed and bimodal rows and touches the intended normal-control rows only lightly; on rankings it is strongest exactly where scores are quiet (the bimodal rows, a dissociation Chapter 12 explains); on tails most cells are blue. The localized red cells, and especially the two largest exceptions, are traced in Chapter 13 to mass near a cut and dense bands rather than to a general tail effect. The materiality counts are tabulated in Table 9.1: of the seven bimodal cells, one is material on scores and seven on the distribution; of the nine skewed cells, five and nine.

Table 9.1: Material case-cells per output family under the prior swap (DP-focused vs Gaussian, PM held), among the 21 case-cells whose standardized shape class is one of the three the simulation grid contains.
Shape class Case-cells Scores Distribution Rankings Tails
bimodal 7 1 7 6 3
normal 5 0 2 3 1
skewed 9 5 9 7 3
Two compact heatmaps count material classified case-cells by output family and lever.
Figure 9.2: No single answer to ‘does the choice matter’ survives a change of output family: each lever has its own geography over the same case-cells. Tiles count the case-cells whose change clears the pre-declared materiality threshold (printed as count over cells; fill is the share), by standardized shape class (rows) and output family (columns) over the 21 case-cells whose class has a counterpart among the simulation’s generating conditions. The left panel is the prior swap (DP-focused against Gaussian, summary held at PM; score, distribution and tail flags come from the frozen rowset, while ranking membership is re-derived with the exact top-k rule). The right panel is the summary swap (GR against PM under the Gaussian prior), the contrast the v1 book specified and never reported. The prior’s materiality concentrates in the skewed rows for scores and in the skewed and bimodal rows for the distribution; the summary swap is material for the distribution family almost everywhere, including the normal-control rows. The five remaining case-cells are excluded, the four the screen leaves undetermined and the one it reports non-normal without a class, because neither status identifies a generating condition to group them by. Every count is a consequence measurement and carries no direction: a material change is a change, not an improvement. Flags from the frozen comparison rowset; classes from the standardized shape manifest.

The right panel is the contrast the first edition of this book specified, froze two claims about, and then never displayed: the summary swap. Its geography is different in both directions. It reaches the distribution family almost everywhere, including the normal controls, because shrinkage repair is not a non-normality phenomenon; and it leaves rankings largely blue, because an affine rescaling (CB) cannot reorder anyone and GR’s reordering is concentrated in dense near-rank bands. The two panels together are the simulation’s two-lever anatomy (Chapter 2), photographed on real data.

9.2 The two levers compared, cell by cell

Figure 9.3 sharpens the comparison to one point per case-cell: the size of the summary swap against the size of the prior swap, for scores and for the distribution. It also forces a distinction the rest of the book depends on, between displacement (how far a swap moves the report) and improvement (how much closer the report gets to the truth), of which a case can measure only the first.

Scatter of summary-swap size against prior-swap size per case-cell, for scores and for the KS distance.
Figure 9.3: Which lever displaces the report more depends on output family and on shape class: the summary swap moves scores more on every normal-consistent and undetermined row, while the prior swap moves the reported curve more on every non-normal cell. Each point is one of the 26 case-cells, placed by the size of the prior swap (horizontal) against the size of the summary swap (vertical) for the same quantity; points above the dashed identity line are cells where changing the summary displaces the output more. Left, median individual-score movement (SD units): all 9 of 9 normal-and-undetermined cells sit above the line, and seven of nine skewed cells below it. Right, the KS distance between the two reported curves: the prior swap equals or exceeds the summary swap on all 16 cells carrying a non-normal class, because at this portfolio’s reliabilities (point area) the summary’s displacement, which is shrinkage repair, is modest (its size tracks reliability, r = -0.59, from 0.12 SD below .70 to 0.04 above .90, the frozen C-08 claim), while the prior contributes the shape change the summary cannot. Displacement is not accuracy: the matched sim-v3 verdicts give different lever rankings under different estimands (chapter 15). Both summaries use the same posterior draws, so the vertical coordinate is paired and has reduced, but not zero, Monte Carlo variation. Classes from the standardized shape manifest; displacements from the frozen comparison rowset.

For individual scores the levers divide the portfolio cleanly: on every normal-consistent and undetermined cell the summary swap displaces scores more than the prior swap, and on seven of the nine skewed cells the prior swap displaces them more, which is where the simulation locates the flexible prior’s individual-level action. For the reported curve the ordering reverses: the prior swap produces the larger between-report distance on all sixteen skewed and bimodal cells. Both levers clear the distribution family’s materiality bar on most of the portfolio (the summary swap on 23 of 26 cells, largely through its spread repair), so the reversal is about which material change is larger, not about either being quiet.

The reversal is not a contradiction of the simulation’s larger-lever result; it is that result’s own moderator doing its work. The frozen claims the first edition never presented say how: the summary swap’s size tracks reliability (correlation -0.59 across cells, median movement 0.118 SD below reliability .70 falling to 0.044 above .90; claim C-08), and this portfolio lives at the high-reliability end, where shrinkage, the thing the summary repairs, is smallest. At fixed reliability the summary swap is largest on the normal-consistent cells (0.124 SD against 0.081 skewed and 0.069 bimodal in the .65–.80 band; claim C-09, a blueprint reversal), with a clean mechanism: on a non-normal case the flexible prior supplies part of the spread-and-shape repair on its own, so the summary has less left to do; on a normal case the summary is the only repair there is. The improvement direction cannot be read from these axes. The exploratory sim-v3 join in Chapter 15 compares truth-based losses, but it does not support a universal larger-share claim for the summary on these rows: the answer changes with the factorial, condition-mean, or paired-geometric estimand. How much movement is toward truth is the simulation’s question, not the case data’s.

9.3 The menu against the default

The last display answers the practitioner’s literal question, K1: not “prior versus summary” but “any of the eight alternatives versus what my software already does”.

Dot plot of the eight alternative combinations against the default, by shape class, for score movement and top-decile churn.
Figure 9.4: What moves a report away from its default differs by shape class: under skew the prior swap moves scores as much as a summary swap, while on normal-consistent rows the summaries dwarf the priors, and the top decile is disturbed by everything except a pure prior swap on those same normal-consistent rows. Each panel row is a shape class, each panel column an output; within a panel, the eight alternative prior + summary combinations are compared with the default report (Gaussian + PM), showing the median over that class’s case-cells of the median score movement (left, SD units) and of the share of the reported top decile that changes membership (right). Color gives the prior, glyph the summary; dashed lines mark the materiality conventions (0.10 SD; the 5.3% churn corresponding to a top-decile Jaccard of 0.90). This is the K1 contrast the v1 book specified and never reported: the practitioner’s own question, what changes if I abandon the default. Each point is a median over five to nine case-cells, so the panels are read as orderings rather than as calibrated effect sizes, and a difference from the default is a displacement, not an improvement. The three classes hold the same 21 case-cells under the standardized screen as under the frozen one. Score movement comes from the frozen K1 medians table; top-decile membership is re-derived using exact k=\lceil0.10n\rceil sets with person-ID boundary tie-breaking.

Three readings of Figure 9.4 matter for what follows. On the normal-consistent cases the summary rows dwarf the prior rows: a reader who moves from Gaussian + PM to Gaussian + CB has changed their reported scores by a median 0.14 SD while a pure prior swap changes them by 0.04. Under skew the prior swap joins the summaries in materiality, and the combined swap (DP + GR) is not the sum of its parts. And the right column shows that almost any departure from the default disturbs a top decile by an amount comparable to the paired seed comparison, a warning Chapter 12 makes precise. The family chapters now take the columns of Figure 9.1 one at a time.