12  Rankings and selection

The third family is the ordering: who ranks above whom, and who makes a selected group. It is the family where aggregate statistics mislead most reliably. It is also where the independent A/B refit gives an essential reference for interpreting method contrasts.

12.1 The correlation that conceals everything

Every between-prior rank correlation in the portfolio exceeds 0.996. Reported alone, that number would close the family: the ordering is preserved, nothing to see. It conceals two facts. First, small rank movements are the norm, not the exception: under skew, 50% of respondents move more than one percentile point. Second, selection operates at cuts, and cuts aggregate small movements into membership changes: on the bimodal cells the prior swap changes about 10% of a reported top decile (top-decile Jaccard 0.818), about 12% of a top-5% group, and under 5% of a top quartile: the tighter the cut, the larger the churn. So the bimodal cells, quiet in scores, are loud in selection, the second half of the dissociation Chapter 9 mapped.

12.2 The paired seed reference

Before attributing churn to the prior, the design measures the churn in one paired refit of the identical method. Re-running it with a new seed changes 2.0% of a top decile on the skewed cells, 4.0% on the normal-consistent rows, 6.2% on the bimodal cells, and 2.1% on the undetermined ones, a stratum the first edition’s headline range omitted. In the worst single cell the same method agreed with itself about only half its selections. Across the three PM prior arms, the worst bimodal overlap is 0.556, or 29% churn. Across all nine prior-summary combinations the exact top-k maximum is 32.0% (minimum Jaccard 0.515). The Gaussian-PM subset in Figure 12.1 has a separate maximum of 28.0%. The figure puts every case-cell’s prior churn against its own seed churn.

Prior-swap churn against same-method seed churn per case-cell, with several cells below the identity line.
Figure 12.1: A reported top decile is partly arbitrary before any method question is asked: re-running the identical method at a new seed changes 0 to 28% of it, and for several case-cells the prior swap churns no more than that paired seed comparison. Each point is one case-cell, placed by the share of its reported top decile (Gaussian + PM) that changes membership when only the MCMC seed changes (horizontal) and when only the prior changes at a fixed seed (vertical); the dashed line marks equality, and points in the shaded region churn less by changing the prior than by re-running the sampler. Re-derived class values: seed churn 2.0% (skewed), 4.0% (normal), 6.2% (bimodal) and 2.1% (undetermined, a group the v1 headline range omitted); the one case-cell reported non-normal without a class is plotted but not summarized here. Churn is 100(1 − 2J/(1+J)) for top-decile Jaccard J. Per-cell and class values are re-derived from exact top-k membership sets, k=\lceil0.10n\rceil, with person ID breaking a boundary tie. One A/B pair per cell does not estimate the full seed-to-seed uncertainty distribution.

The reading is not that selection is meaningless; overlap remains far above the chance benchmark (raw agreement near 98% against a chance level near 82%, with chance-corrected agreement lowest on the bimodal cells at 0.889). The reading is that a reported top decile carries seed-dependent variation before any methodological question is asked, and that method effects should be compared with that paired refit. Because there is only one A/B pair per cell, it does not estimate a seed-to-seed uncertainty interval or establish a universal lower bound. Figure 12.2 applies the standard to the three available perturbations.

Churn from one paired seed comparison, elicitation swap, and prior-family swap by shape class.
Figure 12.2: Judged against the paired seed comparison (grey), only the prior-family swap on bimodal cases shows clearly more churn, while the elicitation setting changes more memberships than the prior family on the skewed cases. Bars give the class-level share of the reported top decile that changes membership under each of three perturbations: re-running the identical method with one new seed (grey; one paired comparison), swapping the two DP elicitation settings, and swapping prior family (orange). On bimodal cases the values are seed comparison 6.2%, elicitation 6.2% and prior 10.0%; on skewed cases the elicitation churn (6.0%) exceeds the prior churn (4.0%). Every bar is re-derived from exact top-k membership sets, with k=\lceil0.10n\rceil and person ID breaking a boundary tie; PM summary, standardized shape classes over the 21 case-cells that carry one. The grey bar is one paired replicate, not an uncertainty interval or a universal lower bound.

Only the prior-family swap on bimodal cells (10.0% against a 6.2% paired seed comparison) clearly exceeds its seed comparison. The elicitation contrast, focused against broad, is 6.2% on the bimodal cells and exceeds the prior-family contrast on the skewed ones (6.0% against 4.0%), an pattern the first edition displayed without remark. Where the prior family barely reorders anyone, the two near-equivalent elicitation settings still produce measurable membership differences. This is also where the elicitation question, deferred from Chapter 10, gets its answer: on scores, distributions and tails the two elicitations are near-identical (median score movement under 0.01 SD; the frozen side-by-side table puts the elicitation’s materiality at 4 of 78 fits on scores against 22 for the prior family), and their one visible effect is top-group membership churn of the same order as the single paired seed comparison. More seeds would be required to separate those sources by an uncertainty interval.

12.3 Clumps

The geometry that makes selection fragile on short forms is visible in the estimates themselves. The stored posterior means are numerically distinct, but short Rasch forms produce dense near-rank bands associated with limited raw-score information. Using the preregistered 0.02-gap definition, the default report’s scale is a set of bands: on the 12-to-16-item cells the median count is 14 clumps with the largest holding 16.2% of the sample. Figure 12.3 shows three scales side by side. A band does not imply exact ties: it marks many nearly indistinguishable estimates packed into a short interval. Small refit perturbations can reorder people inside the band, and a cut through it can move many memberships together. The consequences for cut-based reporting are the next chapter’s subject, and the display belongs here because selection inherits the same geometry: a top-decile boundary that falls inside a dense band is a boundary drawn through distinctions the fitted report resolves only weakly.

Dot-strip of default-report estimates for three cases, showing clumped short Rasch scales against a continuous long 2PL scale, with fixed cuts overlaid.
Figure 12.3: Short Rasch forms produce dense 0.02-gap bands, so a fixed cut can be locally fragile even though the estimates are numerically distinct. Each panel stacks the default-report estimates of one permissively licensed case as dots; red dashed lines are the fixed cuts at θ = 1.0, 1.5 and 2.0 used by the tail family. The 16-item C4 is partitioned into 18 such bands; the largest holds 19% of the sample (95/500) even though all 95 estimates are distinct, while the 56-item 2PL C3 is effectively continuous at this resolution. C6 supplies a second short-form comparison without exposing a case whose person rows are withheld. C4’s 19.0-point cut movement matches its largest band share and is retained as exploratory observation E-01, not as a general arithmetic law. Estimates from the permissive-case frozen run store; band = a run with no gap above 0.02.

12.4 What this family establishes

Rank correlations near one coexist with selection churn near ten percent, so the correlation is the wrong summary for any decision that operates at a cut. The exact top-k seed audit changes 2.0– 6.2% of a class-median top decile, which supplies a directly measured comparison for method-induced churn. At the class median, only the prior-family swap on the former bimodal rows clearly exceeds the corresponding seed value. The elicitation swap is similar to the seed comparison in that group and somewhat larger in the skewed rows, so it should not be interpreted as a stable selection effect. For a program publishing a selection list, the relevant protocol is to duplicate the intended report and require membership agreement; for GR on a 12–16-item form, Chapter 14 makes that duplicate run mandatory in this portfolio-derived rule.