14  Whether the report reproduces

The last results chapter asks a question logically prior to every comparison in the previous five: whether the identical analysis, run twice, produces the same report. For two of the three summaries the answer is an unqualified yes. For the third it is no on a predictable class of datasets, and the diagnostics every pipeline already computes do not flag the failure.

14.1 Two summaries reproduce; one has a threshold

Across all 78 replicated fits, the PM and CB summaries agree between seeds to within 0.028 reference SD for the median person, including the four fits whose convergence diagnostics were worst (maximum \hat R up to 1.28). The GR summary exceeds a 0.10 SD median seed-to-seed movement on 11 of 78 fits, with a worst case of 0.163 SD and a worst between-seed correlation of 0.938. All 11 failures are Rasch fits of 12 to 16 items: 11 of 24 such fits fail, against 0 of 54 fits with 17 items or more. The failure is a threshold in test length, not a trend.

Seed-to-seed movement by summary against item count, with GR failures confined to 12-16-item Rasch forms, and against R-hat, showing no relation.
Figure 14.1: The triple-goal summary fails to reproduce on short forms, and convergence diagnostics do not detect it: the failure is a property of the test, not of the chains. Top row: for each of the 78 replicated fits, the median absolute difference between the seed-A and seed-B estimates of the same person (reference-SD units), by summary; the red line is the 0.10 SD reproducibility threshold and the shaded band marks 12-16-item forms. PM and CB reproduce within 0.028 SD on every fit; GR exceeds the threshold on 11 fits, all of them Rasch forms of 12-16 items (11 of 24 such fits, against 0 of 54 longer ones). Bottom: the same GR movements against the fit’s worst convergence R-hat; 10 of the 11 failures sit left of the 1.1 screen (which catches 1), while form length separates them cleanly (r = -0.66 within Rasch); 10 failures even have maximum \hat R<1.05. The mechanism audit finds dense near-rank bands rather than stochastic resolution of exact ties: only 2 failing production rows contain an exact posterior-rank tie, while the sorted GR ensembles move just 0.0009–0.0054 SD even as 328–493 people change rank. Small independent- refit perturbations therefore reassign an almost fixed set of GR quantiles across nearby people. Movements re-derived from the frozen seed-A and seed-B estimate stores; diagnostics from the frozen fit records.

The mechanism audit rules out the tie story. Across all production fits, PM has 0 exact ties; only 4 of 78 posterior-rank rows have any exact tie, and only 2 failing rows do. In the failures, the sorted GR ensembles move only 0.001–0.005 reference SD even while aligned persons move above 0.10. Posterior-rank correlations remain 0.99986– 0.99995, but 328–493 persons change rank, with median movement 10–19 positions. The supported explanation is dense near-rank bands: small independent-refit perturbations reassign a nearly fixed set of GR quantiles across nearby people. Random resolution of exact ties is neither necessary nor common.

Nothing in that mechanism requires non-convergence. 10 of the 11 failing fits have maximum \hat R<1.05, while a screen at \hat R>1.1 catches 1 of 11. Item count separates failures from successes in this portfolio (correlation -0.66 within Rasch). Reproducibility of the reported quantity is its own property, to be measured, not inferred from chain health.

14.2 The practical rule

The finding changes practice more concretely than any recommendation in this book, because it survives every branch of the decision tree: on a form of 12 to 16 items, do not report GR from a single run. Either run the summary at two seeds and require agreement (the check costs one refit and is the only known screen), or use CB, which repairs ensemble spread, reproduces everywhere here, and concedes shape repair that a 14-item form was never going to deliver anyway. Above about 17 items the portfolio gives no reason for the extra caution (0 failures in 54 fits). The stored v1 headline used a different per-fit denominator and counted 9. This edition uses the book’s common Gaussian-PM cell reference SD throughout and counts 11; the threshold’s location, between 16 and 17 items in this portfolio, should be read as this portfolio’s gap, not a universal constant.