2  The two levers

A Bayesian item response analysis ends with a joint posterior over abilities. Everything a report contains, scores, orderings, shares below a cutoff, is a functional of that posterior, and the analyst controls two separable choices on the way to it: the prior placed on the ability distribution, and the summary that turns each person’s posterior into a reported number. This chapter develops the two choices far enough that the results can be read with their mechanics in view. Formal statements and derivations are collected in Appendix C; the theory volume of this project treats them fully.

2.1 Shrinkage, and what it does to an ensemble

Write \(\theta_i\) for person \(i\)’s ability and \(G\) for the population distribution the study wants recovered. Under the hierarchical model the posterior mean \(\hat\theta_i^{\mathrm{PM}} = E(\theta_i \mid \mathbf{y})\) minimizes expected squared error for each person, and it does so by shrinking: each estimate is a compromise between the person’s data and the prior, weighted by their precisions. Shrinkage is not a defect; it is the mechanism by which the posterior mean earns its optimality. The defect appears only when the ensemble of shrunken estimates is treated as an estimate of \(G\). Averaging pulls every estimate toward the center, so the collection \(\{\hat\theta_i^{\mathrm{PM}}\}\) is systematically narrower than \(G\), and any functional of \(G\) that depends on spread or shape, a variance, a tail share, a percentile cutoff, inherits the compression. The size of the compression is governed by measurement precision, a fact quantified in Chapter 3 and approximated by a particularly simple reference calculation: under a normal–normal working model with constant measurement-error variance, the posterior-mean ensemble’s standard deviation is \(\sqrt{\rho}\) times the latent standard deviation. The general law-of-total-variance result is underdispersion; the \(\sqrt{\rho}\) factor is not an exact identity for the nonlinear, heteroskedastic IRT likelihood or for a DPM posterior.

Figure 2.1 shows the phenomenon and both of the book’s levers acting on it in two selected bimodal replications from the study. The display is an exploratory mechanism illustration; its caption reports where the selected loss ratios fall among their sibling replications.

Two selected-replication density panels. Color and line type distinguish truth and estimates; the low-reliability panel shows PM compression and score-associated ripples, while the high-reliability panel compares Gaussian-GR and DP-GR shape.
Figure 2.1: In these replications the summary repairs spread at low reliability, and the prior separates the two fitted densities only at high reliability. Point-estimate kernel densities are compared with the shaded true-ability density for a bimodal N = 500 Rasch test; color and line type identify series. At reliability 0.5, Gaussian-PM estimates are strongly compressed and visibly rippled by score-associated bands from the short form. The within-band variation visible in the companion scatter is consistent with joint item/population estimation. Gaussian-GR restores spread, but none of the displayed curves recovers two modes in this selected replication. At reliability 0.9, the GR curves have similar spread and differ in valley shape. These are not condition averages. The selected high- and low-reliability DP-focused/Gaussian GR KS ratios are 0.44 (about the 50th empirical percentile; sibling range 0.33-0.71); 0.99 (about the 80th empirical percentile; sibling range 0.86-1.37) among 20 siblings per condition; the disclosed percentile/range check addresses selection sensitivity but neither supplies pointwise uncertainty bands nor establishes a mechanism.

2.2 Summaries as answers to stated problems

The repair for underdispersion is not to abandon the posterior but to ask it a different question. Each summary in the study is tied to a distinct decision problem, and the study’s estimand map (Figure 2.2) organizes them by the question they answer. PM is the exact personwise squared-error action; the production CB and GR calculations are the finite-sample or finite-MCMC implementations described below and in Appendix C.

Diagram mapping three inferential questions through their target parameters and loss functions to the summary that minimizes each loss.
Figure 2.2: The estimand map. Each inferential question fixes a target and a loss, and each is associated with a different action on the same posterior. For rank loss, the summaries studied here are empirically almost indistinguishable (Chapter 13); that result is bounded to this candidate set and simulation design.

The posterior mean (PM) answers the individual question. The general constrained-Bayes (CB) action minimizes squared error subject to matching the joint posterior’s target ensemble mean and variance (Ghosh, 1992). The production implementation uses the posterior-independence specialization based on marginal variances rather than the full cross-person posterior covariance. Both forms are positive affine rescalings of the PM set, so they preserve PM ordering and standardized shape while changing spread; only the joint form exactly enforces the general posterior constraint. The triple-goal estimator (GR) of Shen & Louis (1998) goes further: it estimates ranks and the ensemble distribution, then assigns the corresponding estimated distribution quantile to each person. In this study, expected ranks and the pooled posterior EDF are approximated from retained MCMC draws, with type-8 interpolation used for the final quantiles. The resulting GR set has the estimated distribution’s shape and the estimated ranks’ order; its individual estimates can have higher squared error while improving ensemble recovery, a tradeoff this study measures (Section 12.3). All three summaries are computed from the same fitted posterior, so in the study’s design the summary is a within-fit factor: nothing about the model, the data or the sampler changes between them.

2.3 The prior lever: Dirichlet process mixtures

The second lever replaces the normality assumption on \(G\) itself. The study’s flexible arms place a Dirichlet process mixture (DPM) prior on the ability distribution: a DP prior (Ferguson, 1973) draws a random discrete measure, and a normal kernel at each atom smooths it into a density. Figure 2.3 walks the construction.

Three panels showing the stick-breaking construction, random discrete measures at two concentration values, and smooth multimodal density draws from the DP mixture prior.
Figure 2.3: A Dirichlet process prior is a recipe for random distributions. Panel (a): the stick-breaking construction takes a stick of length one and breaks off a Beta-distributed fraction at a time; the fragments are the weights of a discrete random measure. Panel (b): each weight is attached to a location drawn from the base distribution, giving a random discrete measure G; the concentration parameter alpha governs how many atoms carry visible weight (a few at alpha = 1, many small ones at alpha = 10). Panel (c): the study’s model places a normal kernel at each atom, so the latent density is a countable mixture that can be unimodal, skewed or multimodal without committing to any of them; four independent prior draws are shown. The focused arm calibrates a Gamma prior on alpha so the implied number of occupied clusters concentrates on small values.

Two properties matter for reading the results. First, a normal-kernel DPM has enough support to represent densities close to normal, and a normal base-and-kernel construction can center its prior mean density on a normal shape. That support and centering do not make a DPM fit identical to the parametric Gaussian analysis: the parameter spaces, prior weights and posterior averaging remain different even when one cluster dominates. Whether the flexible model changes loss appreciably under a normal population is therefore an empirical question, which the study’s negative control addresses (Section 7.1). Second, the concentration parameter \(\alpha\) governs how readily the prior supports additional occupied clusters (Antoniak, 1974; Escobar & West, 1995), which makes its own prior a genuine modeling decision. The study fits two versions: a focused arm, whose Gamma prior on \(\alpha\) is calibrated by a numerical exact-moment procedure (Lee, 2026) so that the implied number of occupied clusters concentrates on small values, and a broad arm calibrated to a higher and more diffuse cluster-count target, so that the elicitation’s value is itself measured (H5; Section 8.5). The calibration machinery is summarized in Appendix C and treated in full in the theory volume; sampling uses the truncated stick-breaking representation with truncation levels recorded per fit (Section 14.2).

2.4 Why the levers must be studied jointly

Each lever has a literature; their pairing is less studied in IRT, and the results chapters show why the pairing is where the action is. The flexible prior changes what the posterior knows about shape; the summary changes whether that knowledge reaches the reported numbers. A DPM prior under PM summarization has its better shape estimate shrunk away before reporting (Section 9.4 shows this arithmetic cell by cell); a GR summary under a Gaussian prior faithfully reproduces a shape estimate that cannot bend away from normality. Only the combination among those studied lets both forms of flexibility reach the report, and only when the data carry such information. In this design, the prior controls the modeled range of population shapes, the summary controls how that posterior information enters the reported estimate set, and the measurement controls how strongly the data distinguish those shapes. The measurement component is the subject of the next chapter.

Antoniak, C. E. (1974). Mixtures of Dirichlet processes with applications to Bayesian nonparametric problems. The Annals of Statistics, 2(6), 1152–1174. https://doi.org/10.1214/aos/1176342871
Escobar, M. D., & West, M. (1995). Bayesian density estimation and inference using mixtures. Journal of the American Statistical Association, 90(430), 577–588. https://doi.org/10.1080/01621459.1995.10476550
Ferguson, T. S. (1973). A Bayesian analysis of some nonparametric problems. The Annals of Statistics, 1(2), 209–230. https://doi.org/10.1214/aos/1176342360
Ghosh, M. (1992). Constrained Bayes estimation with applications. Journal of the American Statistical Association, 87(418), 533–540. https://doi.org/10.1080/01621459.1992.10475236
Lee, J. (2026). Design-conditional prior elicitation for Dirichlet process mixtures: A unified framework for cluster counts and weight control. arXiv preprint arXiv:2602.06301.
Shen, W., & Louis, T. A. (1998). Triple-goal estimates in two-stage hierarchical models. Journal of the Royal Statistical Society Series B: Statistical Methodology, 60(2), 455–471. https://doi.org/10.1111/1467-9868.00135