Appendix C — Condensed theory

This appendix states, without proof, the formal results the main text leans on. Full developments are in the project’s theory volume; derivations specific to this study’s design are in the methods papers cited below.

C.1 Reliability functionals and the targeting problem

For a form \(F\) applied to population \(G\), let \(J_F(\theta)\) denote test information and define the study’s operational design coefficient as

\[ \rho_{\mathrm{info}}(F,G) = \frac{\sigma^2_\theta} {\sigma^2_\theta + E_G\{J_F(\theta)^{-1}\}}. \]

This coefficient treats inverse information as a local error-variance approximation. It is not algebraically equal to \(E_G\{\operatorname{Var}(\hat\theta\mid\theta)\}\) for the fitted finite-sample estimator, to conventional observed-score reliability, or to an exact EAP shrinkage coefficient. It omits, among other things, estimator bias and item-parameter uncertainty. Jensen’s inequality gives \(E_G\{J_F(\theta)^{-1}\}\geq E_G\{J_F(\theta)\}^{-1}\), so the displayed coefficient is no larger than the coefficient formed from reciprocal mean information; equality holds only when information is constant \(G\)-almost surely.

The V3 ladder targets and independently certifies \(\rho_{\mathrm{info}}\). It is not a one-dimensional length-only intervention: tiers use nested item subsets and each model-by-shape-by-tier-by-form calibration cell has a fitted global discrimination multiplier c_star, which is applied when responses are generated. The resulting contrasts therefore concern achieved reliability along this joint length-plus-scale ladder. A uniqueness result for a discrimination multiplier at fixed length does not imply uniqueness of an integer test length (Lee, 2025).

C.2 Identification under a DPM ability prior

With \(G\) unrestricted, the IRT likelihood is invariant to location transfers between \(G\) and item difficulties and, in the 2PL, to corresponding scale transfers involving item discriminations. The fitted Rasch models impose mean item difficulty zero. The fitted 2PL models impose both mean item difficulty zero and mean log discrimination zero, equivalently a geometric mean discrimination of one. These constraints select a representative of the likelihood-equivalent location or location-scale orbit; they do not prove identification of an unrestricted mixing distribution.

For loss calculation, the simulation uses the known generating item difficulties and scaled discriminations stored with each frozen dataset to map constrained-item posterior summaries back to the raw standardized DGP scale (Section 5.3). This is a simulation-specific alignment to known truth, not an identification procedure available when the generating parameters are unknown. Finite-item Rasch data identify anchored item parameters and only limited functionals of \(G\); no general finite-item semiparametric 2PL identification theorem is asserted here. The study therefore evaluates summaries of the realized ability ensemble on a common known scale rather than claiming recovery of every feature of an unrestricted population mixing law.

C.3 The DP prior and cluster-count calibration

Under \(G \sim \mathrm{DP}(\alpha, G_0)\) with the stick-breaking representation \(G = \sum_k w_k \delta_{\phi_k}\), \(w_k = v_k \prod_{j<k}(1 - v_j)\), \(v_k \sim \mathrm{Beta}(1, \alpha)\), the number of occupied clusters among \(J\) draws has the Antoniak distribution, with exact expectation \(\sum_{j=1}^J \alpha/(\alpha+j-1)\) and approximation \(\alpha \log\{(\alpha + J)/\alpha\}\) (Antoniak, 1974). V3 constructs both DP-arm Gamma priors with DPprior_fit(..., method = "A2-MN", M = 200): the numerical A2 procedure uses exact induced cluster-count moments with quadrature and nonlinear solving (Lee, 2026). Production does not rely on a closed-form inversion of the logarithmic approximation, and it supplies the resulting Gamma parameters explicitly rather than entering the package’s automatic fallback path. The mapping targets cluster-count beliefs, not atom weights; dominance-tail behavior (the probability that one atom carries most of the mass) is a separate monitored functional, and H5 measures the empirical difference between the focused and broad calibrations.

C.4 The summaries as decision rules

Let \(p(\boldsymbol\theta \mid \mathbf{y})\) be the joint posterior, let \(\boldsymbol\eta=E(\boldsymbol\theta\mid\mathbf y)\), and let \(H_2=\sum_i(\eta_i-\bar\eta)^2\). PM minimizes \(E\{\sum_i(\hat\theta_i-\theta_i)^2\mid\mathbf y\}\). The general Ghosh CB action minimizes the same loss under posterior ensemble-moment constraints and has the affine form

\[ \hat\theta_i^{\mathrm{CB,joint}} =\bar\eta+a_{\mathrm{joint}}(\eta_i-\bar\eta),\qquad a_{\mathrm{joint}}=\sqrt{1+H_1/H_2}, \]

where \(H_1\) is the trace of the joint posterior covariance after centering each posterior draw by its ensemble mean (Ghosh, 1992). The production code instead computes

\[ a_{\mathrm{impl}} =\sqrt{1+\bar v/s^2_{\eta}}, \]

from the mean marginal posterior variance \(\bar v\) and the sample variance of the posterior means \(s^2_{\eta}\). This is the posterior-independence specialization. Using the sample variance retains the corresponding finite-\(N\) factor, but cross-person posterior covariances are omitted. It remains a positive affine map, so it preserves PM ordering and standardized shape, but it need not exactly satisfy the general joint-posterior variance constraint when the person posteriors are dependent.

GR estimates posterior expected ranks, estimates the distribution by the posterior mean EDF, and assigns person \(i\)

\[ \hat\theta_i^{\mathrm{GR}} =\hat F^{-1}\{(\hat r_i-1/2)/N\}. \]

In production, expected ranks and \(\hat F\) are finite-MCMC approximations based on retained draws, \(\hat r_i\) is the integer rank of the estimated expected rank, and the final inverse EDF uses type-8 quantile interpolation. Type 8 is an implementation convention rather than part of the abstract Shen–Louis action (Shen & Louis, 1998).

The theoretical incompatibility claim needed here is an existence statement. There are nondegenerate overlapping posterior configurations for which the unique personwise PM vector does not supply the ISEL-optimal equal-mass distribution; thus one reported vector need not optimize both individual and ensemble losses. This counterexample does not imply conflict for every nondegenerate posterior and does not make the direction or size of a trade universal. Registered H7 is the empirical test for this simulation grid: it asks whether the PM-versus-GR loss gap favors GR on KS more than it favors PM on MSEL. The estimated contrast is 0.281 (95% CI [0.210, 0.351], Holm-adjusted \(p = 8.6 \times 10^{-6}\)). H7 supplies evidence about the studied design; it is not a proof of a universal theorem.

C.5 Loss definitions

For estimate set \(\hat{\boldsymbol\theta}\) and realized abilities \(\boldsymbol\theta\) of size \(N\): \(\mathrm{MSEL} = N^{-1}\sum_i(\hat\theta_i - \theta_i)^2\); \(\mathrm{MSELR} = N^{-1}\sum_i\{R(\hat\theta_i)/N - R(\theta_i)/N\}^2\) with ties averaged; \(\mathrm{KS} = \sup_t \lvert \hat F_{\hat\theta}(t) - \hat F_{\theta}(t)\rvert\) over right-continuous EDFs; the quantile loss is \(\sum_q w_q\{\hat Q(q) - Q(q)\}^2\) over \(q \in \{.10,.25,.50,.75,.90\}\) (empirical type-8 quantiles); the tail loss is \(\sum_c w_c\{\hat p(c) - p(c)\}^2\) over cutoffs \(c \in \{-2,-1.5,1.5,2\}\). For the registered replicate-level H1–H5 ratios governed by Addendum 002, a floor \(f = \max(0.001,\, 0.05\,b)\), where \(b\) is the cell’s equal-form mean comparator loss, is applied symmetrically to both operands, so that \(y = \log\{\max(L_{\mathrm{cand}}, f) / \max(L_{\mathrm{comp}}, f)\}\) (Addendum 002). This floor does not automatically carry into descriptive ratios of condition means; H7 uses the separate share-loss floor specified by the amendment to Addendum 003. The beta loss is MSEL on item difficulties under PM.

Antoniak, C. E. (1974). Mixtures of Dirichlet processes with applications to Bayesian nonparametric problems. The Annals of Statistics, 2(6), 1152–1174. https://doi.org/10.1214/aos/1176342871
Ghosh, M. (1992). Constrained Bayes estimation with applications. Journal of the American Statistical Association, 87(418), 533–540. https://doi.org/10.1080/01621459.1992.10475236
Lee, J. (2025). Reliability-targeted simulation of item response data: Solving the inverse design problem. arXiv preprint arXiv:2512.16012.
Lee, J. (2026). Design-conditional prior elicitation for Dirichlet process mixtures: A unified framework for cluster counts and weight control. arXiv preprint arXiv:2602.06301.
Shen, W., & Louis, T. A. (1998). Triple-goal estimates in two-stage hierarchical models. Journal of the Royal Statistical Society Series B: Statistical Methodology, 60(2), 455–471. https://doi.org/10.1111/1467-9868.00135