| Tier | Entries | Meaning |
|---|---|---|
| primary-held | 292 | Held in a reference library, not read at a locator for this book |
| primary-read | 74 | Independent literature opened at the locator cited; annotated here |
| programme-artifact-read | 8 | Programme-owned book, manuscript, package, code, or table opened directly; supports fidelity to that artifact, not independent corroboration |
| secondary | 3 | Cited as a secondary treatment, and marked as such where used |
Appendix D — Annotated Bibliography
An annotated bibliography is a document made entirely of claims about sources. That is the exact class of claim four external review cycles found errors in — a theorem attributed to the wrong number, an absence certified in a paper whose discussion states the opposite, a condition dropped from a proposition. So this appendix is scoped by a number rather than by a plan.
82 of 377 bibliography entries were opened at a locator and are annotated: 74 primary-read works from the independent literature and 8 programme-artifact-read works produced by this research programme. The second count is deliberately separate. Reading this programme’s own book, manuscript, package, code, or frozen table can establish what that artifact says or does; it cannot supply independent corroboration. The rest are held, or cited as secondary treatments, and the register says which.
Nothing below is written from memory. Each annotation is assembled by code/R/17-annotated-bibliography.R from what the project has already verified about that source — a semantic receipt, a numbered result sourced to it, a correction it grounds, a citation verification recorded at a locator, a bounded narrative receipt, or a programme artifact receipt — and the Evidence column names which. The Tier and Authority columns prevent a programme artifact from being displayed as independent literature.
Citation verification is the narrowest evidence route for an external source. manifest/citation-verification.csv records what happened when a citation was checked against the source itself: which sites were involved, which file was opened, what was found there, and on what date. Only rows marked level = locator and confirmed at that locator are admitted as evidence here. Rows marked ATTRIBUTION SOUND compared a title and its distinctive terms against the citing claim without opening a passage; they are refused, and the sources they cover stay primary-held, so they never reach this appendix at all.
That register also carries the checks that did not come out clean. Of 60 checks, 44 were made at a locator and 6 of those forced a correction to this book — a rate worth stating, because it is the argument for doing the checks at all. They were not all the same kind of error:
- A locator that points at the wrong page. van der Ark (2025) was cited in Chapter 8 at two page numbers and both were one page late — the Lord and Novick lower-bound result is on p. 1679, not p. 1680, and the \(\lambda_3 = \alpha\) identity is in § 3.7 on p. 1686, not p. 1687, where an unrelated simulation section begins. The content was exactly as described; only the locators were wrong.
- A number that was not the source’s number. Chapter 1 gave a Grade 9 battery as reliable at “.83 to .92 overall”; Abedi’s Table 9 runs .80 to .92, a bound that excluded four of seven subscales. The corrected figures make the point better, since the table also gives .53 to .83 for English language learners.
- A claim stronger than the source’s. Chapter 8 put Hussey et al.’s excess of \(\alpha\) values “just above” \(.70\); they locate it at \(.70\), which is also their title.
- A citation the source does not support at all. Chapter 8 grouped Cronbach and Shavelson (2004) with Sijtsma as making the point about Nunnally’s threshold. A full-text scan finds no occurrence of Nunnally, .70, or threshold anywhere in it.
- An attribution to the wrong work by the right author. Chapter 12 credited the held Robbins reprint with the compound-decision framing; that phrase appears in it only in its reference list, pointing at a different paper.
- A disclaimer that was itself wrong. Chapter 8 declined to give a page for Equation 8.3 because “the held copy is unpaginated OCR”. The held copy is a clean paginated scan and the equation is Cronbach’s (2) on p. 299.
Only the first of those six is the error a careless reader would predict. The rest are failures of what was claimed, not of where it was looked for, and none would have been caught by any check short of opening the source.
One thing is deliberately not an evidence type. manifest/citation-context-match.csv records, for 66 narrative citations, which page of a PDF shares the most words with the citing sentence. Review 021 used that table to promote 60 sources to primary-read; the promotion was reverted, because keyword overlap is not reading — several matches landed on OCR noise, and one covered a source this book cites at equation level. The table survives as a navigation aid for whoever reads those sources next, it declares is_reading = NO on every row, and V1 refuses to let it claim otherwise.
D.1 What the tiers mean
primary-held is not a criticism of a source; it is a statement about this book. It means the work is in one of the nine reference libraries and was available, but no claim here rests on having read it at a locator. Where such a work is named in the text, the convention (Chapter 21, Chapter 22) is to name it without an @ citation and say that what was read is another author’s use of it.
programme-artifact-read is a different statement: the artifact was opened directly, but it belongs to the same research programme. Its receipt must state both what was checked and the authority boundary. It may establish cross-book or code-to-prose fidelity; it may not be counted as independent confirmation.
The 82 annotated below are the sources and programme artifacts this book reads directly.
D.2 The annotations
Ordered by the literature group each belongs to, which is assigned by the first chapter that cites it rather than by judgement. What was verified is the receipt’s own text where one exists, so it records what was checked, not what the source is generally about.
D.2.1 Bayesian hierarchy, shrinkage, and empirical Bayes
Albert (1992). Bayesian estimation of normal ogive item response curves using {Gibbs} sampling
Title and abstract (p. 251): ‘Bayesian Estimation of Normal Ogive Item Response Curves Using Gibbs Sampling’, two-parameter normal ogive, Gibbs sampling of the joint posterior of ability and item parameters. P. 258 gives the augmentation itself: the Z_ij ‘can be viewed as continuous underlying or latent variables’, and ‘with the introduction of Z, we are interested in simulating from the joint posterior distribution of (Z, theta, xi)’. That is exactly ch. 10’s description – latent continuous responses imputed, conditional conjugacy, Gibbs throughout.
primary-read · Independent literature read at a locator · cited in ch. 10 · read at abstract p. 251; augmentation construction p. 258 · citation verification VERIF-ALBERT-1992 · read on 2026-08-06
Bock and Mislevy (1982). Adaptive {EAP} estimation of ability in a microcomputer environment
Title: ‘Adaptive EAP Estimation of Ability in a Microcomputer Environment’ (Bock and Mislevy). The abstract describes expected a posteriori estimation of ability ‘based on numerical evaluation of the mean and variance’ over quadrature points. Ch. 11 cites it as the standard psychometric reference for EAP scoring by quadrature, which is what the paper is.
primary-read · Independent literature read at a locator · cited in ch. 11 · read at abstract, p. 431 · citation verification VERIF-BOCKMISLEVY-1982 · read on 2026-08-06
Bürkner (2021). Bayesian item response modeling in with and
Title: ‘Bayesian Item Response Modeling in R with brms and Stan’, Journal of Statistical Software, November 2021, Volume 100, Issue 5, matching the bibliography. Ch. 10 cites it for the Stan route for the parametric model, which is the paper’s subject.
primary-read · Independent literature read at a locator · cited in ch. 10 · read at title and abstract · citation verification VERIF-BURKNER-2021 · read on 2026-08-06
Efron (2016). Empirical {Bayes} deconvolution estimates
Summary: ‘An unknown prior density g(theta) has yielded realizations Theta_1, …, Theta_N. They are unobservable, but each Theta_i produces an observable value X_i according to a known probability mechanism.’ Ch. 12’s sentence is a close paraphrase of this, and the g-modelling proposal ch. 12 also attributes to the paper is one of its stated key words and the subject of sec. 1.
primary-read · Independent literature read at a locator · cited in ch. 12 · read at summary and sec. 1, p. 1 · citation verification VERIF-EFRON-2016 · read on 2026-08-06
Fox (2025). Redefining item response models for small samples
Abstract: ‘The two-parameter IRT model is redefined by analytically integrating out the random person factor in the latent response formulation of the model to make it suitable for small sample applications.’ Ch. 10’s description – integrating the person factor out analytically to obtain a model that behaves in small samples – is this sentence, and ch. 10 states it is read at the level of stated aims.
primary-read · Independent literature read at a locator · cited in ch. 10 · read at abstract · citation verification VERIF-FOX-2025 · read on 2026-08-06
Gelman (2006). Prior distributions for variance parameters in hierarchical models (comment on article by {Browne} and {Draper})
Sec. 2.2 is titled ‘Improper limit of a prior distribution’ and states that the inverse-gamma(epsilon, epsilon) model ‘does not have any proper limiting posterior distribution’. That is the claim ch. 10 attributes to it, at the cited section.
primary-read · Independent literature read at a locator · cited in ch. 10 · read at secs. 2.1-2.2 · citation verification VERIF-GELMAN-2006 · read on 2026-08-06
Gilholm et al. (2021). Bayesian hierarchical multidimensional item response modeling of small sample, sparse data for personalized developmental surveillance
Title: ‘Bayesian Hierarchical Multidimensional Item Response Modeling of Small Sample, Sparse Data for Personalized Developmental Surveillance’, EPM 2021, 81(5), 936-956, matching the bibliography. Ch. 10’s description – hierarchical multidimensional IRT fitted to small, sparse samples – is the title’s own content, and ch. 10 states it is read at the level of stated aims, which is what was checked.
primary-read · Independent literature read at a locator · cited in ch. 10 · read at title and abstract, p. 936 · citation verification VERIF-GILHOLM-2021 · read on 2026-08-06
Goldstein and Spiegelhalter (1996). League tables and their limitations: Statistical issues in comparisons of institutional performance
league-table uncertainty and the distinction between point estimates and rank decisions
primary-read · Independent literature read at a locator · cited in ch. 12, 17, 20 · read at sec. 3 · receipt PV-SRC-GOLDSTEIN; result PROP-17-2 · read on 2026-08-05
James and Stein (1992). Estimation with quadratic loss
Sec. 2, ‘Inadmissibility of the Usual Estimator …’, gives the explicit estimator phi_1(X) = (1 - (p-2)/||X||^2) X, which is the form ch. 12 quotes at the cited section. Citation of the 1992 reprint for the 1961 original follows DECISIONS (a4) and is stated in the text.
primary-read · Independent literature read at a locator · cited in ch. 12 · read at sec. 2 · citation verification VERIF-JAMESSTEIN-1992 · read on 2026-08-06
Laird and Louis (1989). Empirical {Bayes} ranking methods
generic empirical-Bayes rank warning and stochastic-order exception
primary-read · Independent literature read at a locator · cited in ch. 12, 17 · read at sec. 1 and Appendix A, Theorem A; sec. 1 gives the generic warning and Appendix A the stochastic-order exception; Shen and Louis (1998) Theorem 2 corroborates the exception; the Rasch/known-2PL monotone-likelihood-ratio verification is inline · receipt PV-SRC-LAIRD; result PROP-12-3, PROP-17-2 · read on 2026-08-05
Louis (1984). Estimating a population of parameter values using {Bayes} and empirical {Bayes} methods
Gaussian compound-model underdispersion and moment-matching lineage
primary-read · Independent literature read at a locator · cited in ch. 11, 12, 18 · read at sec. 2.1, pp. 394-395; sec. 2.1 Gaussian compound model; adapted to the IRT normal-working likelihood and proved inline by completing the square; sec. 2.1 gives the Gaussian ensemble under-dispersion and moment repair; the general law-of-total-variance identity and constant-error specialization are proved inline · receipt PV-SRC-LOUIS; result THM-11-1, PROP-11-3, THM-18-1 · read on 2026-08-05
Patz and Junker (1999). A straightforward approach to {Markov} chain {Monte} {Carlo} methods for item response models
Abstract reads: ‘In contrast to this two-stage E-M approach, MCMC methods treat item and subject parameters at the same time; this allows us to incorporate standard errors of item estimates into trait inferences, and vice versa.’ Ch. 10 quotes both fragments, with ‘[s]’ correctly marking the only alteration. The Metropolis-within-Gibbs attribution is the paper’s own stated methodology in the same abstract.
primary-read · Independent literature read at a locator · cited in ch. 10 · read at abstract, p. 146 · citation verification VERIF-PATZ-1999 · read on 2026-08-06
Robbins (1992). An empirical {Bayes} approach to statistics
The held document is the reprint of ‘An Empirical Bayes Approach to Statistics’. Its opening states ch. 12’s construction directly: X with a known distribution given an unknown parameter, the parameter itself random with prior G, the unconditional distribution as the mixture, and squared-error risk. But the phrase ‘compound decision problem’ appears in this document ONLY in its reference list, as reference [4], H. Robbins, ‘Asymptotically subminimax solutions of compound statistical decision problems’ – a different, unheld paper. The book’s provenance line attributed the compound-decision framing to the held reprint; corrected to attribute only the empirical Bayes argument to it and to name where the framing actually comes from.
primary-read · Independent literature read at a locator · cited in ch. 12 · read at opening derivation, p. 1 of the reprint; reference [4] · citation verification VERIF-ROBBINS-1992 · read on 2026-08-06
Rubin (1981). Estimation in parallel randomized experiments
Abstract: ‘In the example considered here, randomized experiments were conducted in eight schools to determine the effectiveness of special coaching programs for the SAT. The purpose here is to illustrate Bayesian and empirical Bayesian techniques.’ That is the eight-schools setup and the stated aims ch. 12 says it uses, and nothing more is drawn from it. Bibliographic check: the article’s own OCR-garbled first-page header reads ‘pp. 377-400’, but the running head on the last text page reads 401, so the bibliography’s 377–401 is correct and was left alone.
primary-read · Independent literature read at a locator · cited in ch. 12 · read at abstract, p. 377 · citation verification VERIF-RUBIN-1981 · read on 2026-08-06
Stein (1956). Inadmissibility of the usual estimator for the mean of a multivariate normal distribution
Title: ‘Inadmissibility of the Usual Estimator for the Mean of a Multivariate Normal Distribution’. Sec. 1 states that under summed squared-error loss the usual estimator ‘is admissible for n <= 2, but inadmissible for n >= 3’. Ch. 12 cites Stein for the inadmissibility priority only, which is exact. NOTE: the held copy is the Berkeley Symposium proceedings and its pages carry no usable running numbers, so the locator is by section.
primary-read · Independent literature read at a locator · cited in ch. 12 · read at title and sec. 1 Introduction · citation verification VERIF-STEIN-1956 · read on 2026-08-06
D.2.2 Bayesian nonparametrics and identification
Antonelli et al. (2016). Mitigating bias in generalized linear mixed models: The case for {Bayesian} nonparametrics
Sec. 3 synthesises the strategies ch. 15 lists, and all six are present: moment matching (choosing (psi_1, psi_2) ‘based on the moments of a Gamma distribution’, p. 83), a diffuse Gamma, Kullback-Leibler divergence between the elicited prior on K and the induced prior (p. 83), a fixed value of alpha chosen from an a priori guess at the number of clusters (p. 84), and empirical Bayes and importance sampling, which p. 85 groups together with the diffuse and KL priors as the analyses compared. The abstract also singles out importance sampling and empirical Bayes as broadly reasonable.
primary-read · Independent literature read at a locator · cited in ch. 15 · read at sec. 3 Selection of alpha, pp. 83-85 · citation verification VERIF-ANTONELLI-2016 · read on 2026-08-06
Antoniak (1974). Mixtures of {Dirichlet} processes with applications to {Bayesian} nonparametric problems
Source for THM-15-4: P(K_J = k) via unsigned Stirling numbers of the first kind, E[K_J] = alpha{psi(alpha+J) - psi(alpha)}, and K_J sufficient for alpha
primary-read · Independent literature read at a locator · cited in ch. 15 · read at sec. 4, printed pp. 1160-1163; the digamma closed form is one line from his sum and is verified numerically in V8 · result THM-15-4 · read on 2026-08-04
Blackwell and MacQueen (1973). {Ferguson} distributions via {Pólya} urn schemes
Source for THM-15-3: A Polya sequence’s empirical measure converges to a discrete P distributed as Ferguson’s DP, and the draws are iid given P; the CRP is that partition structure
primary-read · Independent literature read at a locator · cited in ch. 15 · read at abstract and sec. 1; the reduction of the restaurant metaphor to this theorem is this book’s presentation and answers R2-4b · result THM-15-3 · read on 2026-08-04
Dorazio (2009). On selecting a prior for the precision parameter of {Dirichlet} process mixture models
The paper derives, for a Gamma(a, b) prior on the concentration parameter, the probability mass function of the prior induced on the number of clusters K by integrating Pr(K = k | alpha, n) against that Gamma. That is exactly the induced-cluster route ch. 15 attributes to it.
primary-read · Independent literature read at a locator · cited in ch. 15 · read at the induced prior on K · citation verification VERIF-DORAZIO-2009 · read on 2026-08-06
Duncan and MacEachern (2008). Nonparametric {Bayesian} modelling for item response
Title: ‘Nonparametric Bayesian modelling for item response’ (Duncan and MacEachern), Statistical Modelling 2008; 8(1): 41-66 – matching the bibliography’s volume and page range exactly. Ch. 16 cites it as the nonparametric-Bayes treatment of the ability distribution in the item-response setting, which is the paper’s stated subject.
primary-read · Independent literature read at a locator · cited in ch. 16 · read at title and abstract, p. 41 · citation verification VERIF-DUNCAN-2008 · read on 2026-08-06
Ferguson (1973). A {Bayesian} analysis of some nonparametric problems
Source for DEF-15-1: The Dirichlet process with parameter a finite non-null measure alpha: every measurable partition has Dirichlet-distributed probabilities
primary-read · Independent literature read at a locator · cited in ch. 15 · read at secs. 1 and 3; alpha is a MEASURE, whose total mass is the concentration and whose normalization is the base measure · result DEF-15-1 · read on 2026-08-04
Giordano et al. (2023). Evaluating sensitivity to the stick-breaking prior in {Bayesian} nonparametrics (with discussion)
Abstract: ‘due to the flexibility of these models, the consequences of prior choices can be opaque’ and ‘prior choice can have a substantial effect on posterior inferences’. Ch. 15 quotes both, and the paper’s stated contribution is sensitivity to the stick-breaking prior, as ch. 15 says. NOTE: the held PDF is the advance copy paginated 1-34; the bibliography’s 287-366 is the published discussion-paper range, so no page locator is given.
primary-read · Independent literature read at a locator · cited in ch. 15 · read at abstract (advance copy; the held file does not carry the final pagination) · citation verification VERIF-GIORDANO-2022 · read on 2026-08-06
Greve et al. (2022). Spying on the prior of the number of data clusters and the partition distribution in {Bayesian} cluster analysis
Abstract: ‘a major empirical challenge involving the use of these models is in the characterisation of the induced prior on the partitions’, and the paper introduces an approach to compute descriptive statistics of that prior. Ch. 15’s quotation and its gloss both match. NOTE: the held PDF is the arXiv preprint, paginated 1-35, while the bibliography records the published range 205-229, so no page locator is given.
primary-read · Independent literature read at a locator · cited in ch. 15 · read at abstract (preprint copy; the held file does not carry the journal’s pagination) · citation verification VERIF-GREVE-2022 · read on 2026-08-06
Ishwaran and Zarepour (2000). Markov chain {Monte} {Carlo} in approximate {Dirichlet} and beta two-parameter process hierarchical models
Abstract: the paper considers ‘a truncation approximation as well as a weak limit approximation’ for the Dirichlet process, states that ‘the adequacy of the approximation can be easily computed from the output of the Gibbs sampler’, and reports that the truncation approximation ‘offers an exponentially higher degree of accuracy over the weak limit approximation for the same computational effort’. Chs. 15 and 16 cite the paper for studying what the truncation costs, which is exactly this.
primary-read · Independent literature read at a locator · cited in ch. 15, 16 · read at abstract, p. 371 · citation verification VERIF-ISHWARAN-2000 · read on 2026-08-06
Lee (2026) [lee_dpprior_2026]. {DPprior}: Principled prior elicitation for {Dirichlet} process mixture models
The frozen package source implements DPprior_fit, DPprior_dual, and prob_w1_exceeds for design-conditional concentration and weight elicitation. Authority boundary: Direct inspection of programme-owned code establishes implementation fidelity only; claims of calibration quality require separate evidence.
programme-artifact-read · Programme-owned artifact; fidelity evidence, not independent corroboration · cited in ch. 15, 29, F · read at targeted-DPMirt-simulation-codebase-v3/frozen-packages/DPprior: R/16_DPprior_fit.R line 289; R/15_dual_anchor.R line 230; R/08_weights_w1.R line 232 · programme receipt PAR-004 · read on 2026-08-07
Lo (1984). On a class of {Bayesian} nonparametric estimates: {I}. Density estimates
Lo writes the Dirichlet-process mixture density estimator as a posterior average of the mixing distribution and gives the no-sample Bayes estimator centred on the parametric base model; this is the attribution used around Equation 15.4.
primary-read · Independent literature read at a locator · cited in ch. 15 · read at Section 2, pp. 352-354 · narrative receipt NSR-001 · read on 2026-08-04
Miyazaki and Hoshino (2009). A {Bayesian} semiparametric item response model with {Dirichlet} process priors
Abstract (p. 375): ‘since only limited patterns of shapes can be obtained from logistic models or normal ogive models, there is a possibility that the model applied does not fit the data’, restated on p. 376. Ch. 16’s quotation is exact and its gloss – that the construction relaxes the item side – matches the paper’s stated target.
primary-read · Independent literature read at a locator · cited in ch. 16 · read at abstract p. 375; restated p. 376 · citation verification VERIF-MIYAZAKI-2009 · read on 2026-08-06
Murugiah and Sweeting (2012). Selecting the precision parameter prior in {Dirichlet} process mixture models
Abstract: a framework for ‘the specification of the hyperparameters associated with the prior for the precision parameter that can be used both in the presence or absence of subjective prior information about the level of clustering’. Sec. 1 adds that in the absence case the hyperparameters are chosen ‘in a scalable way in order to produce reasonable performance characteristics’ – the defaults ch. 15 refers to. Both halves of ch. 15’s claim match.
primary-read · Independent literature read at a locator · cited in ch. 15 · read at abstract p. 1947; sec. 1 p. 1948 · citation verification VERIF-MURUGIAH-2012 · read on 2026-08-06
Neal (2000). Markov chain sampling methods for {Dirichlet} process mixture models
Title ‘Markov Chain Sampling Methods for Dirichlet Process Mixture Models’; the abstract presents auxiliary-parameter methods that are ‘more efficient than previous ways of handling general Dirichlet process mixture models with non-conjugate priors’, and sec. 1 sets out why Gibbs sampling is easy under conjugacy and hard without it. Ch. 15’s description – the standard catalogue, including the algorithms for non-conjugate kernels – matches.
primary-read · Independent literature read at a locator · cited in ch. 15 · read at abstract and sec. 1, p. 249 · citation verification VERIF-NEAL-2000 · read on 2026-08-06
Paganin et al. (2022). Computational strategies and estimation performance with {Bayesian} semiparametric item response theory models
Paganin et al. write logit(pi_ij) = lambda_i(eta_j - beta_i) with eta_j the latent ability of person j, lambda_i > 0 the discrimination, beta_i the difficulty, y_ij the response with i the item and j the person, G the latent distribution traditionally standard normal, and gamma_i = -lambda_i beta_i the slope-intercept form. Their eta_j is this books theta_p, which collides with this books eta_p for the posterior mean.
primary-read · Independent literature read at a locator · cited in ch. 16, B · read at sec. 2, eq. (1) and the paragraph defining lambda_i and beta_i; eq. (3) for the slope-intercept form · receipt PVII-SRC-PAGANIN · read on 2026-08-05
Rodríguez (2013). On the {Jeffreys} prior for the multivariate {Ewens} distribution
Abstract: ‘We derive the Jeffreys prior for the parameter of the Multivariate Ewens Distribution and study some of its properties. In particular, we show that this prior is proper and has no [finite moments]’. Ch. 15’s two claims – proper, and no finite moments – are the source’s own, from the abstract, as ch. 15 says.
primary-read · Independent literature read at a locator · cited in ch. 15 · read at abstract, p. 1539 · citation verification VERIF-RODRIGUEZ-2013 · read on 2026-08-06
Sethuraman (1994). A constructive definition of {Dirichlet} priors
Source for THM-15-2, PROP-15-5: Stick-breaking: Beta(1,alpha) sticks with iid base-measure locations give a Dirichlet process; A Dirichlet process draw is discrete with probability one
primary-read · Independent literature read at a locator · cited in ch. 15 · read at sec. 1 construction; Theorems 3.4 and 4.3 for the Dirichlet marginals and conjugacy; immediate from the stick-breaking form; Ferguson sec. 4 gives an independent construction · result THM-15-2, PROP-15-5 · read on 2026-08-04
Vicentini and Jermyn (2025). Prior selection for the precision parameter of {Dirichlet} process mixtures
Abstract: ‘Our goal is to specify the prior distribution p(alpha | eta), including its fixed parameter vector eta, in a way that is meaningful’, followed by a three-group categorisation of existing approaches, the first of which links p(alpha|eta) to the induced prior on the cluster count. Ch. 15’s statement of the question, and its quotation of ‘meaningful’, are the source’s.
primary-read · Independent literature read at a locator · cited in ch. 15 · read at abstract · citation verification VERIF-VICENTINI-2025 · read on 2026-08-06
D.2.3 Information, error, and reliability
Adams (2005). Reliability as a measurement design effect
Eq. (5) is Var(theta-hat)=sigma^2_theta + (1/N)sum sigma^2_{l,n}, the decomposition the book cites. Eq. (6) is person separation reliability, and Adams’s own text credits Wright and Stone (1979) exactly as the book says. Both equation NUMBERS confirmed against the page.
primary-read · Independent literature read at a locator · cited in ch. 08 · read at eqs. (5) and (6), p. 165 · citation verification VERIF-ADAMS-2005 · read on 2026-08-06
Andersson and Xin (2018). Large sample confidence intervals for item response theory reliability coefficients
Eq. (13) is rho_Theta(alpha) = integral I(theta;alpha)/(I(theta;alpha)+1) g(theta) d theta, on the unit-variance convention – the average-information functional ch. 8 attributes to it. The source itself attributes the form to Cheng et al. (2012) with Green et al. (1984) as origin, which is precisely the attribution chain ch. 8 states.
primary-read · Independent literature read at a locator · cited in ch. 08 · read at eq. (13), p. 35 · citation verification VERIF-ANDERSSON-2018 · read on 2026-08-06
Ark (2025). Standard errors for reliability coefficients
Both claims confirmed in the source, but BOTH page citations in the book were wrong by one page. (a) The Lord and Novick (1968, Thm 4.4.3) alpha-lower-bound statement is on p. 1679 (sec. 1 Introduction), not p. 1680; p. 1680 holds Table 1 and the software discussion. (b) ‘Guttman’s lambda_3 equals Cronbach’s alpha and also equals (J/(J-1))lambda_1 (Guttman, 1945)’ is on p. 1686, sec. 3.7 ‘Lambda coefficients’, not p. 1687; p. 1687 is sec. 4.1 Method (the simulation population model), unrelated. Both citations corrected in ch. 8.
primary-read · Independent literature read at a locator · cited in ch. 08 · read at p. 1679 (sec. 1); p. 1686 (sec. 3.7) · citation verification VERIF-ARK-2025 · read on 2026-08-06
Cheng et al. (2012). Comparison of reliability measures under factor analysis and item response theory
Abstract: ‘With increasing popularity of item response theory, a parallel reliability measure rho has been introduced using the information function’, and the article studies its relationship to the factor-analytic coefficients. Ch. 8 cites Cheng et al. only as the antecedent Andersson and Xin follow for the unit-variance form, which is the role the paper actually plays; the information-based rho is developed at its eqs. (12)-(14).
primary-read · Independent literature read at a locator · cited in ch. 08 · read at abstract, p. 52 · citation verification VERIF-CHENG-2012 · read on 2026-08-06
Cronbach (1951). Coefficient alpha and the internal structure of tests
Cronbach’s equation (2) on p. 299 is alpha = n/(n-1) (1 - sum_i V_i / V_t), with his own text ‘Here V_t is the variance of test scores, and V_i is the variance of item scores after weighting’ – identical in symbols and content to the book’s eq-alpha. The book’s provenance note claimed ‘the held copy is unpaginated OCR, so no page locator is given’; that is false, the held scan carries running page numbers throughout. Replaced with the real locator.
primary-read · Independent literature read at a locator · cited in ch. 08 · read at eq. (2), p. 299 · citation verification VERIF-CRONBACH-1951 · read on 2026-08-06
Cronbach and Shavelson (2004). My current thoughts on coefficient alpha and successor procedures
The book grouped this source with Sijtsma as making ‘the point at length’ about Nunnally’s threshold and its dropped qualifications. A full-text scan of all 29 pages finds ZERO occurrences of ‘Nunnally’, ‘.70’, ‘cutoff’, ‘cut-off’ or ‘threshold’; the four ‘rule of thumb’ mentions are about interpreting standard deviations and standard errors, not about an alpha criterion. The citation was unsupported as written. What the source does say, in its own abstract, is that alpha ‘covers only a small perspective of the range of measurement uses for which reliability information is needed and that it should be viewed within a much larger system of reliability analysis, generalizability theory’. Ch. 8 now says that, and says it is a different point from the threshold one.
primary-read · Independent literature read at a locator · cited in ch. 08 · read at abstract, p. 391 · citation verification VERIF-CRONBACH-2004 · read on 2026-08-06
Green et al. (1984). Technical guidelines for assessing computerized adaptive tests
Source defines sigma^2_em as the g-weighted average of sigma^2_e(theta) and marginal reliability as rho=(sigma^2_theta - sigma2_em)/sigma2_theta. That is exactly rho_Theta as the book uses it.
primary-read · Independent literature read at a locator · cited in ch. 08 · read at marginal reliability definition, Reliability section · citation verification VERIF-GREEN-1984 · read on 2026-08-06
Guttman (1945). A basis for analyzing test-retest reliability
Source: ‘The bound lambda_2 (sec. 13 below) is computed from the sum of squares of the covariances between items.’ Matches the book’s section number AND its description (squared inter-item covariances).
primary-read · Independent literature read at a locator · cited in ch. 08 · read at sec. 13 · citation verification VERIF-GUTTMAN-1945 · read on 2026-08-06
Hussey et al. (2025). An aberrant abundance of {Cronbach’s} alpha values at .70
The book said the excess was ‘just above .70’. The source locates it AT .70: the title is ‘An Aberrant Abundance of Cronbach’s Alpha Values at .70’, and the abstract reports ‘> 67,000 alpha values taken from > 60,000 measures’ showing ‘robust evidence of excesses at the alpha = .70 rule-of-thumb threshold that cannot be explained by justifiable measurement practices’; p. 10 likewise says ‘an excess at .70’. ‘Just above’ claimed a precision the source does not state. Corrected to the source’s own wording, and the corpus size added since it is what makes the signature convincing.
primary-read · Independent literature read at a locator · cited in ch. 08 · read at title and abstract; discussion p. 10 · citation verification VERIF-HUSSEY-2025 · read on 2026-08-06
Lee (2026) [lee_irtsimrel_2026]. {IRTsimrel}: Reliability-Targeted Simulation for Item Response Data
The package states the outer Jensen inequality and implements the two reliability functionals used by this programme. Authority boundary: Direct inspection of programme-owned package source establishes implementation and internal theorem provenance; it is not independent published corroboration.
programme-artifact-read · Programme-owned artifact; fidelity evidence, not independent corroboration · cited in ch. 09, 25, F · read at IRTsimrel v0.2.0: vignettes/theory-reliability.Rmd Theorem 1 and Corollary 1; R/reliability_utils.R compute_rho_bar and compute_rho_tilde; vignettes/theory-reliability.Rmd Theorem 1 gives only the outer inequality; the placement of rho_Theta between the two and the organized three-term proof are this book’s Jensen adaptation; vignettes/theory-reliability.Rmd Corollary 1; the non-monotonicity caveat is the package’s own wording and is preserved · programme receipt PAR-001; result THM-08-3, PROP-09-1 · read on 2026-08-04
Lee (2026) [lee_reliability-targeted_2026]. Reliability-targeted simulation of item response data: Solving the inverse design problem
The paper defines the inverse reliability-design problem and presents Empirical Quadrature Calibration and Stochastic Approximation Calibration. Authority boundary: A direct read of this programme-authored paper supports its stated algorithms and scope; it is not independent evidence that their empirical claims generalize.
programme-artifact-read · Programme-owned artifact; fidelity evidence, not independent corroboration · cited in ch. 09, 25, F · read at arXiv 2512.16012 v2; OCR sidecar lines 123-207 (Sections 3.2-3.3 and Algorithms 1-2) · programme receipt PAR-002 · read on 2026-08-07
Lee (2026) [lee_simulation_2026]. A simulation study of {Dirichlet} process mixture priors and goal-specific posterior summaries in {Bayesian} {IRT}
The theory book reads the companion’s frozen hypothesis, census, and lever records without recomputing its fits. Authority boundary: This is cross-book fidelity to a programme-owned companion, not independent replication or external authority.
programme-artifact-read · Programme-owned artifact; fidelity evidence, not independent corroboration · cited in ch. 09, 16, 24, 25, 26, 27 · read at DPMirt-simulation-study-v3/data/derived/book-facts.rds, cond.rds, and evidence.rds; pinned in manifest/external-source-ledger.csv · programme receipt PAR-005 · read on 2026-08-07
Sijtsma (2009). On the use, the misuse, and the very limited usefulness of {Cronbach’s} alpha
All three of ch. 8’s claims are the paper’s own. Bound not estimate: p. 107, ‘alpha is a lower bound to the reliability, in many cases, even a gross underestimate’, developed at p. 111 against the glb. A statement about the sum score: reliability is defined throughout for the total score X+ and its parallel form. Not a measure of unidimensionality: conclusion 4, p. 119, ‘Alpha is not a measure of internal consistency. Neither is it a measure of the degree of unidimensionality.’
primary-read · Independent literature read at a locator · cited in ch. 08 · read at pp. 107, 111; conclusion 4, p. 119 · citation verification VERIF-SIJTSMA-2009 · read on 2026-08-06
Wright and Masters (1982). Rating scale analysis
Source for PROP-08-2: Person separation index S = sigma/RMSEM and strata H = (4S+1)/3, under the three-measurement-error convention
primary-read · Independent literature read at a locator · cited in ch. 07, 08 · read at sec. 5.5; read directly at intake, C-002 and C-003 confirm the manuscript attribution · result PROP-08-2; correction C-001, C-002, C-003, C-004, C-005 · read on 2026-08-04
Zhang et al. (2025). Realistic simulation of item difficulties
Sec. 1 states verbatim that ‘test reliability increases as the variance in item difficulty distribution decreases (Gulliksen, 1945; Lord, 1952)’ – both the claim and the two attributed authors match ch. 9. Sec. 3.1 reports the tighter grouping of realistic difficulties relative to N(0,1) and U(-2,2), and gives the summary as a figure, exactly as ch. 9 notes. Sec. 4 treats each difficulty estimate as the mean of a normal with SD equal to its SE, generating a mixture – the construction ch. 9 describes.
primary-read · Independent literature read at a locator · cited in ch. 09 · read at secs. 1, 3.1, 4 · citation verification VERIF-ZHANG-2025 · read on 2026-08-06
D.2.4 Non-normal latent distributions
Blanca et al. (2013). Skewness and kurtosis in real data samples
Abstract (p. 78): 693 distributions, sample sizes 10 to 30; skewness ranged between -2.49 and 2.33; kurtosis between -1.92 and 7.41; ‘only 5.5% of distributions were close to expected values under normality’. Table 3 (p. 80) confirms the minima and maxima independently. Every figure ch. 13 quotes matches.
primary-read · Independent literature read at a locator · cited in ch. 13 · read at abstract p. 78; Table 3 p. 80 · citation verification VERIF-BLANCA-2013 · read on 2026-08-06
Cain et al. (2017). Univariate and multivariate skewness and kurtosis for measuring nonnormality: Prevalence, influence and estimation
Abstract (p. 1716): 1,567 univariate and 254 multivariate distributions ‘collected from authors of articles published in Psychological Science and the American Education Research Journal’, with ‘74 % of univariate distributions and 68 % multivariate distributions deviated from normality’. The p. 1722 summary repeats both percentages against the same counts. Counts, percentages and both journal names match ch. 13.
primary-read · Independent literature read at a locator · cited in ch. 13 · read at abstract p. 1716; results p. 1722 · citation verification VERIF-CAIN-2017 · read on 2026-08-06
Lee (2026) [lee_theta-nonnormality_2026]. How common are estimated latent-distribution departures from normality? {Evidence} from 504 item-response data sets
The manuscript reports the corrected 504-unit latent-shape analysis and treats between-calibration differences as sensitivity rather than truth recovery. Authority boundary: This programme manuscript and corrected frame are primary empirical inputs, not independent published corroboration of the theory book.
programme-artifact-read · Programme-owned artifact; fidelity evidence, not independent corroboration · cited in ch. 13, 14, 28 · read at manuscript-v1.9-arXiv-public/sections/02_method.tex and 04_results.tex; corrected 504-unit frame pinned as SRC-IRW-SHAPE504 · programme receipt PAR-008 · read on 2026-08-07
Li and Cai (2018). Summed score likelihood–based indices for testing latent variable distribution fit in item response theory
Section 1 supplies the two non-statistical motivations the book attributes to it, in the form the book gives: following Woods (2006), latent distributions that are positively skewed rather than normal. Locator ‘sec. 1’ as cited is correct. Supersedes this key’s earlier attribution-only row.
primary-read · Independent literature read at a locator · cited in ch. 13, 14 · read at sec. 1 · citation verification VERIF-LICAI-2018 · read on 2026-08-06
Micceri (1989). The unicorn, the normal curve, and other improbable creatures
Abstract (p. 156) states verbatim: an investigation of 440 large-sample achievement and psychometric measures ‘found all to be significantly nonnormal at the alpha .01 significance level’, and that several classes of contamination were found. Count, alpha level, universality and the contamination catalogue all match ch. 13. Micceri’s own description of the measures as ‘these discrete, bounded, measures’ is on pp. 156-157, as ch. 13 states.
primary-read · Independent literature read at a locator · cited in ch. 13 · read at abstract and intro, pp. 156-157 · citation verification VERIF-MICCERI-1989 · read on 2026-08-06
Paddock et al. (2006). Flexible distributions for triple-goal estimates in two-stage hierarchical models
triple-goal motivation, DP-1/SBR and DP-2 loss patterns, heterogeneous-variance rank comparison, and model-choice warning Paddock et al. combined flexible distributions with triple-goal estimation in two-stage hierarchical models before Lee et al., and their results do not uniformly favour flexible methods. Their threshold and tail-area summaries are adjacent precedent for fixed-cut distributional decisions, so this work does not claim to originate that target.
primary-read · Independent literature read at a locator · cited in ch. 13, 16, 17, 18, 19, 21, 27 · read at abstract; secs. 2.1, 4.2.2.1, 4.2.2.2, and 6; Whole article; two-stage setting; rank-loss results; threshold and tail summaries · receipt PV-SRC-PADDOCK, PVII-SRC-PADDOCK; result THM-17-3 · read on 2026-08-05
D.2.5 Orientation and assessment context
Abedi (2002). Standardized achievement tests and {English} language learners: Psychometrics issues
The book claimed a Grade 9 battery ‘reliable at .83 to .92 overall’. Table 9 (Site 2, Grade 9, Stanford 9) gives English-only subscale reliabilities of .835 vocabulary, .916 reading comprehension, .898 math, .803 mechanics, .823 expression, .805 science, .805 social science – a range of .80 to .92, not .83 to .92; the stated lower bound excluded four of the seven subscales. Abedi states no summary range in his prose (p. 247), so there was no alternative source for .83. Corrected to the table’s own figures, which also make the ELL contrast exact: .53 to .83 for ELL students, with the gap widening from reading comprehension (.92 against .83) to social science (.81 against .53).
primary-read · Independent literature read at a locator · cited in ch. 01 · read at Table 9, p. 248 · citation verification VERIF-ABEDI-2002 · read on 2026-08-06
Bock and Aitkin (1981). Marginal maximum likelihood estimation of item parameters: Application of an {EM} algorithm
Title: ‘Marginal Maximum Likelihood Estimation of Item Parameters: Application of an EM Algorithm’. Abstract: marginal maximum likelihood ‘becomes practical when computing procedures based on an EM algorithm are used’, and ‘by characterizing the ability distribution empirically, arbitrary assumptions about its form are avoided’. Ch. 5’s claim that this algorithm is what made MML practical is the abstract’s own word.
primary-read · Independent literature read at a locator · cited in ch. 02, 05 · read at abstract, p. 443 · citation verification VERIF-BOCKAITKIN-1981 · read on 2026-08-06
Conoyer et al. (2022). Meta-analysis of validity and review of alternate form reliability and slope for curriculum-based measurement in science and social studies
Grounds correction C-007: The earlier correction was itself wrong because it mixed rows and units. Table 1 contains 20 assessment-level N values: 25, 51, 51, 58, 58, 63, 106, 106, 117, 146, 148, 153, 198, 202, 205, 367, 547, 746, 799, and 967. R type-7 quartiles are 61.75, 147, and 245.5, so the manuscript’s rounded 62/147/245 and range 25–967 are supported. The value 1,545 is the combined sample for the Ford and Hosp study, whereas Table 1 reports separate grade-level assessment samples of 799 and 746. WITHDRAW the 2026-08-03 correction; retain the manuscript claim with a precise Table 1 locator.
primary-read · Independent literature read at a locator · cited in ch. 01 · read at — · correction C-007 · read on 2026-08-04
Debelak et al. (2022). An Introduction to the Rasch Model with Examples in R
Debelak et al. define theta_p as the ability of person p, beta_i as the difficulty of item i, and write Pr(U_pi = 1 | theta_p, beta_i) = exp(theta_p - beta_i)/(1 + exp(theta_p - beta_i)). Their notation is this books notation, which is why this book adopted it.
primary-read · Independent literature read at a locator · cited in ch. 02, 03, 05, 07 · read at ch. 2, eq. (2.1) and the paragraph defining theta_p and beta_i; secs. 2.4.1 and 2.4.3; direct attribution to Debelak et al. (2022), not Andersen (1970) · receipt PVII-SRC-DEBELAK; result THM-03-1 · read on 2026-08-03
Fox (2010). Bayesian Item Response Modeling: Theory and Applications
Fox indexes persons by i and items by k: the response is Y_ik, the ability theta_i, the item parameters are collected in xi, and the population hyperparameters in theta_P. Both indices are therefore reversed relative to this books p for persons and i for items. His sec. 2.1 shrinkage discussion is the same statement as ch. 11s with i where this book writes p.
primary-read · Independent literature read at a locator · cited in ch. 02, 10, B · read at ch. 2, eq. (2.1) and secs. 2.1-2.2; ch. 2 sec. 2.1 eqs. 2.1-2.2 give the person/item conditional-independence factorizations; adapted here to the four-level centered-item model · receipt PVII-SRC-FOX; result PROP-10-1 · read on 2026-08-04
Lee (2026) [lee_reliability-distribution_2026]. How reliable are psychological measurements? {The} distribution of marginal reliability across 889 item-response datasets
The manuscript defines the full-corpus reliability estimands; the theory book keeps the frozen 889-unit analysis target distinct from the 879-row public release. Authority boundary: This programme manuscript and its data products are primary empirical inputs, not independent published corroboration of the theory book.
programme-artifact-read · Programme-owned artifact; fidelity evidence, not independent corroboration · cited in ch. 01, 08, 25, 28 · read at manuscript-v3-arXiv-public/sections/body.tex; frozen 889-row case snapshot and public 879-row release pinned as SRC-IRW-REL-FULL889 and SRC-IRW-REL-PUBLIC879 · programme receipt PAR-007 · read on 2026-08-07
Lee et al. (2025). Improving the estimation of site-specific effects and their distribution in multisite trials
The first stage is a Gaussian sampling approximation that plugs in an estimated se-hat-squared and then treats it as known. Equations 6-7 define MSELR and its percentile-scaled MSELP counterpart. Informativeness I is an average reliability determined by n-bar and sigma; it spans .01-.71 with reported mean .25 and equals .04-.06 in the application. DP-versus-Gaussian ISEL crossovers are conditional on both I and J: about I=.20 at J=300, about .40 at J=75-100, and no significant DP advantage at J=25 even at I=.71. PM is best for RMSE, CB/GR are better for ISEL, MSELP is essentially unaffected by model and summary choice, the three goals are attributed to Shen and Louis, DP-inform uses Lee et al.s chi-squared K-matched alpha elicitation, and zero correlation between tau_j and se-hat-squared is a named limitation. For the crosswalk: tau_j is the site effect, hat-tau_j its ML estimate, se-hat-squared_j the estimated first-stage variance that is plugged in and then conditioned on, G the prior on tau_j, sigma^2 the cross-site variance, S_j the shrinkage factor, V_j the posterior variance, J the number of sites, and I the informativeness – an average reliability in [0,1], colliding with this books I for the number of items. Note 2 records that they call G a prior to match Bayesian hierarchical practice while a frequentist frame reads the same object as a distributional assumption.
primary-read · Independent literature read at a locator · cited in ch. 01, 15, 17, 21, B · read at pp. 731-764 read end to end; eqs. (1)-(7) and (9)-(12); Table 1 p. 754; Figures 1-8; ‘A Simulation Study’ pp. 741-744; ‘Results: Effects of Data-Generating Factors’ pp. 744-747; ‘Results: Case Studies’ pp. 747-752; ‘Results: Real-Data Example’ pp. 752-756; Discussion pp. 756-758; notes 1-2 p. 759; eqs. (1)-(4) and (9)-(12); note 2 · receipt PVII-SRC-LEE, PVII-SRC-LEE-B · read on 2026-08-05
Nelson et al. (2023). Review of curriculum-based measurement in mathematics: An update and extension of the literature
Title: ‘Review of curriculum-based measurement in mathematics: An update and extension of the literature’, Journal of School Psychology 97 (2023) 1-42, matching the bibliography. The abstract states it updates and extends the Foegen et al. (2007) progress-monitoring review across 99 studies. Ch. 1 cites it only for the curriculum-based measurement context, which is what it is.
primary-read · Independent literature read at a locator · cited in ch. 01 · read at title and abstract, p. 1 · citation verification VERIF-NELSON-2023 · read on 2026-08-06
Rutkowski and Rutkowski (2013). Measuring socioeconomic background in {PISA}: One size might not fit all
The table reports reliabilities for three PISA socioeconomic subscales (Wealth, Cultural Possessions, Home Educational Resources) by educational system. The lowest value in the table is .41 (Estonia, Home Educational Resources); all 23 parsed rows and 69 reliability entries were checked, and no value falls below it. Ch. 1’s ‘as low as .41 … in some groups’ is exact, and .41 is genuinely the minimum rather than merely an instance.
primary-read · Independent literature read at a locator · cited in ch. 01 · read at reliability table, p. 270 · citation verification VERIF-RUTKOWSKI-2013 · read on 2026-08-06
Shen and Louis (1998). Triple-goal estimates in two-stage hierarchical models
expected-rank action, midpoint-quantile ISEL solution, GR construction and assignment, stochastic-order rank result, CB rank invariance, and sorted regret The three inferential goals and triple-goal estimator are Shen and Louiss, and Lee et al. attribute them there. GR orders scalar units and reads midpoint quantiles from an estimated EDF. On p. 468 the authors explicitly say all approaches generalize to multivariate unit-specific parameters and point to Ghosh for multivariate CB. This supports multivariate use but does not supply a canonical joint transformation-aware ordering or EDF action. Shen and Louis write the posterior variance of unit k as lambda_k, which under TD-3 collides with this books lambda_i for discrimination on most pages of Parts III and VI. This is the collision that forced this books single notation departure, resolved as v_p.
primary-read · Independent literature read at a locator · cited in ch. 01, 12, 17, 18, 19, 21, 22, B · read at secs. 2.1, 2.3, 2.4, 3, and 4.2; Theorems 1-2; eq. 14; pp. 457-462; secs. 2.1, 2.3, 2.4 and 3; Theorems 1-2; eq. (3); p. 468 multivariate statement; sec. 2 notation; sec. 2, read directly; the decomposition including the irreducible posterior-variance term is theirs; source lambda_p is translated to registered v_p under TD-3; sec. 2.1 p. 457 gives the expected-rank action and its generally noninteger values; sec. 3 Theorem 2 p. 459 gives induced-order agreement under stochastic ordering; Goldstein & Spiegelhalter (1996) supplies the rank-uncertainty context; secs. 2.1-2.4 supply the three Bayes actions and Theorem 1’s midpoint-quantile ISEL solution; Paddock et al. abstract/sec. 2.1 supplies the general incompatibility motivation; the narrower existence statement and correctly computed normal example are this book’s and checked against a frozen table in V8; sec. 3 p. 459 states rank invariance; the standardized-moment and shape statements are the book’s elementary positive-affine extension; sec. 2.1 p. 457 gives expected ranks; sec. 2.3 pp. 457-458 gives ISEL and midpoint quantiles; sec. 2.4 pp. 458-459 gives the three-step construction and optimal assignment; uniform randomization within exact tie blocks is the book’s exchangeability-preserving completion; sec. 2.4 p. 459; the permutation argument is exact, while the explicit tie-block completion and verification are supplied here; sec. 4.2 eq. (14) p. 462 gives the sorted form after adopting posterior-mean ranks, which the source says agree with optimal ranks for a wide class; the book states the actual-permutation identity first and checks both cases in V8 · receipt PV-SRC-SHEN, PVII-SRC-SHEN-VII, PVII-SRC-SHEN-B; result PROP-17-1, PROP-17-2, THM-17-3, PROP-18-2, THM-19-1, PROP-19-2, PROP-19-3 · read on 2026-08-05
D.2.6 Positioning and scope
Gu and Xu (2019). The sufficient and necessary condition for the identifiability and estimability of the {DINA} model
The sufficient AND necessary condition for identifying all DINA parameters is a condition on the design alone: the Q-matrix complete, each of the K attributes required by at least three items, and any two columns of the submatrix Q* distinct. When it holds the MLEs are consistent as N grows.
primary-read · Independent literature read at a locator · cited in ch. 22 · read at Conditions 1-2 pp. 4-5; Theorem 1; Corollary 1 and its surrounding discussion · receipt PVII-SRC-GUXU19 · read on 2026-08-05
Gu and Xu (2020). Partial identifiability of restricted latent class models
With known item parameters Proposition 3.2 identifies sums of p over profile equivalence classes defined by identical Gamma columns. With unknown item parameters inseparability alone is insufficient. The structural conditions C1-C2 in Theorem 3.1 and the adjusted-Gamma route with the conditions in Proposition 3.3 are sufficient routes to p-partial identifiability; this receipt does not assert that either route is generally necessary. The comparison between these finite equivalence-class masses and San Martin et al.s continuous integral functionals is this books cross-source synthesis and is not a statement attributed to either source.
primary-read · Independent literature read at a locator · cited in ch. 22 · read at sec. 2; Proposition 3.2; Definition 3.2; Theorem 3.1 conditions C1-C2; Proposition 3.3; Gamma-column equivalence relation · receipt PVII-SRC-GUXU20 · read on 2026-08-05
Masters (1982). A {Rasch} model for partial credit scoring
Three claims. (a) In the Rasch dichotomous model the number of successes is sufficient for the person parameter. (b) The PCM preserves sufficiency: conditioning on the total count of completed steps removes beta_n, so CML estimates step difficulties from calibration responses without specifying G. (c) In Masters’ PCM-versus-GRM comparison the graded-response form prevents algebraic separation even with no discrimination parameter; this local obstruction is the category-boundary construction rather than a slope. Masters’ characterization of Samejima was checked separately against Samejima (1969).
primary-read · Independent literature read at a locator · cited in ch. 22 · read at sec. 3 p. 152; sec. 4 p. 155; sec. 5 eq. (10) p. 158 and eqs. (11)-(14) pp. 159-160; summary p. 172 · receipt PVII-SRC-MASTERS · read on 2026-08-05
Muraki (1992). A generalized partial credit model: Application of an {EM} algorithm
The GPCM adds a slope a_j to the PCM and leaves the Rasch family, whose separability and minimal sufficient statistics permit CML. Muraki differentiates category probability with respect to theta: dP_jk/dtheta = a_j P_jk [k - sum_c c P_jc]. Thus a_j multiplies the theta-gradient; this is not a derivative with respect to the slope. Muraki specifies a normal population density in the marginal-likelihood model and uses Gauss-Hermite quadrature to compute its integrals: normality is a modelling assumption and quadrature is the computational method.
primary-read · Independent literature read at a locator · cited in ch. 22 · read at p. 160 (Rasch separability and minimal sufficient statistics); eq. (13) p. 163; eqs. (17), (22)-(23) pp. 165-167 · receipt PVII-SRC-MURAKI · read on 2026-08-05
Samejima (1969). Estimation of latent ability using a response pattern of graded scores
Samejima constructs graded-response category probabilities from cumulative category-boundary response functions. This independently supports Masters characterization of the GRM construction; it does not by itself establish a universal boundary for every ordered-response model.
primary-read · Independent literature read at a locator · cited in ch. 22 · read at graded-response construction and category-boundary operating characteristics · receipt PVII-SRC-SAMEJIMA · read on 2026-08-05
Xu et al. (2025). Robustness of identifying item–trait relationships under non-normality in {MIRT} models
Across EIFA, EM-L1 and EMS, generated skewness and excess kurtosis generally reduce item-trait F1 recovery and increase parameter MSE, with method- and condition-specific exceptions including cells where EIFA benefits. The discussion proposes semi-nonparametric, skew-normal and skew-t latent distributions as future responses. The paper supports a robustness warning, not a literature-wide claim that flexible priors are absent.
primary-read · Independent literature read at a locator · cited in ch. 22 · read at Abstract; results; discussion, Mathematics 13(23), 3858 · receipt PVII-SRC-XU25 · read on 2026-08-05
D.2.7 Posterior summaries and inferential goals
Ghosh (1992). Constrained {Bayes} estimation with applications
general constrained-Bayes action, inflation conditions, and nonexistence boundary Constrained Bayes is Ghoshs and its scalar moment-matching action transfers to this setting. The introduction explicitly motivates CB by identifying parameters above or below a cutoff and by category classification, so fixed-cut decisions are precedent rather than a novelty of this work. The source also treats multivariate CB, while an unordered-profile action is not supplied.
primary-read · Independent literature read at a locator · cited in ch. 18, 21 · read at Theorem 1 and proof pp. 534-535; Remarks 1-2; Introduction; Theorem 1 and proof pp. 534-535; Remarks 1-2; Theorem 1 and proof pp. 534-535 give the general rule; Remarks 1-2 give inflation and nonexistence conditions; V8 checks the independent-posterior form, its exact equality to the package expression through R’s P-1 sample-variance denominator, and the omitted cross-person covariance term · receipt PV-SRC-GHOSH, PVII-SRC-GHOSH; result THM-18-1 · read on 2026-08-05
Lee (2026) [lee_casestudy_2026]. Case studies of {Dirichlet} process mixture priors and goal-specific posterior summaries in {Bayesian} {IRT}: Thirteen real tests from the {Item} {Response} {Warehouse}
The cited second edition supplies the consequence vocabulary and correction status while the frozen first-edition analysis layer supplies the registered numbers. Authority boundary: This is cross-book fidelity to programme-owned companions; it does not convert exploratory case evidence into confirmatory or independent evidence.
programme-artifact-read · Programme-owned artifact; fidelity evidence, not independent corroboration · cited in ch. 19, 20, 24, 25, 26, 28, 29, F · read at DPMirt-case-study-v2 book for narrative conclusions; DPMirt-case-study first-edition frozen P-series tables and replicate audit for numbers; pinned in manifest/external-source-ledger.csv; replicate-audit record and its widened claim C-11: 9 of 24 fits at 12-16 items exceed 0.10 SD self-reproduction; 0 of 54 at 17 or more; 8 of the 9 failures have max R-hat below 1.05; table P3-T14 (selection-noise within method): same-method different-seed top-decile churn 6.4516% bimodal / 4.0% normal / 2.0% skewed / 8.6957% undetermined; table P3-T15 and correction C-06 locate prior-family and elicitation contrasts against that floor · programme receipt PAR-006; receipt PV-SRC-CASE-GR, PV-SRC-CASE-SEED · read on 2026-08-07
Lee (2026) [lee_dpmirt_2026]. {DPMirt}: {Bayesian} semiparametric item response theory models using {Dirichlet} process mixture priors
The source implements PM, CB, and GR summaries and exposes the default and constrained item-model branches read by Chapters 18 and 26. Authority boundary: Direct inspection of programme-owned code supports what the code does; it does not independently validate the method or the book’s interpretation.
programme-artifact-read · Programme-owned artifact; fidelity evidence, not independent corroboration · cited in ch. 18, 26, 29, F · read at DPMirt v0.2.0: R/estimates.R .triple_goal() lines 264-296; R/model_spec.R model and constraint branches; R/estimates.R, .triple_goal() lines 264-296: theta_pm <- colMeans(s) at line 272; lambda_k <- theta_psd^2 at line 281; var_pm <- var(theta_pm) at line 284; cb_factor <- sqrt(1 + mean(lambda_k) / var_pm) at line 296; base R var uses divisor P-1 · programme receipt PAR-003; receipt PV-SRC-DPMIRT-CB · read on 2026-08-07
Lin et al. (2006). Loss function based ranking in two-stage, hierarchical models
bibliographic lineage for the Louis-Ghosh-Paddock comparison and loss-based ranking context
primary-read · Independent literature read at a locator · cited in ch. 18 · read at reference list and loss definitions in secs. 1-2 · receipt PV-SRC-LIN · read on 2026-08-05
Lockwood et al. (2002). Uncertainty in rank estimation: Implications for value-added modeling accountability systems
Gaussian stability-ratio statement and subject-specific empirical stability values Lockwood, Louis and McCaffrey give the Gaussian calibration relating rank recovery to a stability coefficient used in ch. 20 and discuss classification/reporting decisions for extreme percentiles. This is the correct source for the rank-stability point previously misattributed to Lockwood et al. (2018).
primary-read · Independent literature read at a locator · cited in ch. 20, 21 · read at secs. 2, 5, and 6; sec. 2 rank-stability calibration; secs. 5-6 classification and reporting context; sec. 2 gives the Gaussian statement under its model conditions; the Rasch calculation uses full-support midpoint-quantile quadrature, common-wbar interpolation, a frozen 400-versus-800-node convergence receipt, and a unit-variance logistic heavy-tail working shape in code/R/02-rank-concordance.R and tables/F-rank-reliability-.rds · receipt PV-SRC-LOCKWOOD02, PVII-SRC-LOCKWOOD02; result PROP-20-1 · read on 2026-08-05*
Lockwood et al. (2018). Flexible {Bayesian} models for inferences from coarsened, group-level achievement data
stated aim and scope of flexible Bayesian inference from coarsened group achievement data Lockwood et al. fit a bivariate group-parameter model and apply CB and TG coordinatewise. The paper is a multivariate-use precedent; it is not the source of ch. 20s Gaussian rank-stability calibration.
primary-read · Independent literature read at a locator · cited in ch. 20, 21, 22 · read at abstract and sec. 1; Bivariate FH-HETOP model and coordinatewise CB/TG, pp. 666-672 · receipt PV-SRC-LOCKWOOD18, PVII-SRC-LOCKWOOD · read on 2026-08-05
D.2.8 Rasch, the 2PL, and estimation
Baker and Kim (2004). Item Response Theory: Parameter Estimation Techniques
Baker and Kim write the two-parameter normal ogive as Z_ij = alpha_i(theta_j - beta_i), so alpha_i is the item discrimination, beta_i the difficulty, theta_j the trait with j indexing persons and i items, and zeta_i = -alpha_i beta_i the slope-intercept form. They also state that Z_ij = alpha_i theta_j - gamma_i is an atypical parameterization of item difficulty.
primary-read · Independent literature read at a locator · cited in ch. 03, 06, 07, 24, B · read at ch. 12, eqs. (12.6) and (12.7) and the surrounding text · receipt PVII-SRC-BAKERKIM · read on 2026-08-05
Haberman (2016). Models with nuisance and incidental parameters
Source for THM-05-1: JML is inconsistent for fixed I as P grows; in the two-item Rasch case the estimated difficulty difference converges a.s. to twice its true value. Consistent if I also grows with log(I)/P -> 0
primary-read · Independent literature read at a locator · cited in ch. 05 · read at sec. 9.2.1 (pp. ~152-154), reporting Andersen (1973, pp. 66-69) · result THM-05-1 · read on 2026-08-03
Lord (1983). Unbiased estimators of ability parameters, of their variance, and of their parallel-forms reliability
Eq. (20), p. 236, gives I = sum_i P_i’^2/(P_i Q_i) – the information the book calls script-J. Sec. 1.4, eqs. (27)-(29) on p. 237, give the bias as B_1(theta-hat) = I^-2 sum_i A_i I_i (phi_i - 1/2) with phi_i = (P_i - c_i)/(1 - c_i), and Lord states there that B_1 is of order n^-1 because I is of order n, which is the book’s order claim in his notation. Lord’s display is the THREE-PARAMETER form; the guessing-free -J/(2 script-J^2) the book displays is Warm’s rendering. Setting c_i = 0 with P_i’ = A_i P_i Q_i and P_i’’ = A_i^2 P_i Q_i (Q_i - P_i) reduces Lord’s sum to -J/2, so the two agree; the equivalence was checked by hand. Ch. 6’s provenance section now states this, so a reader opening Lord is not surprised by a formula that is not on his page.
primary-read · Independent literature read at a locator · cited in ch. 06 · read at eqs. (20), (27)-(29), pp. 236-237 · citation verification VERIF-LORD-1983 · read on 2026-08-06
San Mart{'i}n (2016). Identification of item response theory models
Ch. 22s 2PL consolidation: the semiparametric 2PL identification problem is open – marked with question marks in Table 8.1 and stated as open in sec. 8.7 – so every identification statement this book proves about G is proved for the Rasch case. Also the basis of C-009: the scale identification of sec. 8.5.1 holds only when the shape of G is known and sigma alone is unknown.
primary-read · Independent literature read at a locator · cited in ch. 04, 16, 22, 26 · read at sec. 8.5.1 and Table 8.1; sec. 8.7; sec. 8.2.2 (p. 131); secs. 8.4-8.7 and Table 8.1 for the Rasch/2PL scale distinction; Appendix F sec. F.2.3 and inline push-forward extend it to G · receipt PVII-SRC-SANMARTIN16; result DEF-04-1, PROP-16-2; correction C-009, C-010 · read on 2026-08-03
San Mart{'i}n et al. (2011). On the {Bayesian} nonparametric generalization of {IRT}-type models
A finite Rasch test of I items identifies I+1 integral functionals of G (Theorem 5 in the anchor-item form), not G as an unrestricted distribution. Theorem 3 is the Rasch Poisson-count model and is not this result; see C-014. In ch. 22 this supports only the bounded lesson that a finite design identifies selected aspects of a latent distribution. These integral functionals are not equivalence-class masses or a partition of the real line and are not structurally the same object as Gu and Xu’s grouped proportions.
primary-read · Independent literature read at a locator · cited in ch. 04, 16, 21, 22, 26 · read at Theorems 3, 4, 5 and 6 and expression (11); sec. 1 identification problem, adapted with the explicit shift orbit from Appendix F sec. F.2.2; sec. 1 supports the ordinary-base-centering statement; the chapter explicitly does not generalize it to every constrained or transformed nonparametric prior; Theorem 5 and expression (11), not Theorem 3 as the Appendix F draft states (see C-014); Theorem 6 gives the asymptotic conditions; no equivalence with a selected pair’s total-variation separation is claimed · receipt PVII-SRC-SANMARTIN; result PROP-16-1, PROP-16-3, PROP-16-4 · read on 2026-08-04
Warm (1989). Weighted likelihood estimation of ability in item response theory
Source for THM-06-1: With known item parameters and Warm’s boundedness, smoothness, and replicated-item asymptotic conditions, the WLE with d log w/d theta = J/(2 J-info) has bias o(I^-1), is asymptotically normal, and has the same first-order asymptotic variance as ML.
primary-read · Independent literature read at a locator · cited in ch. 06 · read at eqs. (6)-(10) and theorem, pp. 430-432; Appendix assumptions and separate variance proof, pp. 444-448 · result THM-06-1 · read on 2026-08-04
D.3 What this appendix exposes
Generating rather than writing it surfaced two things worth stating plainly.
0 locator-read annotations are unbound (none). This is an executable release condition, not an aspiration: the generator stops if a locator-read entry lacks a receipt. The inherited six were resolved without inventing evidence. Lo (1984) now has a bounded Chapter 15 narrative receipt. Five programme-owned works that already had direct artifact reads now carry programme receipts, and all eight programme-owned reads are classified under the separate programme-artifact-read tier.
Conoyer’s correction binding remains explicit as well: corrections join through the bibliography’s verified_claims, rather than by searching a prose locator for a BibTeX key.
The annotations are uneven in length, and the unevenness is informative. Sources with long entries are those whose receipts record several distinct verified claims — Lee et al. (2025), Masters (1982), Gu and Xu (2020). Sources with a one-line entry supply a single result. The length of an annotation therefore measures how much of this book rests on that source, which is more useful than a uniform paragraph would be.
D.4 What is deliberately not here
No evaluations. This appendix does not say whether a source is good, influential, or superseded. It says what was read in it and where. A judgement about a literature would need a survey this book has not conducted, and Chapter 23 records where that absence matters.
No entries for the 292 held-but-unread works. Listing them by title would fill pages and make none of them checkable. They are in refs/references.bib with their tier, and code/R/03-citation-check.R refuses any numbered result that cites one as though it had been read. The same checker validates every programme receipt, narrative receipt, date, bibliography key, and verified_claims binding.
No reading plan. Which of the held works should be read next is a question about the next piece of research, not about this book. The live, owner-assigned acquisition debt is shown separately in Appendix E; an entry there is a request to acquire or disambiguate a work, not evidence that it has been read.