2  Notation, Terminology, and Two Frameworks

A reader who opens this book at chapter 8 and finds \(\bar w\) needs to know what it means, and a reader who moves between this book and the manuscript it stands behind needs to know where the two differ. This chapter fixes both. It also settles a vocabulary dispute that turns out not to be about vocabulary: whether the assumption \(\theta_p \sim G\) is a prior or a population distribution, and what follows from the answer.

2.1 The data

A test of \(I\) items is administered to \(P\) persons. Write \(p = 1,\dots,P\) for persons and \(i = 1,\dots,I\) for items. The response of person \(p\) to item \(i\) is a random variable \(U_{pi} \in \{0,1\}\), with \(1\) denoting a correct response or an endorsement; its realization is \(u_{pi}\). Capitals denote random variables and lowercase their realizations, and the distinction is held throughout: it is the difference between \(\operatorname{Var}(U_{pi})\), a property of the model, and the observed \(u_{pi}\), a number.

Two collections of responses matter, and they are not interchangeable. The within-person response vector

\[ \mathbf{u}_p = (u_{p1}, \dots, u_{pI})^\top \in \{0,1\}^I \]

collects one person’s responses across items, and it is the object every ability estimator in this book takes as input. The item response vector \(\mathbf{u}^i = (u_{1i},\dots,u_{Pi})^\top\) collects one item’s responses across persons, and it is what item calibration consumes. Their margins are the person total score \(r_p = \sum_i u_{pi}\) and the item total score \(c_i = \sum_p u_{pi}\). Whether \(r_p\) carries all the information about \(\theta_p\) that the data contain is not a notational question but a modelling one, and Chapter 3 shows that the answer separates the Rasch model from the 2PL.

What the submitted manuscript says. The manuscript writes \(\mathbf{u}_p\) in its central equations without ever defining it, which Reviewer 2 flagged. It also writes responses as \(y_{ip}\), with the item index first. This book uses \(u_{pi}\), person index first, matching the person-major layout of a response matrix and the convention of the appendix drafts that seed Chapter 3 through Chapter 11.

2.2 Notation

Table 2.1 is the master table. It is generated from manifest/notation-register.csv, which the build lints every chapter against: a symbol that appears in a numbered equation but not in the register fails check V3.

Table 2.1: Master notation. Source: tables/T-notation.rds.
Symbol Meaning Note
\(P\) number of persons actually analyzed (index bound) P = N in a realized design cell with no exclusions or missing persons
\(N\) planned simulation-design sample size a design factor; distinguished from the realized analytic index bound P
\(I\) number of items (index bound)
\(p\) person index p = 1..P
\(i\) item index i = 1..I
\(U_{pi}\) response random variable capital = random, lowercase = realized, held throughout
\(u_{pi}\) realized response
\(y_{ip}\) legacy manuscript response notation used only in the notation crosswalk; replaced by u_{pi} with person-first indexing
\(\mathbf{U}\) response matrix (random) P x I; realized as
\(\mathbf{u}\) realized response matrix the full data; distinct from the within-person vector _p
\(\mathbf{u}_p\) within-person response vector DEFINED in ch02 — R2’s second bullet; the manuscript uses it undefined
\(\mathbf{u}^i\) item response vector across persons
\(r_p\) person total score sufficient for theta_p under Rasch only (PROP-03-4)
\(c_i\) item total score
\(\mathbf{r}\) vector of person total scores the statistic CML conditions on
\(\theta_p\) person latent trait
\(\beta_i\) item difficulty
\(\lambda_i\) item discrimination TD-3: present from ch02. Rasch is lambda_i == 1. See v_p
\(\pi_{pi}\) probability of a correct response
\(P_i(\theta)\) item response probability as a function of ability local notation in ch06; equals Pr(U_{pi}=1 | theta) with person index suppressed
\(Q_i(\theta)\) complementary item response probability Q_i(theta) = 1 - P_i(theta)
\(P_i'(\theta)\) first derivative of the item response probability prime denotes differentiation with respect to theta; P_i’’ is the second derivative
\(J(\theta)\) Warm bias numerator plain J is not information; J(theta) = sum_i P_i’ P_i’’/(P_i Q_i) in ch06
\(\mathcal{I}\) retired information symbol appears only when documenting the manuscript collision; this book uses mathcal J
\(L_C\) Rasch conditional likelihood conditions on person total scores
\(L_M\) marginal likelihood integrates person traits against G
\(T_{\boldsymbol\lambda}\) known-discrimination weighted score a sufficient statistic only when the discrimination vector is fixed and known
\(X_p\) standardized latent variable known-shape location-scale representation theta_p = sigma X_p
\(S_p(\theta\,\boldsymbol\psi)\) person scoring estimating function zero defines a regular plug-in ability estimator
\(\boldsymbol\psi\) vector of calibrated item parameters contains beta under Rasch and beta plus lambda under 2PL; distinct from item hyperparameters phi
\(V_\psi\) covariance matrix of calibrated item estimates from the calibration design; independent-calibration formula in PROP-06-2
\(\mathbf{g}_\psi\) implicit-function sensitivity of ability to item parameters = -(partial S/partial theta)^(-1) partial S/partial psi
\(g_{\beta_j}\) Rasch component of plug-in sensitivity = P_j Q_j / mathcal J for the Rasch ML score
\(\mathcal{J}_{2\mathrm{PL}}(\theta)\) 2PL test information local specialization = sum_i lambda_i^2 P_i Q_i
\(J_{2\mathrm{PL}}(\theta)\) 2PL Warm bias numerator local specialization = sum_i lambda_i^3 P_i Q_i (1-2P_i)
\(\mathcal{J}(\theta)\) test information ONE symbol only; general form sum_i lambda_i^2 pi(1-pi)
\(\mathrm{se}(\hat\theta_p)^2\) conditional sampling variance of the ability estimate never called ‘the variance of theta_p’ — ch11 sec6 explains why
\(\mathrm{MSEM}\) mean square measurement error = P^{-1} sum_p se(hat theta_p)^2; W&M write MSE_P
\(\mathrm{RMSEM}\) root mean square measurement error W&M write SE_P = (MSE_P)^{1/2}
\(G\) latent-trait distribution distinct from G_N and bar G_N — the manuscript conflates them
\(G_N\) empirical distribution of the realized theta
\(\bar G_N\) posterior estimate of G_N a probability in [0,1] evaluated at t — NOT a trait estimate. R1’s confusion
\(\boldsymbol\delta\) person hyperparameters (mu_theta sigma^2_theta) ‘=’ not ‘in’ — R2’s sixth bullet
\(\boldsymbol\phi\) item hyperparameters (mu_beta sigma^2_beta)
\(\mu_\theta\) person population mean
\(\sigma^2_\theta\) person population variance a design parameter in the simulation; an estimator (SA^2 = SD^2 - MSE) in Wright & Masters — C-004
\(w_p\) person shrinkage weight
\(\bar w\) Rasch test reliability = sigma2/(sigma2+MSEM) = W&M’s R_P. NOT equal to mean(w_p) — PROP-08-1
\(S\) person separation index = sigma_theta / RMSEM
\(H\) number of person strata = (4S+1)/3, from strata three measurement errors apart
\(\eta_p\) posterior mean of theta_p
\(v_p\) posterior variance of theta_p DEPARTURE from Shen & Louis, forced by TD-3. Stated at first use in ch11, repeated at the Part VI first use in ch17 and in ch18, and recorded in the Appendix B crosswalk
\(\alpha\) DP concentration parameter NEVER ‘precision’ — R2-4a. The manuscript uses ‘precision’ throughout
\(F\) the DP-distributed mixing measure
\(G_0\) DP base measure
\(K_J\) number of occupied clusters among J draws
\(M\) finite truncation level
\(\theta^{CB}_p\) constrained Bayes estimate
\(\theta^{GR}_p\) triple-goal estimate
\(R_p\) true rank of theta_p
\(\hat R_p\) integer-valued estimated rank
\(\mathbf{I}_I\) identity matrix of order I the subscript is the test length; bold upright distinguishes it from the italic scalar I, and it appears only in PROP-10-2
\(\mathbf{C}\) centering matrix I_I - I^{-1} 11’ symmetric idempotent of rank I-1; used only to derive the induced item prior in PROP-10-2
\(\mathcal{X}\) sample space on which the Dirichlet process lives Ferguson writes the pair (X, A); the latent-trait application takes X = R
\(\mathcal{A}\) sigma-field of subsets of the sample space the partitions in DEF-15-1 are measurable with respect to it
\(p_n\) nth stick-breaking weight = theta_n prod_{m<n}(1 - theta_m); sums to one a.s.
\(Y_n\) nth stick-breaking atom location iid from the base measure G_0, independent of the sticks
\(W_i\) indicator that the ith draw is a new distinct value P(W_i = 1) = alpha/(alpha + i - 1)
\(Z_n\) number of distinct values among n draws Antoniak’s symbol; this book writes K_J for the same quantity when J is the number of persons
\(z_p\) cluster label of person p under the CRP representation the DPM’s mixture-component assignment; theta_p | z_p ~ N(muTilde_{z_p}, s2Tilde_{z_p})
\(\mathbf{z}\) vector of cluster labels over persons z ~ CRP(alpha, P); its number of distinct values is K
\(\Theta\) generic latent-trait random variable uppercase form is used when the trait is explicitly random
\(\rho\) generic reliability value or target subscripts distinguish classical and population reliability functionals
\(\varepsilon_p\) person-level measurement error appears in the working decomposition hat(theta)_p = theta_p + epsilon_p
\(\chi^2\) chi-squared distribution used only in a prior-density comparison
\(\epsilon\) positive limiting or boundary constant local analytic placeholder rather than a model parameter
\(\tau_L\) working-likelihood precision inverse conditional measurement-error variance
\(\tau_\pi\) prior precision inverse latent population variance in the Gaussian working model
\(\tau^2\) canonical variance in the compound-decision formulation local translation of the source notation
\(\xi\) canonical mean vector in the James–Stein formulation used only in the compound-decision comparison
\(\varphi_1\) James–Stein estimator canonical shrinkage rule in the source parameterization
\(\nu_1\) shape hyperparameter of the inverse-gamma base component package-facing name retained in the model specification
\(\nu_2\) scale hyperparameter of the inverse-gamma base component package-facing name retained in the model specification
\(\xi_G(k)\) identified Laplace-transform functional of G indexed by k = 0 through I
\(\mathbf{a}\) action vector in a decision problem the quantity a loss is minimized over; distinct from any estimator until a loss is named
\(a_p\) pth coordinate of the action vector the number reported for person p
\(\mathbf{R}\) vector of person ranks R_p = sum_q 1{theta_q <= theta_p}
\(A_p\) estimated rank assigned to person p used in the rank-loss definitions
\(\mathrm{MSELR}\) mean squared error loss of ranks rank-scale loss averaged over persons
\(\mathrm{MSELP}\) mean squared error loss of percentile ranks equals MSELR divided by P squared
\(\mathrm{KS}\) Kolmogorov–Smirnov distance supremum EDF discrepancy rather than integrated EDF loss
\(Q_p\) percentile rank of person p equals R_p divided by P
\(\hat{U}_j\) jth mass point of the ISEL-optimal discrete EDF estimate = Gbar^-1((2j-1)/2K); on the trait scale, unlike Gbar itself
\(\bar{G}_K\) posterior expectation of the realized EDF a cumulative PROBABILITY at each t; this book writes Gbar_N when K = N persons
\(m_i\) number of ordered steps in polytomous item i local to the polytomous section; the dichotomous case is m_i = 1
\(x_{ni}\) person n’s score on polytomous item i, a count of completed steps Masters’ own notation, kept so the sufficiency derivation can be quoted at his locators; this book’s dichotomous response is u_{pi} and his person index n is this book’s p
\(Q\) Q-matrix: which attributes each diagnostic item requires bare and unsubscripted, unlike ch. 6 complementary response probability Q_i(theta); the two never appear in the same section
\(\Gamma\) design matrix whose columns index distinguishable latent profiles separability of Gamma decides whether individual profile proportions are identified
\(B_j\) indicator that draw j opens a new CRP cluster local to the Antoniak proof; the B_j are independent Bernoulli with P = alpha/(alpha+j-1), which is what puts E[K_J] in closed form
\(n_h\) size of occupied CRP block h the family also licenses n_1 through n_k in the partition likelihood
\(e_r\) elementary symmetric polynomial of degree r used in the finite Rasch normalization identity
\(\gamma_i\) slope-intercept item parameter, gamma_i = -lambda_i beta_i foreign notation quoted in the crosswalk; this book never reparameterizes to slope-intercept
\(\zeta_i\) slope-intercept item parameter, zeta_i = -alpha_i beta_i the same object as Paganin et al.s gamma_i under a different letter; both quoted in the crosswalk
\(\gamma\) Baker and Kim’s atypical difficulty parameter in Z = alpha theta - gamma named in the crosswalk only to warn that two different gammas circulate in the Gibbs-sampling literature
\(\lambda_p\) posterior variance of the unit parameter NOT used by this book; appears only in the crosswalk to record the collision that forced the v_p departure
\(\beta_i^{\mathrm{tmp}}\) working (uncentred) item difficulty inside the sampler, before the identification centring the fitted beta_i is beta_i^tmp minus the mean of the beta^tmp vector; ch. 26 eq-2pl-constraints
\(\beta^{\mathrm{tmp}}\) the vector of working item difficulties (appears under an overline as its mean) companion of beta_i^tmp
\(\lambda_i^{\mathrm{tmp}}\) working (uncentred) log-scale discrimination inside the sampler, before the log-mean centring log lambda_i is log lambda_i^tmp minus the mean of the log lambda^tmp vector; prior N(0.5, 0.5) before centring
\(\lambda^{\mathrm{tmp}}\) the vector of working discriminations (appears under an overline as its log-mean) companion of lambda_i^tmp
\(d_s\) draw-s inverse geometric-mean discrimination multiplier used in 2PL post-hoc normalization equals (product_i lambda_is)^(-1/I); the positive-affine slope is a_s = 1/d_s
\(q_{\mathbf u}(\theta)\) known-item 2PL conditional probability kernel for response pattern u product of Bernoulli item probabilities; integrating it against G gives pi_u(G)

Four choices in that table depart from the manuscript, and each is deliberate.

\(P\) for persons, \(N\) for the design. The manuscript uses \(N\) for the number of examinees. This book reserves \(N\) for the simulation-design sample size that the companion volume varies, and writes \(P\) for the index bound, so that a statement made across the two books is unambiguous about which quantity it means. In a realized design cell with no exclusions or missing persons, \(P=N\); \(N\) names the planned design factor and \(P\) names the number of response vectors actually analyzed.

One symbol for information. Information is \(\mathcal{J}(\theta)\) everywhere. The manuscript’s text alternates between \(\mathcal{J}\) and \(\mathcal{I}\) within a single paragraph, and the appendix drafts use \(\mathcal{I}\) exclusively; the simulation-study book uses \(\mathcal{J}\). Cross-book consistency decides it.

\(\boldsymbol\delta = (\mu_\theta, \sigma^2_\theta)^\top\), with an equals sign. The manuscript writes \(\{\mu_\theta,\sigma^2_\theta\} \in \delta\), which says the hyperparameters are elements of \(\boldsymbol\delta\) rather than being \(\boldsymbol\delta\). Reviewer 2 was right to correct it.

Posterior variance is \(v_p\), not a person-indexed lambda. Shen and Louis, whose estimators Chapter 18 and Chapter 19 develop, use the latter convention for the posterior variance of \(\theta_p\). This book carries the 2PL alongside the Rasch model, so \(\lambda_i\) is the item discrimination on most of the pages where that source symbol would appear. Rather than subscript-disambiguate a Greek letter across two hundred pages, the posterior variance becomes \(v_p\). The departure is restated at its first use in Chapter 11 and Chapter 18 and recorded in the crosswalk of Appendix B.

\(\bar w\) is a reliability coefficient, not the mean of \(w_p\). The overbar is historical notation inherited from the manuscript and Wright and Masters, not an averaging operator. This book uses \(P^{-1}\sum_p w_p\) when it means the average shrinkage weight. The two quantities coincide only under constant conditional error; Chapter 8 states the approximation and its Jensen direction.

2.3 Two models, and what separates them

The book carries the Rasch model and the two-parameter logistic model together. Writing \(\pi_{pi} = \Pr(U_{pi} = 1 \mid \theta_p, \beta_i, \lambda_i)\), the 2PL is

\[ \operatorname{logit}(\pi_{pi}) = \lambda_i(\theta_p - \beta_i), \]

and the Rasch model is the special case \(\lambda_i \equiv 1\). Almost everything in this book holds for both. Three things do not, and they are worth naming here because they organize several later chapters:

  1. Sufficiency. Under the Rasch model \(r_p\) is sufficient for \(\theta_p\). Under the 2PL weighted score \(\sum_i \lambda_i u_{pi}\) is sufficient only conditional on fixed, known discriminations. With unknown \(\boldsymbol\lambda\) it is parameter-indexed and is not an observed statistic available for Rasch-style conditioning (Chapter 3).
  2. Conditional maximum likelihood. Ordinary Rasch CML conditions on person totals to eliminate the person parameters. The 2PL has no analogous parameter-free total-score reduction, so this exact Rasch construction does not carry over (Chapter 5). This matters because Rasch CML assumes nothing about \(G\) and therefore provides a reference point for measuring the effect of a distributional assumption.
  3. Scale. The Rasch model has a location indeterminacy and no scale indeterminacy; the 2PL has both. The 2PL therefore needs two identifying constraints where the Rasch model needs one, and the second constraint pins the scale of \(G\) itself (Chapter 4, Chapter 16).

These three properties make the Rasch model useful for isolating the effect of assuming a shape for \(G\). They are substantive restrictions, not an appeal to tradition.

2.4 A distributional assumption is not the same thing as a prior

The manuscript’s exposition moves from maximum likelihood to Bayesian estimation without saying why, and in doing so leaves the impression that assuming \(\theta_p \sim N(\mu_\theta,\sigma^2_\theta)\) is something Bayesians do and frequentists do not. Reviewer 2 objected, and the objection is correct.

Joint maximum likelihood treats the \(\theta_p\) as fixed incidental parameters, and Rasch CML conditions them away; neither method needs a population distribution \(G\). Marginal maximum likelihood and hierarchical Bayes do require one. In MML the person parameters are integrated out against \(G\),

\[ L_M(\boldsymbol\beta, G \mid \mathbf{u}) = \prod_{p=1}^{P} \int L(\theta, \boldsymbol\beta \mid \mathbf{u}_p)\, dG(\theta), \tag{2.1}\]

and without that assumption this marginal likelihood is undefined. Within MML the assumption is not optional and it is not Bayesian; it is what makes that formulation work. What is distinctively Bayesian in the comparison below is the second assumption — a prior on the item parameters. Under this book’s item-centering convention it is implemented through iid Gaussian auxiliaries that are mean-centered to form \(\boldsymbol\beta\); the resulting constrained difficulties are dependent. MML does not make this prior assumption, treating \(\boldsymbol\beta\) as fixed structural parameters.

Table 2.2 sets the two side by side.

Table 2.2: The role of \(G\) in the two frameworks. Source: tables/T-frameworks.rds.
Question Frequentist (MML) Bayesian (hierarchical)
What is \(G\)? A population distribution: the law of \(\theta\) in the population persons were sampled from A prior for \(\theta_p\) — but a hierarchical one, which is why the distinction is subtle
Are the \(\theta_p\) random? Yes — random effects, integrated out Yes — assigned a prior
Are the \(\beta_i\) random? No — fixed structural parameters, no distribution assigned Yes — assigned a prior, which is the extra assumption
What is done with \(G\)’s parameters? Estimated by maximising the marginal likelihood Assigned hyperpriors and integrated over
How is \(G\) relaxed? Nonparametric MML: estimate \(G\) as a discrete or smoothed distribution Semiparametric priors: a Dirichlet process mixture over \(G\)
Does shrinkage occur? Yes — via empirical Bayes, plugging \((\hat\mu_\theta,\hat\sigma^2_\theta)\) into the same weight Yes — the same weight, with hyperparameter uncertainty carried

Read down the columns and the practical difference is smaller than the rhetorical one. The functional form of \(G\) is identical; the marginal likelihood Equation 2.1 is the same integral the Bayesian model computes. What differs is what happens next: MML maximizes over \((\boldsymbol\beta, \mu_\theta, \sigma^2_\theta)\) and reports point estimates, while the Bayesian model assigns hyperpriors and integrates. MML is the limiting case of the hierarchical model as the hyperpriors flatten and the hyperparameter posteriors concentrate.

The consequence for this book is that “assuming normality” is a modelling choice, not a Bayesian sin — and, more usefully, that the ways of relaxing it run in parallel. The frequentist route replaces the parametric \(G\) with a nonparametric maximum-likelihood estimate of the mixing distribution; the Bayesian route replaces it with a prior over distributions. Chapter 14 traces the first and Chapter 15 develops the second, and reading them as two answers to one question is the correct way to read them.

Shrinkage follows from the hierarchy, not from the framework. Under either, the estimate of \(\theta_p\) borrows from the population: with a Gaussian \(G\) the posterior mean is a weighted average of the within-person estimate and \(\mu_\theta\), with weight

\[ w_p = \frac{\sigma^2_\theta}{\sigma^2_\theta + \operatorname{se}(\hat\theta_p)^2}. \tag{2.2}\]

Frequentist empirical Bayes computes the same weight from MML point estimates. Chapter 11 derives Equation 2.2 and Chapter 12 places it in the tradition it comes from; what matters here is only that the hierarchical structure, and not the choice of framework, is what produces it.

2.5 Levels, not stages

The manuscript describes its model in terms of a “first stage” and a “second stage.” Reviewer 2 called the usage non-standard, and it is worse than non-standard: it is ambiguous, because the same word is the ordinary name for the steps of a two-step estimation procedure — calibrate the items, then score the persons. A reader cannot tell whether “second stage” names a layer of the model or a step of the fitting.

This book writes levels and reserves the word for model architecture.

Table 2.3: Level vocabulary. Source: tables/T-levels.rds.
Level Describes Why not “stage”
Level 1 — measurement The observation process conditional on the parameters: the likelihood “Stage” names a step in an estimation procedure, not a layer of the model
Level 2 — population The distribution of the Level-1 parameters — \(\theta_p \sim G\); under item centering, iid Gaussian auxiliaries induce a dependent prior on constrained \(\boldsymbol\beta\) MML followed by ability estimation is two stages of estimation within one Level-2 model
Level 3 — hyperprior The distribution of the parameters governing Level 2 The two vocabularies cut the model differently, and conflating them is what made the manuscript’s usage unreadable

Under this vocabulary, MML followed by ability estimation is two stages of estimation applied to a model with two levels, and saying so is not a quibble — Chapter 6 shows that the properties of the resulting ability estimate depend on exactly this distinction, because the second stage treats a quantity estimated in the first as though it were known.

2.6 Conventions

Numbered results. Definitions, theorems, propositions, lemmas, and corollaries are numbered by chapter and carry a provenance tag in the header:

Theorem 8.3 (Jensen gap) · restated from Lee (2026, Prop. 2)

Restated means the result is established in the cited source and reproduced here, possibly in different notation. Adapted means it is established there for a different setting and modified here, and the modification is stated. This book’s derivation means a search for prior art was run and logged and found none — a claim about the literature, which fails the same way a citation does, and which the reader should therefore treat as the least secure of the three.

Every numbered result carries a page- or theorem-level locator to its source. A result cited without one has not been checked, and the build refuses to release with any such citation outstanding.

Proof coverage is claim-level. Short arguments run inline; longer book-supplied arguments move to Appendix A; published theorems not reproduced here remain source-only at a precise locator. The Appendix A index distinguishes these from partial derivations, so a locator is never described as though it were a stand-alone proof.

Cross-references to the companion volume are spelled out — DPMirt Simulation Study, ch. 04 § 3 — rather than given as a bare pointer, so that a printed copy of this book remains usable.

2.7 Sources and provenance

The data-structure notation and the level vocabulary of Section 2.5 follow the Appendix A draft prepared for the APM revision (§§ A.2, A.4.3, A.5), which in turn follows Fox (2010) and Debelak et al. (2022) for the hierarchical framing and Gelman et al. (2013) for the level terminology. That draft carries no citation authority in this book — it is machine-generated and unreconciled — and every statement taken from it has been re-derived or checked against the source named here.

The parallel between the marginal likelihood and the hierarchical model in Section 2.4 restates Bock and Aitkin (1981) for the frequentist side and Fox (2010) for the Bayesian, and follows the Appendix D draft (§ D.5) in setting them against each other. The observation that MML is the flat-hyperprior limit of the hierarchical model is standard and is stated here without novelty claim.

The notation departures are this book’s own decisions, recorded in manifest/notation-register.csv with the collision each resolves.

Bock, R. Darrell, and Murray Aitkin. 1981. “Marginal Maximum Likelihood Estimation of Item Parameters: Application of an EM Algorithm.” Psychometrika 46 (4): 443–59. https://doi.org/10.1007/BF02293801.
Debelak, Rudolf, Carolin Strobl, and Matthew D. Zeigenfuse. 2022. An Introduction to the Rasch Model with Examples in r. Chapman; Hall/CRC.
Fox, Jean-Paul. 2010. Bayesian Item Response Modeling: Theory and Applications. Statistics for Social and Behavioral Sciences. Springer.
Gelman, Andrew, John B. Carlin, Hal S. Stern, David B. Dunson, Aki Vehtari, and Donald B. Rubin. 2013. Bayesian Data Analysis. 3rd ed. Chapman; Hall/CRC.