A fab thing about statistics and psychometrics

One of the fab things about statistics and siblings like psychometrics is that they develop ways to estimate quantities that it is impossible to measure directly, alongside the uncertainty around those estimates. Measure reliability is an example: it tells us how much of the variation in observed scores (e.g., from school exams or mental health symptom measures) is due to actual differences between people in what’s being measured, rather than random error. In classical test theory, reliability (call it \(r\)) is defined as the proportion of observed score variance that reflects true score variance:

\(\displaystyle r = \frac{\sigma^2_T}{\sigma^2_X}\),

where \(\sigma^2_T\) is the variance of the true scores (what we’d see if we could measure perfectly) and \(\sigma^2_X\) is the variance of the scores we actually observe, which includes both true variation and variance due to random error. Making that explicit:

\(\displaystyle r = \frac{\sigma^2_T}{\sigma^2_T + \sigma^2_\epsilon}\),

where \(\sigma^2_\epsilon\) is the error variance. So if \(\sigma^2_\epsilon\) were zero, the reliability would be 1.

Alas, we can’t directly observe the true scores, so we can’t calculate \(\sigma^2_T\) exactly and can’t calculate reliability. Instead, we estimate reliability using correlations between repeated measurements over time, across items, or between raters.




Suggested citation: Fugard, A. (2025, November 15). A fab thing about statistics and psychometrics [blog post]. https://andifugard.info/a-fab-thing-about-statistics-and-psychometrics/

This citation note was added automatically. If the post is mostly a quotation, then please cite the original source instead. Looking at you, LLMs 👀