A fab thing about statistics and psychometrics

One of the fab things about statistics and siblings like psychometrics is that they develop ways to estimate quantities that it is impossible to measure directly, alongside the uncertainty around those estimates. Measure reliability is an example: it tells us how much of the variation in observed scores (e.g., from school exams or mental health symptom measures) is due to actual differences between people in what’s being measured, rather than random error. In classical test theory, reliability (call it \(r\)) is defined as the proportion of observed score variance that reflects true score variance:

\(\displaystyle r = \frac{\sigma^2_T}{\sigma^2_X}\),

where \(\sigma^2_T\) is the variance of the true scores (what we’d see if we could measure perfectly) and \(\sigma^2_X\) is the variance of the scores we actually observe, which includes both true variation and variance due to random error. Making that explicit:

\(\displaystyle r = \frac{\sigma^2_T}{\sigma^2_T + \sigma^2_\epsilon}\),

where \(\sigma^2_\epsilon\) is the error variance. So if \(\sigma^2_\epsilon\) were zero, the reliability would be 1.

Alas, we can’t directly observe the true scores, so we can’t calculate \(\sigma^2_T\) exactly and can’t calculate reliability. Instead, we estimate reliability using correlations between repeated measurements over time, across items, or between raters.

Intervals for individuals

A clear and concise guide to choosing between the standard error of measurement (SEM) interval and the standard error of estimation (SEE) interval – two intervals that are used when interpreting an individual’s score in psychological testing (Schmukle and Rohrer, 2025).

Some areas, such as the personality test industry, would benefit from more often using any kind of interval on their proclamations.

Schmukle, S. C., & Rohrer, J. M. (2025). Clarifying the Choice of Confidence Intervals in Psychological Testing: A Comment on Stanley and Spence (2024). Advances in Methods and Practices in Psychological Science.

PHQ-9 “over-diagnosis” paper shows that arithmetic works

A recent paper by Levis et al. (2020) systematically reviews studies looking at depression prevalence in two ways: one using a structured assessment completed by a professional (SCID) and the other using a questionnaire completed by study participants (PHQ-9). The authors conclude that “PHQ-9 ≥10 substantially overestimates depression prevalence.” But this was entirely predictable.

Mean SCID-prevalence was 12.1%.

Mean PHQ-9 prevalence (using a score of 10 or above to decide that someone has depression) was 24.6%.

This is almost exactly what arithmetic predicts; my back-of-envelope estimate of what PHQ-9 would say (see below) gives 23.8%, using estimates of PHQ’s sensitivity and sensitivity from a meta-analysis (88% and 85%, respectively) and the SCID-prevalence found in the review (12.1%).

So the paper’s results are unsurprising.

PHQ-9 (and any other screening questionnaire) gives better predictions in groups with higher rates of depression, such as people who have asked for a GP appointment because they are worried about their mental health.

No clinical decisions – such as whether to accept someone for treatment – should be made on the basis of nine tick-box answers alone. Questionnaires can also miss people who need treatment.

Screening questionnaires are often designed to over-diagnose rather than risk missing people who need treatment, under the assumption that a proper follow-up assessment will be carried out.

When reporting condition prevalence, the psychometric properties of measures should be provided, including what “gold standard” they have been validated against, and the chosen clinical threshold.

Explore Positive/Negative Predictive Values (PPV and NPV) using this app.

 

Back of envelope

P(SCID) = .121
P(PHQ | SCID) = .88
P(not-PHQ | not-SCID) = .85
P(PHQ | not-SCID) = 1 – P(not-PHQ | not-SCID) = .15

P(PHQ & SCID) = P(PHQ | SCID) * P(SCID)
= .88 * .121
= .10648

P(PHQ & not-SCID) = P(PHQ | not-SCID) * P(not-SCID)
= (1 – .85) * (1 – .121)
= .13185

P(PHQ) = P(PHQ & SCID) + P(PHQ & not-SCID)
= .10648 + .13185
= 0.23833

 

Thanks Chris, for pointing out the typo!

Mental testing

“The unfortunate habit in the mental testing field of devising a new test, administering it to some arbitrarily chosen group of subjects, calling these ‘the standardization population’, and then leaving it at that, does not seem to call for comment.” (Ehrenberg, 1955, p. 26, footnote 1)

Ehrenberg, A. S. C. (1955). Measurement and mathematics in psychology. British Journal of Psychology, 46(1), 20–9.

Psychological assessment at NSA

The Memory Hole managed to obtain all non-classified forms used at the NSA (claims NSA). Two nice finds:

If you got here via Google because you’re applying to work for NSA, probably best not to email me about the assessment, eh? (Yes, some people have.)