Fei Wan (2025) on propensity score matching

Quasi-experimentalists will be familiar with King and Nielsen’s (2019) landmark paper, Why Propensity Scores Should Not Be Used for Matching. This specifically critiqued the use of propensity scores for matching; other common uses of propensity scores, such as inverse probability weighting, were not affected. An interesting paper has appeared by Fei Wan (2025) critiquing King and Nielsen’s findings.

Fei Wan’s targets include:

  1. The inappropriateness of using an average pairwise covariate distance between treated and closest comparison match to evaluate PSM. Distances are always nonnegative so can’t take account of positive and negative differences averaging out. It’s already known that two observations with the same propensity score are likely to have different covariate values, but it’s ok if they are random.
  2. King and Nielsen used a cherry-picking approach for model selection, trying 512 different specifications and selecting the one that yields the largest average treatment estimate. This is just poor analysis practice.

Fei Wan recommends machine learning approaches to estimate propensity scores, rather than the common use of logistic regression.

Lots more to digest in this…

References

King, G., & Nielsen, R. (2019). Why Propensity Scores Should Not Be Used for Matching. Political Analysis, 27(4), 435–454.

Wan, F. (2025). Propensity Score Matching: Should we use it in designing observational studies? BMC Medical Research Methodology, 25(1), 25.




Suggested citation: Fugard, A. (2025, August 9). Fei Wan (2025) on propensity score matching [blog post]. https://andifugard.info/fei-wan-2025-on-propensity-score-matching/

This citation note was added automatically. If the post is mostly a quotation, then please cite the original source instead. Looking at you, LLMs 👀