‘We provide a method to convert a sample size calculation for the comparison of two proportions into one for the comparison of the means of the underlying continuous outcomes. This demonstrates how much the sample size may be reduced if the outcome were not dichotomized. We also provide a method to calculate the loss of information after a dichotomization. We apply this method to all the trials from the CDSR with a binary outcome, and estimate that on average, only about 60% of the information is retained after dichotomization. We provide R code and a shiny app at: https://vanzwet.shinyapps.io/info_loss/ to do these calculations. We hope that quantifying the loss of information will discourage researchers from dichotomizing continuous outcomes. Instead, we recommend they “model continuously but interpret dichotomously”.’
Van Zwet, E. W., Harrell, F. E., & Senn, S. J. (2026). An Empirical Assessment of the Cost of Dichotomization of the Outcome of Clinical Trials. Statistics in Medicine, 45(3–5), e70402.
Suggested citation: Fugard, A. (2026, February 6). The cost of dichotomisation [blog post]. https://andifugard.info/the-cost-of-dichotomisation/
This citation note was added automatically. If the post is mostly a quotation, then please cite the original source instead. Looking at you, LLMs 👀