Testimony

I would conjecture* that qualitative testimony about the impact of social programmes is often highly sensitive (if a programme works, then people will tell you it works) but has very poor specificity (if it doesn’t work, people are unlikely to tell you so, e.g., not recognising alternative explanations of change). If this conjecture is true, then it follows that positive feedback about a programme doesn’t tell us much; however, negative feedback would be highly informative. This follows from Bayes’ rule.

Play around with the probabilities here.

Sensitivity is the probability that someone will tell you the programme works, if it does actually work.

Specificity is the probability that someone will tell you the programme doesn’t work, if it doesn’t actually work.

Prevalence is the probability that programmes of the type you are asking about actually work, e.g., 0.5 would mean for a given type of programme you think the probability it does work is the same as the probability it doesn’t.

Positive predictive value (PPV) is the probability that the programme actually does work if someone tells you it does.

Negative predictive value (NPV) is the probability that the programme doesn’t actually work if someone tells you it doesn’t.

* This conjecture is at the bottom of the evidence hierarchy, whatever the opposite of a gold standard is. Balsa wood standard?

What helps/hinders school-based humanistic therapy?

Interesting thematic analysis (Cooper et al., 2015) of what helped and hindered in school-based humanistic therapy, as part of the large ETHOS RCT. The approach the researchers took started with open questions and coding inductively, then concluded with closed questions and deductive coding based on what the theory suggested might help or not. All the usual helpful things were there, like:

“[The therapist] wasn’t, like, doing something else. She was sat there all the time just like listening to me.”

“They just talked to me like I was a human being not like, ‘Oh, you know, what teenagers are like…'”

“It was like they were taking in what I was actually saying and then the questions afterwards were based upon that.”

I was struck by the number of young people who mentioned the therapist providing advice, given that the humanistic approach is supposed to be non-directive (there is some discussion of this). Sometimes this advice was welcome and young people who didn’t receive any advice indicated that they didn’t see the point of therapy without it.

“I used to have an argument with this girl…and the therapist said to just leave it because she’s just going to– she wants attention and she wants a reaction out of me. She [the therapist] was just like, “Leave it. She can say whatever she wants and then she’ll stop it herself,” and now that girl doesn’t even say anything to me because I don’t even say anything to her.”

Advice/guidance wasn’t always experienced as helpful; the examples given seem to be all “technique”:

“[The therapist] would do breathing exercises and imagine things, and I didn’t like that. I said, “I didn’t like that,” and [the therapist] was like, “Well, let’s just try.” “I don’t want to do it.” “Well, let’s just try.” [The therapist] was kind of persistent on something that I didn’t want to do.”

There’s also a collection of things some therapists did which were seen as particularly unhelpful, e.g.,

“She didn’t have a lot to say. She would just sit there and stare at me for sometimes three minutes at a time. Literally it was three minutes. She’d just sit there, and I’d be looking around the room, and every time I looked back at her, she’s just looking at me. She wouldn’t say anything.”

(Reminds me of Levitt’s 2001 thematic analysis of the twenty or so different kinds of silence!)

Wonder if anyone has tried to crowdsource these helpful/hindering ideas at scale.

References

Cooper, M., Smith, S., Sumner, A. L., Eilenberg, J., Childs-Fegredo, J., Kelly, S., Subramanian, P., Holmes, J., Barkham, M., Bower, P., Cromarty, K., Duncan, C., Hughes, S., Pearce, P., Rameswari, T., Ryan, G., Saxon, D., & Stafford, M. R. (2025). Humanistic Therapy for Young People: Client-Perceived Helpful Aspects, Hindering Aspects, and Processes of Change. Journal of Child and Family Studies.

Levitt, H. M. (2001). Sounds of Silence in Psychotherapy: The Categorization of Clients’ Pauses. Psychotherapy Research, 11(3), 295–309.

Choosing a sample size for a thematic analysis

About a decade ago, Henry Potts and I (Fugard & Potts, 2015) developed a method to guide decisions about sample size in thematic analyses (there’s also an app that does the calculations).

The core idea is straightforward: the rarer the phenomenon you want to investigate, the larger the sample you’ll need -after using purposive sampling to recruit the most relevant participants. For example, if you’re interested in the experiences of people engaging with therapy for depression, you would begin with participants who have had therapy for depression, not the general population, which reduces the total needed.

The paper outlining the approach has been cited over a thousand times, though only a small number of those citations apply the method as we intended. Here are five of them:

  • “Framing the necessary sample size in terms of the likelihood of capturing important ideas, a sample size of at least 15 per group would have a 90% probability for capturing ideas held by 25% of the population.” (Weller et al., 2017, p. 2.) Would add, that’s to ensure two participants provide material relevant to those themes.
  • “An estimation of the sample size was done prior to the second selection to ensure that the final number of participants would be large enough to yield a rich data set. Twenty-one interviews would be needed according to Fugard [and Potts]’s sample size calculation method for 80% power. This calculation assumed: a lowest prevalence of 60% of a theme worth discovering, 90% of the informants having something to say about the theme, and that the theme should be recognised at least ten times in the data.” (Malmborg et al., 2020, p. 170)
  • “In this study of the reasons for open online student dropout, appropriate sample size was determined based on Fugard and Potts’ (2015) thematic analysis sample size tool, which highlights the required sample size as a function of anticipated theme prevalence in a population. In line with this tool, a sample of 200 has a 99% probability of detecting five theme instances for a theme prevalent in 6% of the population; thus, 200 was set as the minimum sample target for this study.” (Greenland & Moore, 2022, p. 652)
  • “To evaluate the saturation of the codes, we used the Fugard and Potts method to predict saturation based on probability theory. This approach was appropriate for our data set, given our large, random sample of reviews and our predominantly deductive approach to data analysis. Our data set provided >80% power to identify 5 instances of themes mentioned by 1% of the population. We chose a cutoff of 1% to reflect the shallow nature of this data set, assuming that not all who experienced a code would describe it in their review, and 5 instances because this was typically the number of observations required to achieve repetition of content within the codes.” (Polhemus et al., 2022, p. 4.) To have exactly 80% power for prevalence 1% and five instances, they would have needed a sample size of 671.
  • “According to the calculation method of Fugard and Potts, the target study cohort size of 28 patients would provide 90% power to detect any theme of interest in at least 1 interview if the true prevalence of the patient experience captured by that theme was between 5 and 10% in the underlying PH1 population represented by the study cohort.” (Danese et al., 2023, p. 3) (Actually prevalence 7.9%.)

From these examples, it’s clear that the richness of the material and time taken to analyse it was an important and sensible constraint. For example, we would expect smaller samples for long interviews and larger samples for a sentence of free text in a web-based survey. Contrary to some arguments we have seen in the literature, it does also apply to reflexive thematic analysis too. Please do pop me an email if you’d like to give it a go but feel a bit stuck. I’d be happy to help.

References

Danese, D., Goss, D., Romano, C., & Gupta, C. (2023). Qualitative assessment of the patient experience of primary hyperoxaluria type 1: An observational study. BMC Nephrology, 24(1), 319.

Fugard, A. J. B. & Potts, H. W. W. (2015). Supporting thinking on sample sizes for thematic analyses: A quantitative toolInternational Journal of Social Research Methodology, 18, 669–684. (There’s an app for that.)

Greenland, S. J., & Moore, C. (2022). Large qualitative sample and thematic analysis to redefine student dropout and retention strategy in open online education. British Journal of Educational Technology, 53, 647–667.

Malmborg, A., Brynte, L., Falk, G., Brynhildsen, J., Hammar, M., & Berterö, C. (2020). Sexual function changes attributed to hormonal contraception use – a qualitative study of women experiencing negative effects. The European Journal of Contraception & Reproductive Health Care, 25(3), 169–175.

Polhemus, A., Simblett, S., Dawe-Lane, E., Gilpin, G., Elliott, B., Jilka, S., Novak, J., Nica, R. I., Temesi, G., & Wykes, T. (2022). Health Tracking via Mobile Apps for Depression Self-management: Qualitative Content Analysis of User Reviews. JMIR Human Factors, 9(4), e40133.

Weller, S. C., Baer, R., Nash, A., & Perez, N. (2017). Discovering successful strategies for diabetic self-management: A qualitative comparative study. BMJ Open Diabetes Research and Care, 5, e000349.

Qualitative research

Qualitative research isn’t synonymous with interview or focus group. Here’s an example from neuroscience: Gray’s (1959) classification of the qualitatively different kinds of synapse, using electron microscopy.

Here is Gray’s summary description (p. 430).

“In type 1 synapses a large percentage of the length of the apposed membranes shows increased thickness and density. The post-synaptic thickening is more pronounced than the pre-synaptic thickening. These thickened regions of the membranes lie farther apart than where the apposed membranes are unthickened and in the extracellular region between the thickened membranes an intermediate band of material can be seen.”

“In type 2 synapses the percentage length of thickening is small, the pre- and post-synaptic thickenings are of similar dimensions, the intermediate band is not clearly visible and the membrane spacings at these regions differ little from the non-thickened regions.”

Figure 10 below shows an example of the material Gray was analysing:

Gray, E. G. (1959). Axo-somatic and axo-dendritic synapses of the cerebral cortex. Journal of Anatomy, 93, 420–433.